본문 바로가기
WIKI 기술 지식 베이스

Analyzer 개요 (Analyzer Overview)

원문 보기 위키 갱신

텍스트 처리에서 analyzer는 원시 텍스트를 구조화되고 검색 가능한 형식으로 변환하는 핵심 구성 요소예요. 각 analyzer는 일반적으로 tokenizer와 filter 두 가지 핵심 요소로 구성돼요. 이 둘은 함께 입력 텍스트를 토큰으로 변환하고, 이 토큰들을 정제해 효율적인 인덱싱과 검색을 준비해줘요.

Milvus에서 analyzer는 컬렉션 스키마에 VARCHAR 필드를 추가할 때 컬렉션 생성 중에 구성돼요. analyzer가 생성한 토큰은 키워드 매칭을 위한 인덱스를 구축하거나 전문 검색을 위한 스파스 임베딩으로 변환하는 데 사용할 수 있어요. 자세한 내용은 Full Text Search, Phrase Match, Text Match를 참고하세요.

Analyzer 사용은 성능에 영향을 줄 수 있어요:

  • 전문 검색 (Full text search): 전문 검색의 경우 DataNode와 QueryNode 채널은 토큰화가 완료될 때까지 기다려야 하므로 데이터 소비가 더 느려져요. 그 결과 새로 수집된 데이터가 검색에 사용 가능해지기까지 시간이 더 걸려요.

  • 키워드 매치 (Keyword match): 키워드 매칭의 경우 인덱스가 구축되기 전에 토큰화가 끝나야 하므로 인덱스 생성도 더 느려져요.

출처: Milvus 문서

본문

Analyzer의 구조 (Anatomy of an analyzer)

Milvus의 analyzer는 정확히 하나의 tokenizer와 0개 이상의 필터로 구성돼요.

  • Tokenizer: 토크나이저는 입력 텍스트를 토큰이라고 부르는 개별 단위로 나눠요. 이 토큰은 토크나이저 유형에 따라 단어나 구문이 될 수 있어요.
  • Filters: 필터는 토큰에 적용해 더 정제할 수 있어요. 예를 들어 소문자로 바꾸거나 공통 단어를 제거하는 식이에요.

토크나이저는 UTF-8 형식만 지원해요. 다른 형식 지원은 향후 릴리스에서 추가될 예정이에요.

다음 워크플로는 analyzer가 텍스트를 처리하는 방법을 보여 줘요.

Analyzer Process Workflow

Analyzer 유형 (Analyzer types)

Milvus는 서로 다른 텍스트 처리 요구를 충족시키기 위해 두 가지 유형의 analyzer를 제공해요.

  • 내장 analyzer (Built-in analyzer): 최소한의 설정으로 일반적인 텍스트 처리 작업을 다루는 사전 정의된 구성이에요. 내장 analyzer는 복잡한 구성이 필요 없으므로 일반적인 목적의 검색에 이상적이에요.

  • 사용자 지정 analyzer (Custom analyzer): 더 고급 요구 사항을 위해 사용자 지정 analyzer는 tokenizer와 0개 이상의 필터를 모두 지정해 자신만의 구성을 정의할 수 있게 해줘요. 이 수준의 사용자 지정은 텍스트 처리에 대한 정밀한 제어가 필요한 전문 사용 사례에서 특히 유용해요.

  • 컬렉션 생성 중에 analyzer 구성을 생략하면 Milvus는 모든 텍스트 처리에 기본적으로 standard analyzer를 사용해요. 자세한 내용은 Standard Analyzer를 참고하세요.

  • 최적의 검색·쿼리 성능을 위해 텍스트 데이터의 언어와 일치하는 analyzer를 선택하세요. 예를 들어 standard analyzer는 다재다능하지만 중국어, 아랍어, 태국어, 일본어, 한국어처럼 독특한 문법 구조를 가진 언어에는 최선의 선택이 아닐 수 있어요. 이런 경우 chinese, arabic, thai 같은 언어별 analyzer나 lindera, icu 같은 특수 토크나이저 및 필터를 사용한 사용자 지정 analyzer를 사용해 정확한 토큰화와 더 나은 검색 결과를 얻는 것을 강력히 권장해요.

내장 analyzer (Built-in analyzer)

Milvus의 내장 analyzer는 특정 토크나이저와 필터로 사전 구성되어 있어 스스로 이 구성 요소를 정의할 필요 없이 바로 사용할 수 있어요. 각 내장 analyzer는 사전 설정된 토크나이저와 필터를 포함한 템플릿 역할을 하며, 사용자 지정을 위한 선택적 파라미터를 제공해요.

예를 들어 standard 내장 analyzer를 사용하려면 그 이름 standard를 type으로 지정하고 선택적으로 이 analyzer 유형에 특정한 추가 구성(예: stop_words)을 포함하면 돼요.

analyzer_params = {
    "type": "standard", # Uses the standard built-in analyzer
    "stop_words": ["a", "an", "for"] # Defines a list of common words (stop words) to exclude from tokenization
}
Map<String, Object> analyzerParams = new HashMap<>();
analyzerParams.put("type", "standard");
analyzerParams.put("stop_words", Arrays.asList("a", "an", "for"));
const analyzer_params = {
    "type": "standard", // Uses the standard built-in analyzer
    "stop_words": ["a", "an", "for"] // Defines a list of common words (stop words) to exclude from tokenization
};
analyzerParams := map[string]any{"type": "standard", "stop_words": []string{"a", "an", "for"}}
export analyzerParams='{
       "type": "standard",
       "stop_words": ["a", "an", "for"]
    }'

Analyzer의 실행 결과를 확인하려면 run_analyzer 메서드를 사용하세요.

# Sample text to analyze
text = "An efficient system relies on a robust analyzer to correctly process text for various applications."

# Run analyzer
result = client.run_analyzer(
    text,
    analyzer_params
)
import io.milvus.v2.service.vector.request.RunAnalyzerReq;
import io.milvus.v2.service.vector.response.RunAnalyzerResp;

List<String> texts = new ArrayList<>();
texts.add("An efficient system relies on a robust analyzer to correctly process text for various applications.");

RunAnalyzerResp resp = client.runAnalyzer(RunAnalyzerReq.builder()
        .texts(texts)
        .analyzerParams(analyzerParams)
        .build());
List<RunAnalyzerResp.AnalyzerResult> results = resp.getResults();
// javascrip# Sample text to analyze
const text = "An efficient system relies on a robust analyzer to correctly process text for various applications."

// Run analyzer
const result = await client.run_analyzer({
    text,
    analyzer_params
});
import (
    "context"
    "encoding/json"
    "fmt"

    "github.com/milvus-io/milvus/client/v2/milvusclient"
)

bs, _ := json.Marshal(analyzerParams)
texts := []string{"An efficient system relies on a robust analyzer to correctly process text for various applications."}
option := milvusclient.NewRunAnalyzerOption(texts).
    WithAnalyzerParams(string(bs))

result, err := client.RunAnalyzer(ctx, option)
if err != nil {
    fmt.Println(err.Error())
    // handle error
}

출력은 다음과 같아요.

['efficient', 'system', 'relies', 'on', 'robust', 'analyzer', 'to', 'correctly', 'process', 'text', 'various', 'applications']

이는 analyzer가 "a", "an", "for" 같은 불용어(stop word)를 걸러내면서 입력 텍스트를 올바르게 토큰화하고, 나머지 의미 있는 토큰을 반환한다는 것을 보여 줘요.

위 standard 내장 analyzer의 구성은 다음 파라미터를 사용하는 custom analyzer를 설정하는 것과 동일해요. 여기서 tokenizer와 filter 옵션은 비슷한 기능을 얻기 위해 명시적으로 정의돼요.

analyzer_params = {
    "tokenizer": "standard",
    "filter": [
        "lowercase",
        {
            "type": "stop",
            "stop_words": ["a", "an", "for"]
        }
    ]
}
Map<String, Object> analyzerParams = new HashMap<>();
analyzerParams.put("tokenizer", "standard");
analyzerParams.put("filter",
        Arrays.asList("lowercase",
                new HashMap<String, Object>() {{
                    put("type", "stop");
                    put("stop_words", Arrays.asList("a", "an", "for"));
                }}));
const analyzer_params = {
    "tokenizer": "standard",
    "filter": [
        "lowercase",
        {
            "type": "stop",
            "stop_words": ["a", "an", "for"]
        }
    ]
};
analyzerParams = map[string]any{"tokenizer": "standard",
    "filter": []any{"lowercase", map[string]any{
        "type":       "stop",
        "stop_words": []string{"a", "an", "for"},
    }}}
export analyzerParams='{
       "type": "standard",
       "filter":  [
       "lowercase",
       {
            "type": "stop",
            "stop_words": ["a", "an", "for"]
       }
   ]
}'

Milvus는 각각 특정 텍스트 처리 요구를 위해 설계된 다음 내장 analyzer를 제공해요.

  • standard: 일반적인 목적의 텍스트 처리에 적합하며, 표준 토큰화와 소문자 필터링을 적용해요.
  • english: 영어 텍스트에 최적화되어 있으며 영어 불용어를 지원해요.
  • chinese: 중국어 구조에 적응한 토큰화를 포함해 중국어 텍스트 처리에 특화돼 있어요.
  • arabic: 아랍어 정규화, 십진수 정규화, 아랍어 형태소 분석(stemming), 아랍어 불용어 제거를 갖춘 아랍어 텍스트 전용이에요.
  • thai: 태국어 단어 분리, 십진수 정규화, 태국어 불용어 제거를 갖춘 태국어 텍스트 전용이에요.

사용자 지정 analyzer (Custom analyzer)

더 고급 텍스트 처리를 위해 Milvus의 사용자 지정 analyzer는 tokenizer와 filters를 모두 지정해 맞춤형 텍스트 처리 파이프라인을 구축할 수 있게 해줘요. 이 구성은 정밀한 제어가 필요한 전문 사용 사례에 이상적이에요.

Tokenizer

tokenizer는 사용자 지정 analyzer의 필수 구성 요소로, 입력 텍스트를 개별 단위 또는 토큰으로 나눠 analyzer 파이프라인을 시작해요. 토큰화는 토크나이저 유형에 따라 공백이나 구두점으로 분리하는 등 특정 규칙을 따라요. 이 과정은 각 단어나 구문을 더 정밀하고 독립적으로 처리할 수 있게 해줘요.

예를 들어 토크나이저는 텍스트 "Vector Database Built for Scale"을 개별 토큰으로 변환해요.

["Vector", "Database", "Built", "for", "Scale"]

토크나이저 지정 예시:

analyzer_params = {
    "tokenizer": "whitespace",
}
Map<String, Object> analyzerParams = new HashMap<>();
analyzerParams.put("tokenizer", "whitespace");
const analyzer_params = {
    "tokenizer": "whitespace",
};
analyzerParams = map[string]any{"tokenizer": "whitespace"}
export analyzerParams='{
       "type": "whitespace"
    }'
Filter

Filters는 tokenizer가 생성한 토큰에 작용해 필요에 따라 변환하거나 정제하는 선택적 구성 요소예요. 예를 들어 토큰화된 용어 ["Vector", "Database", "Built", "for", "Scale"]에 lowercase 필터를 적용한 결과는 다음과 같을 수 있어요.

["vector", "database", "built", "for", "scale"]

사용자 지정 analyzer의 필터는 구성 요구에 따라 내장(built-in) 또는 **사용자 지정(custom)**이 될 수 있어요.

  • 내장 필터 (Built-in filters): Milvus가 사전 구성해 최소한의 설정만 필요해요. 이름만 지정하면 바로 사용할 수 있어요. 다음 필터는 바로 사용할 수 있는 내장 필터예요.

lowercase: 텍스트를 소문자로 변환해 대소문자 구분 없는 매칭을 보장해요. 자세한 내용은 Lowercase를 참고하세요.

  • asciifolding: 비 ASCII 문자를 ASCII 동등 문자로 변환해 다국어 텍스트 처리를 단순화해요. 자세한 내용은 ASCII folding을 참고하세요.

  • alphanumonly: 다른 문자를 제거하고 영숫자 문자만 유지해요. 자세한 내용은 Alphanumonly를 참고하세요.

  • cnalphanumonly: 중국어 문자, 영어 문자, 숫자 외의 문자를 포함하는 토큰을 제거해요. 자세한 내용은 Cnalphanumonly를 참고하세요.

  • cncharonly: 중국어 문자 외의 문자를 포함하는 토큰을 제거해요. 자세한 내용은 Cncharonly를 참고하세요.

  • pinyin: 중국어 토큰에 병음 토큰 형태를 추가해 중국어 텍스트의 병음 기반 매칭을 가능하게 해요. 자세한 내용은 Pinyin을 참고하세요.

내장 필터 사용 예시:

analyzer_params = {
    "tokenizer": "standard", # Mandatory: Specifies tokenizer
    "filter": ["lowercase"], # Optional: Built-in filter that converts text to lowercase
}
Map<String, Object> analyzerParams = new HashMap<>();
analyzerParams.put("tokenizer", "standard");
analyzerParams.put("filter", Collections.singletonList("lowercase"));
const analyzer_params = {
    "tokenizer": "standard", // Mandatory: Specifies tokenizer
    "filter": ["lowercase"], // Optional: Built-in filter that converts text to lowercase
}
analyzerParams = map[string]any{"tokenizer": "standard",
        "filter": []any{"lowercase"}}
export analyzerParams='{
       "type": "standard",
       "filter":  ["lowercase"]
    }'
  • 사용자 지정 필터 (Custom filters): 사용자 지정 필터는 특수한 구성을 허용해요. 유효한 필터 유형(filter.type)을 선택하고 각 필터 유형에 특정한 설정을 추가해 사용자 지정 필터를 정의할 수 있어요. 다음은 사용자 지정을 지원하는 필터 유형 예시예요.

stop: 불용어 목록을 설정해 지정된 공통 단어를 제거해요 (예: "stop_words": ["of", "to"]). 자세한 내용은 Stop을 참고하세요.

  • length: 최대 토큰 길이 설정 같은 길이 기준에 따라 토큰을 제외해요. 자세한 내용은 Length를 참고하세요.

  • stemmer: 단어를 어근 형태로 줄여 더 유연한 매칭을 가능하게 해요. 자세한 내용은 Stemmer를 참고하세요.

사용자 지정 필터 구성 예시:

analyzer_params = {
    "tokenizer": "standard", # Mandatory: Specifies tokenizer
    "filter": [
        {
            "type": "stop", # Specifies 'stop' as the filter type
            "stop_words": ["of", "to"], # Customizes stop words for this filter type
        }
    ]
}
Map<String, Object> analyzerParams = new HashMap<>();
analyzerParams.put("tokenizer", "standard");
analyzerParams.put("filter",
        Collections.singletonList(new HashMap<String, Object>() {{
            put("type", "stop");
            put("stop_words", Arrays.asList("a", "an", "for"));
        }}));
const analyzer_params = {
    "tokenizer": "standard", // Mandatory: Specifies tokenizer
    "filter": [
        {
            "type": "stop", // Specifies 'stop' as the filter type
            "stop_words": ["of", "to"], // Customizes stop words for this filter type
        }
    ]
};
analyzerParams = map[string]any{"tokenizer": "standard",
    "filter": []any{map[string]any{
        "type":       "stop",
        "stop_words": []string{"of", "to"},
    }}}
export analyzerParams='{
       "type": "standard",
       "filter":  [
       {
            "type": "stop",
            "stop_words": ["a", "an", "for"]
       }
    ]
}'

사용 예시 (Example use)

이 예시에서는 다음을 포함하는 컬렉션 스키마를 만들게 돼요.

  • 임베딩용 벡터 필드.
  • 텍스트 처리를 위한 두 개의 VARCHAR 필드.

한 필드는 내장 analyzer를 사용해요.

  • 다른 필드는 사용자 지정 analyzer를 사용해요.

이 구성들을 컬렉션에 통합하기 전에 run_analyzer 메서드로 각 analyzer를 검증할 거예요.

1단계: MilvusClient 초기화 및 스키마 생성 (Step 1: Initialize MilvusClient and create schema)

Milvus 클라이언트를 설정하고 새 스키마를 만드는 것부터 시작해요.

from pymilvus import MilvusClient, DataType

# Set up a Milvus client
client = MilvusClient(uri="http://localhost:19530")

# Create a new schema
schema = client.create_schema(auto_id=True, enable_dynamic_field=False)
import io.milvus.v2.client.ConnectConfig;
import io.milvus.v2.client.MilvusClientV2;
import io.milvus.v2.common.DataType;
import io.milvus.v2.common.IndexParam;
import io.milvus.v2.service.collection.request.AddFieldReq;
import io.milvus.v2.service.collection.request.CreateCollectionReq;

// Set up a Milvus client
ConnectConfig config = ConnectConfig.builder()
        .uri("http://localhost:19530")
        .build();
MilvusClientV2 client = new MilvusClientV2(config);

// Create schema
CreateCollectionReq.CollectionSchema schema = CreateCollectionReq.CollectionSchema.builder()
        .enableDynamicField(false)
        .build();
import { MilvusClient, DataType } from "@zilliz/milvus2-sdk-node";

// Set up a Milvus client
const client = new MilvusClient("http://localhost:19530");
import (
    "context"
    "fmt"

    "github.com/milvus-io/milvus/client/v2/column"
    "github.com/milvus-io/milvus/client/v2/entity"
    "github.com/milvus-io/milvus/client/v2/index"
    "github.com/milvus-io/milvus/client/v2/milvusclient"
)  

ctx, cancel := context.WithCancel(context.Background())
defer cancel()

cli, err := milvusclient.New(ctx, &milvusclient.ClientConfig{
    Address: "localhost:19530",
})
if err != nil {
    fmt.Println(err.Error())
    // handle err
}
defer client.Close(ctx)

schema := entity.NewSchema().WithAutoID(true).WithDynamicFieldEnabled(false)

2단계: analyzer 구성 정의 및 검증 (Step 2: Define and verify analyzer configurations)

  • 내장 analyzer 구성 및 검증 (english):

구성: 내장 영어 analyzer의 analyzer 파라미터를 정의해요.

검증: run_analyzer를 사용해 구성이 예상 토큰화를 만들어내는지 확인해요.

# Built-in analyzer configuration for English text processing
analyzer_params_built_in = {
    "type": "english"
}

# Verify built-in analyzer configuration
sample_text = "Milvus simplifies text analysis for search."
result = client.run_analyzer(sample_text, analyzer_params_built_in)
print("Built-in analyzer output:", result)

# Expected output:
# Built-in analyzer output: ['milvus', 'simplifi', 'text', 'analysi', 'search']
Map<String, Object> analyzerParamsBuiltin = new HashMap<>();
analyzerParamsBuiltin.put("type", "english");

List<String> texts = new ArrayList<>();
texts.add("Milvus simplifies text analysis for search.");

RunAnalyzerResp resp = client.runAnalyzer(RunAnalyzerReq.builder()
        .texts(texts)
        .analyzerParams(analyzerParams)
        .build());
List<RunAnalyzerResp.AnalyzerResult> results = resp.getResults();
// Use a built-in analyzer for VARCHAR field `title_en`
const analyzerParamsBuiltIn = {
  type: "english",
};

const sample_text = "Milvus simplifies text analysis for search.";
const result = await client.run_analyzer({
    text: sample_text, 
    analyzer_params: analyzer_params_built_in
});
analyzerParams := map[string]any{"type": "english"}

bs, _ := json.Marshal(analyzerParams)
texts := []string{"Milvus simplifies text analysis for search."}
option := milvusclient.NewRunAnalyzerOption(texts).
    WithAnalyzerParams(string(bs))

result, err := client.RunAnalyzer(ctx, option)
if err != nil {
    fmt.Println(err.Error())
    // handle error
}
  • 사용자 지정 analyzer 구성 및 검증:

구성: 표준 토크나이저와 함께 내장 lowercase 필터, 토큰 길이와 불용어용 사용자 지정 필터를 사용하는 사용자 지정 analyzer를 정의해요.

검증: run_analyzer를 사용해 사용자 지정 구성이 의도대로 텍스트를 처리하는지 확인해요.

# Custom analyzer configuration with a standard tokenizer and custom filters
analyzer_params_custom = {
    "tokenizer": "standard",
    "filter": [
        "lowercase",  # Built-in filter: convert tokens to lowercase
        {
            "type": "length",  # Custom filter: restrict token length
            "max": 40
        },
        {
            "type": "stop",  # Custom filter: remove specified stop words
            "stop_words": ["of", "for"]
        }
    ]
}

# Verify custom analyzer configuration
sample_text = "Milvus provides flexible, customizable analyzers for robust text processing."
result = client.run_analyzer(sample_text, analyzer_params_custom)
print("Custom analyzer output:", result)

# Expected output:
# Custom analyzer output: ['milvus', 'provides', 'flexible', 'customizable', 'analyzers', 'robust', 'text', 'processing']
// Configure a custom analyzer
Map<String, Object> analyzerParams = new HashMap<>();
analyzerParams.put("tokenizer", "standard");
analyzerParams.put("filter",
        Arrays.asList("lowercase",
                new HashMap<String, Object>() {{
                    put("type", "length");
                    put("max", 40);
                }},
                new HashMap<String, Object>() {{
                    put("type", "stop");
                    put("stop_words", Arrays.asList("of", "for"));
                }}
        )
);

List<String> texts = new ArrayList<>();
texts.add("Milvus provides flexible, customizable analyzers for robust text processing.");

RunAnalyzerResp resp = client.runAnalyzer(RunAnalyzerReq.builder()
        .texts(texts)
        .analyzerParams(analyzerParams)
        .build());
List<RunAnalyzerResp.AnalyzerResult> results = resp.getResults();
// Configure a custom analyzer for VARCHAR field `title`
const analyzerParamsCustom = {
  tokenizer: "standard",
  filter: [
    "lowercase",
    {
      type: "length",
      max: 40,
    },
    {
      type: "stop",
      stop_words: ["of", "to"],
    },
  ],
};
const sample_text = "Milvus provides flexible, customizable analyzers for robust text processing.";
const result = await client.run_analyzer({
    text: sample_text, 
    analyzer_params: analyzer_params_built_in
});
analyzerParams = map[string]any{"tokenizer": "standard",
    "filter": []any{"lowercase", 
    map[string]any{
        "type": "length",
        "max":  40,
    map[string]any{
        "type": "stop",
        "stop_words": []string{"of", "to"},
    }}}
    
bs, _ := json.Marshal(analyzerParams)
texts := []string{"Milvus provides flexible, customizable analyzers for robust text processing."}
option := milvusclient.NewRunAnalyzerOption(texts).
    WithAnalyzerParams(string(bs))

result, err := client.RunAnalyzer(ctx, option)
if err != nil {
    fmt.Println(err.Error())
    // handle error
}

3단계: 스키마에 필드 추가 (Step 3: Add fields to the schema)

이제 analyzer 구성을 검증했으니 스키마 필드에 추가해요.

# Add VARCHAR field 'title_en' using the built-in analyzer configuration
schema.add_field(
    field_name='title_en',
    datatype=DataType.VARCHAR,
    max_length=1000,
    enable_analyzer=True,
    analyzer_params=analyzer_params_built_in,
    enable_match=True,
)

# Add VARCHAR field 'title' using the custom analyzer configuration
schema.add_field(
    field_name='title',
    datatype=DataType.VARCHAR,
    max_length=1000,
    enable_analyzer=True,
    analyzer_params=analyzer_params_custom,
    enable_match=True,
)

# Add a vector field for embeddings
schema.add_field(field_name="embedding", datatype=DataType.FLOAT_VECTOR, dim=3)

# Add a primary key field
schema.add_field(field_name="id", datatype=DataType.INT64, is_primary=True)
schema.addField(AddFieldReq.builder()
        .fieldName("title")
        .dataType(DataType.VarChar)
        .maxLength(1000)
        .enableAnalyzer(true)
        .analyzerParams(analyzerParams)
        .enableMatch(true) // must enable this if you use TextMatch
        .build());

// Add vector field
schema.addField(AddFieldReq.builder()
        .fieldName("embedding")
        .dataType(DataType.FloatVector)
        .dimension(3)
        .build());
// Add primary field
schema.addField(AddFieldReq.builder()
        .fieldName("id")
        .dataType(DataType.Int64)
        .isPrimaryKey(true)
        .autoID(true)
        .build());
// Create schema
const schema = {
  auto_id: true,
  fields: [
    {
      name: "id",
      type: DataType.INT64,
      is_primary: true,
    },
    {
      name: "title_en",
      data_type: DataType.VARCHAR,
      max_length: 1000,
      enable_analyzer: true,
      analyzer_params: analyzerParamsBuiltIn,
      enable_match: true,
    },
    {
      name: "title",
      data_type: DataType.VARCHAR,
      max_length: 1000,
      enable_analyzer: true,
      analyzer_params: analyzerParamsCustom,
      enable_match: true,
    },
    {
      name: "embedding",
      data_type: DataType.FLOAT_VECTOR,
      dim: 4,
    },
  ],
};
schema.WithField(entity.NewField().
    WithName("id").
    WithDataType(entity.FieldTypeInt64).
    WithIsPrimaryKey(true).
    WithIsAutoID(true),
).WithField(entity.NewField().
    WithName("embedding").
    WithDataType(entity.FieldTypeFloatVector).
    WithDim(3),
).WithField(entity.NewField().
    WithName("title").
    WithDataType(entity.FieldTypeVarChar).
    WithMaxLength(1000).
    WithEnableAnalyzer(true).
    WithAnalyzerParams(analyzerParams).
    WithEnableMatch(true),
)

4단계: 인덱스 파라미터 준비 및 컬렉션 생성 (Step 4: Prepare index parameters and create the collection)

# Set up index parameters for the vector field
index_params = client.prepare_index_params()
index_params.add_index(field_name="embedding", metric_type="COSINE", index_type="AUTOINDEX")

# Create the collection with the defined schema and index parameters
client.create_collection(
    collection_name="my_collection",
    schema=schema,
    index_params=index_params
)
// Set up index params for vector field
List<IndexParam> indexes = new ArrayList<>();
indexes.add(IndexParam.builder()
        .fieldName("embedding")
        .indexType(IndexParam.IndexType.AUTOINDEX)
        .metricType(IndexParam.MetricType.COSINE)
        .build());

// Create collection with defined schema
CreateCollectionReq requestCreate = CreateCollectionReq.builder()
        .collectionName("my_collection")
        .collectionSchema(schema)
        .indexParams(indexes)
        .build();
client.createCollection(requestCreate);
// Set up index params for vector field
const indexParams = [
  {
    name: "embedding",
    metric_type: "COSINE",
    index_type: "AUTOINDEX",
  },
];

// Create collection with defined schema
await client.createCollection({
  collection_name: "my_collection",
  schema: schema,
  index_params: indexParams,
});

console.log("Collection created successfully!");
idx := index.NewAutoIndex(index.MetricType(entity.COSINE))
indexOption := milvusclient.NewCreateIndexOption("my_collection", "embedding", idx)

err = client.CreateCollection(ctx,
    milvusclient.NewCreateCollectionOption("my_collection", schema).
        WithIndexOptions(indexOption))
if err != nil {
    fmt.Println(err.Error())
    // handle error
}

다음 단계 (What's next)

Analyzer를 구성한 뒤 Milvus가 제공하는 텍스트 검색 기능과 통합할 수 있어요. 자세한 내용은 다음을 참고하세요.

더 알아보기 (Learn more)