Jieba 토크나이저 (Jieba)
jieba 토크나이저는 중국어 텍스트를 구성 요소 단어로 분해해 처리해요.
jieba 토크나이저는 구두점을 출력에서 별도의 토큰으로 유지해요. 예를 들어 "你好!世界。"는 ["你好", "!", "世界", "。"]가 돼요. 이런 독립 구두점 토큰을 제거하려면 removepunct 필터를 사용하세요.
출처: Milvus 문서
본문
구성 (Configuration)
Milvus는 jieba 토크나이저에 대해 두 가지 구성 방식을 지원해요. 단순 구성과 사용자 지정 구성이에요.
단순 구성 (Simple configuration)
단순 구성에서는 토크나이저를 "jieba"로만 설정하면 돼요. 예를 들어:
Python
# Simple configuration: only specifying the tokenizer name
analyzer_params = {
"tokenizer": "jieba", # Use the default settings: dict=["_default_"], mode="search", hmm=True
}
Java
Map<String, Object> analyzerParams = new HashMap<>();
analyzerParams.put("tokenizer", "jieba");
NodeJS
const analyzer_params = {
"tokenizer": "jieba",
};
Go
analyzerParams = map[string]any{"tokenizer": "jieba"}
cURL
# restful
analyzerParams='{
"tokenizer": "jieba"
}'
이 단순 구성은 다음 사용자 지정 구성과 동일해요:
Python
# Custom configuration equivalent to the simple configuration above
analyzer_params = {
"type": "jieba", # Tokenizer type, fixed as "jieba"
"dict": ["_default_"], # Use the default dictionary
"mode": "search", # Use search mode for improved recall (see mode details below)
"hmm": True # Enable HMM for probabilistic segmentation
}
Java
Map<String, Object> analyzerParams = new HashMap<>();
analyzerParams.put("type", "jieba");
analyzerParams.put("dict", Collections.singletonList("_default_"));
analyzerParams.put("mode", "search");
analyzerParams.put("hmm", true);
NodeJS
// javascript
Go
analyzerParams = map[string]any{"type": "jieba", "dict": []any{"_default_"}, "mode": "search", "hmm": true}
cURL
# restful
파라미터에 대한 자세한 내용은 Custom configuration을 참고하세요.
사용자 지정 구성 (Custom configuration)
더 많은 제어가 필요하다면 사용자 지정 구성을 제공할 수 있어요. 사용자 지정 사전 지정, 분할 모드 선택, HMM(Hidden Markov Model) 켜고 끄기까지 할 수 있어요. 예를 들어:
Python
# Custom configuration with user-defined settings
analyzer_params = {
"tokenizer": {
"type": "jieba", # Fixed tokenizer type
"dict": ["customDictionary"], # Custom dictionary list; replace with your own terms
"mode": "exact", # Use exact mode (non-overlapping tokens)
"hmm": False # Disable HMM; unmatched text will be split into individual characters
}
}
Java
Map<String, Object> analyzerParams = new HashMap<>();
analyzerParams.put("tokenizer", new HashMap<String, Object>() {{
put("type", "jieba"));
put("dict", Arrays.asList("customDictionary"));
put("mode", "exact");
put("hmm", false);
}});
NodeJS
// javascript
Go
analyzerParams := map[string]interface{}{
"tokenizer": map[string]interface{}{
"type": "jieba",
"dict": []string{"customDictionary"},
"mode": "exact",
"hmm": false,
},
}
cURL
# restful
| 파라미터 | 설명 | 기본값 |
|---|---|---|
type |
토크나이저 타입이에요. "jieba"로 고정돼요. |
"jieba" |
dict |
analyzer가 어휘 소스로 로드할 사전 목록이에요. 내장 옵션: - "_default_": 엔진의 내장 간체 중국어 사전을 로드해요. 자세한 내용은 dict.txt를 참고하세요.- "_extend_default_": "_default_"의 모든 것에 추가로 번체 중국어 보충 사전을 더해 로드해요. 자세한 내용은 dict.txt.big를 참고하세요.내장 사전과 임의 개수의 사용자 지정 사전을 섞을 수도 있어요. 예: ["_default_", "结巴分词器"]. |
["_default_"] |
mode |
분할 모드예요. 가능한 값: - "exact": 문장을 가장 정밀한 방식으로 분할하려고 해요. 텍스트 분석에 이상적이에요.- "search": exact 모드를 기반으로 긴 단어를 더 잘게 나눠 recall을 높여요. 검색 엔진 토큰화에 적합해요.자세한 내용은 Jieba GitHub Project를 참고하세요. |
"search" |
hmm |
사전에 없는 단어를 확률적으로 분할하기 위해 HMM(Hidden Markov Model)을 켤지 여부를 나타내는 불리언 플래그예요. | true |
큰 사용자 지정 어휘를 dict로 인라인하지 않고 외부 파일에서 로드하려면 아래의 Custom configuration with a dictionary file을 참고하세요.
analyzer_params를 정의한 뒤에는 컬렉션 스키마를 정의할 때 VARCHAR 필드에 적용할 수 있어요. 그러면 Milvus가 지정한 analyzer로 해당 필드의 텍스트를 처리해 효율적인 토큰화와 필터링을 수행해요. 자세한 내용은 Example use를 참고하세요.
사전 파일로 사용자 지정 구성하기 (Custom configuration with a dictionary file) — Milvus 3.0.x 호환
도메인 용어집, 제품 용어, 고유명사 목록 같은 큰 사용자 지정 어휘는 단어를 파일에 저장하고 그 파일을 원격 파일 리소스로 등록한 뒤, extra_dict_file 파라미터로 토크나이저에서 참조해요. 그러면 analyzer가 내장 사전 위에 이 단어들을 어휘로 로드해요.
파일은 평범한 UTF-8 텍스트로, 한 줄에 용어 하나씩이에요. 예를 들어:
结巴分词器
向量数据库
파일을 Milvus 클러스터가 사용하도록 구성된 오브젝트 스토어에 업로드한 뒤 등록해요:
Python
from pymilvus import MilvusClient
client = MilvusClient(uri="http://localhost:19530")
# Register the uploaded file under a name you'll reference from analyzer configs.
client.add_file_resource(
name="zh_terms",
path="file/zh_terms.txt", # full S3 object key, including the rootPath prefix
)
Java
// java
NodeJS
// nodejs
Go
// go
cURL
# restful
토크나이저에서 extra_dict_file로 등록된 리소스를 참조해요:
Python
analyzer_params = {
"tokenizer": {
"type": "jieba",
"dict": ["_default_"], # keep the built-in dictionary
"mode": "exact",
"hmm": False,
"extra_dict_file": {
"type": "remote",
"resource_name": "zh_terms",
"file_name": "zh_terms.txt",
},
},
}
client.run_analyzer(["milvus结巴分词器中文测试"], analyzer_params)
# → [['milvus', '结巴', '分词器', '中文', '测试']]
Java
// java
NodeJS
// nodejs
Go
// go
cURL
# restful
extra_dict_file 파라미터는 다음 필드를 가진 객체를 받아요:
| 필드 | 설명 |
|---|---|
type |
리소스 타입이에요. add_file_resource로 등록한 파일에는 "remote"를 사용해요. self-hosted 배포에 쓰는 "local" 변형은 Manage File Resources를 참고하세요. |
resource_name |
add_file_resource로 파일을 등록할 때 사용한 이름이에요. |
file_name |
등록된 리소스의 오브젝트 스토어 경로 중 파일 이름 부분이에요. (예: 리소스를 path="file/zh_terms.txt"로 등록했다면 "zh_terms.txt".) |
extra_dict_file로 추가한 단어는 내장 사전과 병합되므로 jieba의 분할 알고리즘이 기존 항목과 함께 이 단어들을 보게 돼요. 특정 용어가 독립 토큰으로 나타날지 여부는 jieba의 확률 가중 DAG 선택에 달려 있어요. 向量数据库 같은 긴 사용자 지정 용어는 내장 사전에서 向量 + 数据库 같은 짧은 항목들의 빈도가 더 높다면 여전히 그렇게 나뉠 수 있어요.
예시 (Examples)
analyzer 구성을 컬렉션 스키마에 적용하기 전에 run_analyzer 메서드로 동작을 먼저 확인해 보세요.
Analyzer 구성
Python
analyzer_params = {
"tokenizer": {
"type": "jieba",
"dict": ["结巴分词器"],
"mode": "exact",
"hmm": False
}
}
Java
Map<String, Object> analyzerParams = new HashMap<>();
analyzerParams.put("tokenizer", new HashMap<String, Object>() {{
put("type", "jieba"));
put("dict", Arrays.asList("结巴分词器"));
put("mode", "exact");
put("hmm", false);
}});
NodeJS
// javascript
Go
analyzerParams := map[string]interface{}{
"tokenizer": map[string]interface{}{
"type": "jieba",
"dict": []string{"结巴分词器"},
"mode": "exact",
"hmm": false,
},
}
cURL
# restful
run_analyzer로 검증하기
Python
from pymilvus import (
MilvusClient,
)
client = MilvusClient(
uri="http://localhost:19530",
token="root:Milvus"
)
# Sample text to analyze
sample_text = "milvus结巴分词器中文测试"
# Run the standard analyzer with the defined configuration
result = client.run_analyzer(sample_text, analyzer_params)
print("Standard analyzer output:", result)
Java
import io.milvus.v2.client.ConnectConfig;
import io.milvus.v2.client.MilvusClientV2;
import io.milvus.v2.service.vector.request.RunAnalyzerReq;
import io.milvus.v2.service.vector.response.RunAnalyzerResp;
ConnectConfig config = ConnectConfig.builder()
.uri("http://localhost:19530")
.token("root:Milvus")
.build();
MilvusClientV2 client = new MilvusClientV2(config);
List<String> texts = new ArrayList<>();
texts.add("milvus结巴分词器中文测试");
RunAnalyzerResp resp = client.runAnalyzer(RunAnalyzerReq.builder()
.texts(texts)
.analyzerParams(analyzerParams)
.build());
List<RunAnalyzerResp.AnalyzerResult> results = resp.getResults();
NodeJS
// javascript
Go
import (
"context"
"encoding/json"
"fmt"
"github.com/milvus-io/milvus/client/v2/milvusclient"
)
client, err := milvusclient.New(ctx, &milvusclient.ClientConfig{
Address: "localhost:19530",
APIKey: "root:Milvus",
})
if err != nil {
fmt.Println(err.Error())
// handle error
}
bs, _ := json.Marshal(analyzerParams)
texts := []string{"milvus结巴分词器中文测试"}
option := milvusclient.NewRunAnalyzerOption(texts).
WithAnalyzerParams(string(bs))
result, err := client.RunAnalyzer(ctx, option)
if err != nil {
fmt.Println(err.Error())
// handle error
}
cURL
# restful
예상 출력
['milvus', '结巴分词器', '中', '文', '测', '试']
더 알아보기 (Learn more)
- removepunct — 독립 구두점 토큰 제거
- Analyzer 개요 — analyzer 구성과 토크나이저 전반
- Manage File Resources — self-hosted 배포의 파일 리소스 관리