본문 바로가기
WIKI 기술 지식 베이스

Jieba 토크나이저 (Jieba)

원문 보기 위키 갱신

jieba 토크나이저는 중국어 텍스트를 구성 요소 단어로 분해해 처리해요.

jieba 토크나이저는 구두점을 출력에서 별도의 토큰으로 유지해요. 예를 들어 "你好!世界。"는 ["你好", "!", "世界", "。"]가 돼요. 이런 독립 구두점 토큰을 제거하려면 removepunct 필터를 사용하세요.

출처: Milvus 문서

본문

구성 (Configuration)

Milvus는 jieba 토크나이저에 대해 두 가지 구성 방식을 지원해요. 단순 구성과 사용자 지정 구성이에요.

단순 구성 (Simple configuration)

단순 구성에서는 토크나이저를 "jieba"로만 설정하면 돼요. 예를 들어:

Python
# Simple configuration: only specifying the tokenizer name
analyzer_params = {
    "tokenizer": "jieba",  # Use the default settings: dict=["_default_"], mode="search", hmm=True
}
Java
Map<String, Object> analyzerParams = new HashMap<>();
analyzerParams.put("tokenizer", "jieba");
NodeJS
const analyzer_params = {
    "tokenizer": "jieba",
};
Go
analyzerParams = map[string]any{"tokenizer": "jieba"}
cURL
# restful
analyzerParams='{
  "tokenizer": "jieba"
}'

이 단순 구성은 다음 사용자 지정 구성과 동일해요:

Python
# Custom configuration equivalent to the simple configuration above
analyzer_params = {
    "type": "jieba",          # Tokenizer type, fixed as "jieba"
    "dict": ["_default_"],     # Use the default dictionary
    "mode": "search",          # Use search mode for improved recall (see mode details below)
    "hmm": True                # Enable HMM for probabilistic segmentation
}
Java
Map<String, Object> analyzerParams = new HashMap<>();
analyzerParams.put("type", "jieba");
analyzerParams.put("dict", Collections.singletonList("_default_"));
analyzerParams.put("mode", "search");
analyzerParams.put("hmm", true);
NodeJS
// javascript
Go
analyzerParams = map[string]any{"type": "jieba", "dict": []any{"_default_"}, "mode": "search", "hmm": true}
cURL
# restful

파라미터에 대한 자세한 내용은 Custom configuration을 참고하세요.

사용자 지정 구성 (Custom configuration)

더 많은 제어가 필요하다면 사용자 지정 구성을 제공할 수 있어요. 사용자 지정 사전 지정, 분할 모드 선택, HMM(Hidden Markov Model) 켜고 끄기까지 할 수 있어요. 예를 들어:

Python
# Custom configuration with user-defined settings
analyzer_params = {
    "tokenizer": {
        "type": "jieba",           # Fixed tokenizer type
        "dict": ["customDictionary"],  # Custom dictionary list; replace with your own terms
        "mode": "exact",           # Use exact mode (non-overlapping tokens)
        "hmm": False               # Disable HMM; unmatched text will be split into individual characters
    }
}
Java
Map<String, Object> analyzerParams = new HashMap<>();
analyzerParams.put("tokenizer", new HashMap<String, Object>() {{
  put("type", "jieba"));
  put("dict", Arrays.asList("customDictionary"));
  put("mode", "exact");
  put("hmm", false);
}});
NodeJS
// javascript
Go
analyzerParams := map[string]interface{}{
  "tokenizer": map[string]interface{}{
      "type": "jieba",
      "dict": []string{"customDictionary"},
      "mode": "exact",
      "hmm":  false,
  },
}
cURL
# restful
파라미터 설명 기본값
type 토크나이저 타입이에요. "jieba"로 고정돼요. "jieba"
dict analyzer가 어휘 소스로 로드할 사전 목록이에요. 내장 옵션:
- "_default_": 엔진의 내장 간체 중국어 사전을 로드해요. 자세한 내용은 dict.txt를 참고하세요.
- "_extend_default_": "_default_"의 모든 것에 추가로 번체 중국어 보충 사전을 더해 로드해요. 자세한 내용은 dict.txt.big를 참고하세요.
내장 사전과 임의 개수의 사용자 지정 사전을 섞을 수도 있어요. 예: ["_default_", "结巴分词器"].
["_default_"]
mode 분할 모드예요. 가능한 값:
- "exact": 문장을 가장 정밀한 방식으로 분할하려고 해요. 텍스트 분석에 이상적이에요.
- "search": exact 모드를 기반으로 긴 단어를 더 잘게 나눠 recall을 높여요. 검색 엔진 토큰화에 적합해요.
자세한 내용은 Jieba GitHub Project를 참고하세요.
"search"
hmm 사전에 없는 단어를 확률적으로 분할하기 위해 HMM(Hidden Markov Model)을 켤지 여부를 나타내는 불리언 플래그예요. true

큰 사용자 지정 어휘를 dict로 인라인하지 않고 외부 파일에서 로드하려면 아래의 Custom configuration with a dictionary file을 참고하세요.

analyzer_params를 정의한 뒤에는 컬렉션 스키마를 정의할 때 VARCHAR 필드에 적용할 수 있어요. 그러면 Milvus가 지정한 analyzer로 해당 필드의 텍스트를 처리해 효율적인 토큰화와 필터링을 수행해요. 자세한 내용은 Example use를 참고하세요.

사전 파일로 사용자 지정 구성하기 (Custom configuration with a dictionary file) — Milvus 3.0.x 호환

도메인 용어집, 제품 용어, 고유명사 목록 같은 큰 사용자 지정 어휘는 단어를 파일에 저장하고 그 파일을 원격 파일 리소스로 등록한 뒤, extra_dict_file 파라미터로 토크나이저에서 참조해요. 그러면 analyzer가 내장 사전 위에 이 단어들을 어휘로 로드해요.

파일은 평범한 UTF-8 텍스트로, 한 줄에 용어 하나씩이에요. 예를 들어:

结巴分词器
向量数据库

파일을 Milvus 클러스터가 사용하도록 구성된 오브젝트 스토어에 업로드한 뒤 등록해요:

Python
from pymilvus import MilvusClient

client = MilvusClient(uri="http://localhost:19530")

# Register the uploaded file under a name you'll reference from analyzer configs.
client.add_file_resource(
    name="zh_terms",
    path="file/zh_terms.txt",    # full S3 object key, including the rootPath prefix
)
Java
// java
NodeJS
// nodejs
Go
// go
cURL
# restful

토크나이저에서 extra_dict_file로 등록된 리소스를 참조해요:

Python
analyzer_params = {
    "tokenizer": {
        "type": "jieba",
        "dict": ["_default_"],             # keep the built-in dictionary
        "mode": "exact",
        "hmm": False,
        "extra_dict_file": {
            "type": "remote",
            "resource_name": "zh_terms",
            "file_name": "zh_terms.txt",
        },
    },
}

client.run_analyzer(["milvus结巴分词器中文测试"], analyzer_params)
# → [['milvus', '结巴', '分词器', '中文', '测试']]
Java
// java
NodeJS
// nodejs
Go
// go
cURL
# restful

extra_dict_file 파라미터는 다음 필드를 가진 객체를 받아요:

필드 설명
type 리소스 타입이에요. add_file_resource로 등록한 파일에는 "remote"를 사용해요. self-hosted 배포에 쓰는 "local" 변형은 Manage File Resources를 참고하세요.
resource_name add_file_resource로 파일을 등록할 때 사용한 이름이에요.
file_name 등록된 리소스의 오브젝트 스토어 경로 중 파일 이름 부분이에요. (예: 리소스를 path="file/zh_terms.txt"로 등록했다면 "zh_terms.txt".)

extra_dict_file로 추가한 단어는 내장 사전과 병합되므로 jieba의 분할 알고리즘이 기존 항목과 함께 이 단어들을 보게 돼요. 특정 용어가 독립 토큰으로 나타날지 여부는 jieba의 확률 가중 DAG 선택에 달려 있어요. 向量数据库 같은 긴 사용자 지정 용어는 내장 사전에서 向量 + 数据库 같은 짧은 항목들의 빈도가 더 높다면 여전히 그렇게 나뉠 수 있어요.

예시 (Examples)

analyzer 구성을 컬렉션 스키마에 적용하기 전에 run_analyzer 메서드로 동작을 먼저 확인해 보세요.

Analyzer 구성

Python
analyzer_params = {
    "tokenizer": {
        "type": "jieba",
        "dict": ["结巴分词器"],
        "mode": "exact",
        "hmm": False
    }
}
Java
Map<String, Object> analyzerParams = new HashMap<>();
analyzerParams.put("tokenizer", new HashMap<String, Object>() {{
  put("type", "jieba"));
  put("dict", Arrays.asList("结巴分词器"));
  put("mode", "exact");
  put("hmm", false);
}});
NodeJS
// javascript
Go
analyzerParams := map[string]interface{}{
  "tokenizer": map[string]interface{}{
      "type": "jieba",
      "dict": []string{"结巴分词器"},
      "mode": "exact",
      "hmm":  false,
  },
}
cURL
# restful

run_analyzer로 검증하기

Python
from pymilvus import (
    MilvusClient,
)

client = MilvusClient(
    uri="http://localhost:19530",
    token="root:Milvus"
)

# Sample text to analyze
sample_text = "milvus结巴分词器中文测试"

# Run the standard analyzer with the defined configuration
result = client.run_analyzer(sample_text, analyzer_params)
print("Standard analyzer output:", result)
Java
import io.milvus.v2.client.ConnectConfig;
import io.milvus.v2.client.MilvusClientV2;
import io.milvus.v2.service.vector.request.RunAnalyzerReq;
import io.milvus.v2.service.vector.response.RunAnalyzerResp;

ConnectConfig config = ConnectConfig.builder()
        .uri("http://localhost:19530")
        .token("root:Milvus")
        .build();
MilvusClientV2 client = new MilvusClientV2(config);

List<String> texts = new ArrayList<>();
texts.add("milvus结巴分词器中文测试");

RunAnalyzerResp resp = client.runAnalyzer(RunAnalyzerReq.builder()
        .texts(texts)
        .analyzerParams(analyzerParams)
        .build());
List<RunAnalyzerResp.AnalyzerResult> results = resp.getResults();
NodeJS
// javascript
Go
import (
    "context"
    "encoding/json"
    "fmt"

    "github.com/milvus-io/milvus/client/v2/milvusclient"
)

client, err := milvusclient.New(ctx, &milvusclient.ClientConfig{
    Address: "localhost:19530",
    APIKey:  "root:Milvus",
})
if err != nil {
    fmt.Println(err.Error())
    // handle error
}

bs, _ := json.Marshal(analyzerParams)
texts := []string{"milvus结巴分词器中文测试"}
option := milvusclient.NewRunAnalyzerOption(texts).
    WithAnalyzerParams(string(bs))

result, err := client.RunAnalyzer(ctx, option)
if err != nil {
    fmt.Println(err.Error())
    // handle error
}
cURL
# restful

예상 출력

['milvus', '结巴分词器', '中', '文', '测', '试']

더 알아보기 (Learn more)