Analyzer 매핑 파라미터
Analyzer 매핑 파라미터
analyzer 매핑 파라미터는 text 필드를 인덱싱하거나 검색할 때 텍스트 분석에 사용할 분석기(analyzer)를 지정해요. search_analyzer 매핑 파라미터로 재정의하지 않는 한, 이 분석기가 인덱스 시점과 검색 시점 분석을 모두 처리해요.
출처: 문서
본문
analyzer 매핑 파라미터는 text 필드를 인덱싱하거나 검색할 때 텍스트 분석에 사용할 분석기를 지정합니다. search_analyzer 매핑 파라미터로 재정의하지 않는 한, 이 분석기가 인덱스 시점(index-time)과 검색 시점(search-time) 분석을 모두 처리합니다. 분석기에 대한 자세한 내용은 Text analysis를 참고하세요.
analyzer 매핑 파라미터는 text 필드만 지원합니다.
analyzer 파라미터는 Update Mapping API를 사용해 기존 필드에서 갱신할 수 없습니다. 기존 필드의 분석기를 변경하려면 데이터를 재인덱싱해야 합니다.
분석기를 프로덕션 환경에 배포하기 전에 테스트할 것을 권장합니다.
검색 인용 분석기(Search quote analyzer)
search_quote_analyzer 파라미터를 사용하면 구문 쿼리(phrase query)에 대해 다른 분석기를 지정할 수 있습니다. 이는 일반 용어 검색과 비교해 구문 검색에서 중단 단어(stop word)를 다르게 처리해야 할 때 특히 유용합니다.
중단 단어가 있는 구문 쿼리를 효과적으로 처리하려면 세 가지 분석기 설정을 구성하세요.
- 중단 단어를 포함한 모든 용어를 보존하는 인덱싱용 분석기.
- 일반 쿼리에서 중단 단어를 걸러내는
search_analyzer. - 구문 쿼리에서 중단 단어를 유지하는
search_quote_analyzer.
예제
다음 예제는 구문 쿼리에서 중단 단어를 용어 쿼리와 다르게 처리하기 위해 search_quote_analyzer를 사용하는 방법을 보여 줍니다.
먼저 세 가지 분석기 유형을 모두 사용해 인덱스를 만듭니다. index_analyzer는 "the"와 "a" 같은 중단 단어를 포함한 모든 용어를 인덱싱 중 보존합니다. search_analyzer는 일반 용어 쿼리에서 중단 단어를 제거합니다. search_quote_analyzer는 문서 인덱싱에 사용한 것과 같은 분석기를 사용해 정확한 구문 매칭이 올바르게 작동하도록 합니다.
PUT /product_catalog
{
"settings": {
"analysis": {
"analyzer": {
"index_analyzer": {
"type": "custom",
"tokenizer": "standard",
"filter": [
"lowercase"
]
},
"search_analyzer": {
"type": "custom",
"tokenizer": "standard",
"filter": [
"lowercase",
"english_stop"
]
}
},
"filter": {
"english_stop": {
"type": "stop",
"stopwords": "_english_"
}
}
}
},
"mappings": {
"properties": {
"product_name": {
"type": "text",
"analyzer": "index_analyzer",
"search_analyzer": "search_analyzer",
"search_quote_analyzer": "index_analyzer"
}
}
}
}
다음으로 인덱스에 샘플 문서를 추가합니다.
PUT /product_catalog/_doc/1
{
"product_name": "The Smart Watch Pro"
}
PUT /product_catalog/_doc/2
{
"product_name": "A Smart Watch Ultra"
}
인덱스에서 구문 "the smart watch"(따옴표로 묶음)를 검색합니다.
GET /product_catalog/_search
{
"query": {
"query_string": {
"query": "\"the smart watch\"",
"default_field": "product_name"
}
}
}
쿼리가 따옴표로 묶여 있으므로 구문 쿼리가 되며, 이는 다음 쿼리와 동일합니다.
GET /product_catalog/_search
{
"query": {
"match_phrase": {
"product_name": "the smart watch"
}
}
}
구문 쿼리는 중단 단어를 보존하는 search_quote_analyzer를 사용합니다. 그 결과 쿼리 "the smart watch"는 해당 정확한 구문을 포함하는 문서만 매칭하므로 응답에는 첫 번째 문서만 포함됩니다.
{
"took": 263,
"timed_out": false,
"_shards": {
"total": 1,
"successful": 1,
"skipped": 0,
"failed": 0
},
"hits": {
"total": {
"value": 1,
"relation": "eq"
},
"max_score": 0.48081374,
"hits": [
{
"_index": "product_catalog",
"_id": "1",
"_score": 0.48081374,
"_source": {
"product_name": "The Smart Watch Pro"
}
}
]
}
}
이제 "the smart watch" 텍스트(따옴표 없이)를 검색합니다.
GET /product_catalog/_search
{
"query": {
"query_string": {
"query": "the smart watch",
"default_field": "product_name"
}
}
}
쿼리가 따옴표로 묶여 있지 않으므로 용어 수준(term-level) 쿼리가 되며, 이는 다음 쿼리와 동일합니다.
GET /product_catalog/_search
{
"query": {
"match": {
"product_name": "the smart watch"
}
}
}
용어 수준 쿼리는 텍스트를 토큰화하고 중단 단어를 제거하는 search_analyzer를 사용합니다. 그 결과 쿼리는 [smart, watch] 토큰으로 분석되어 두 문서 모두와 매칭됩니다.
{
"took": 38,
"timed_out": false,
"_shards": {
"total": 1,
"successful": 1,
"skipped": 0,
"failed": 0
},
"hits": {
"total": {
"value": 2,
"relation": "eq"
},
"max_score": 0.16574687,
"hits": [
{
"_index": "product_catalog",
"_id": "1",
"_score": 0.16574687,
"_source": {
"product_name": "The Smart Watch Pro"
}
},
{
"_index": "product_catalog",
"_id": "2",
"_score": 0.16574687,
"_source": {
"product_name": "A Smart Watch Ultra"
}
}
]
}
}
관련 문서(Related documentation)
- Text analysis
- Search analyzers
- Query string queries