Analyzer 매핑 파라미터

Analyzer 매핑 파라미터

analyzer 매핑 파라미터는 text 필드를 인덱싱하거나 검색할 때 텍스트 분석에 사용할 분석기(analyzer)를 지정해요. search_analyzer 매핑 파라미터로 재정의하지 않는 한, 이 분석기가 인덱스 시점과 검색 시점 분석을 모두 처리해요.

출처: 문서

본문

analyzer 매핑 파라미터는 text 필드를 인덱싱하거나 검색할 때 텍스트 분석에 사용할 분석기를 지정합니다. search_analyzer 매핑 파라미터로 재정의하지 않는 한, 이 분석기가 인덱스 시점(index-time)과 검색 시점(search-time) 분석을 모두 처리합니다. 분석기에 대한 자세한 내용은 Text analysis를 참고하세요.

analyzer 매핑 파라미터는 text 필드만 지원합니다.

analyzer 파라미터는 Update Mapping API를 사용해 기존 필드에서 갱신할 수 없습니다. 기존 필드의 분석기를 변경하려면 데이터를 재인덱싱해야 합니다.

분석기를 프로덕션 환경에 배포하기 전에 테스트할 것을 권장합니다.

검색 인용 분석기(Search quote analyzer)

search_quote_analyzer 파라미터를 사용하면 구문 쿼리(phrase query)에 대해 다른 분석기를 지정할 수 있습니다. 이는 일반 용어 검색과 비교해 구문 검색에서 중단 단어(stop word)를 다르게 처리해야 할 때 특히 유용합니다.

중단 단어가 있는 구문 쿼리를 효과적으로 처리하려면 세 가지 분석기 설정을 구성하세요.

  • 중단 단어를 포함한 모든 용어를 보존하는 인덱싱용 분석기.
  • 일반 쿼리에서 중단 단어를 걸러내는 search_analyzer.
  • 구문 쿼리에서 중단 단어를 유지하는 search_quote_analyzer.

예제

다음 예제는 구문 쿼리에서 중단 단어를 용어 쿼리와 다르게 처리하기 위해 search_quote_analyzer를 사용하는 방법을 보여 줍니다.

먼저 세 가지 분석기 유형을 모두 사용해 인덱스를 만듭니다. index_analyzer는 "the"와 "a" 같은 중단 단어를 포함한 모든 용어를 인덱싱 중 보존합니다. search_analyzer는 일반 용어 쿼리에서 중단 단어를 제거합니다. search_quote_analyzer는 문서 인덱싱에 사용한 것과 같은 분석기를 사용해 정확한 구문 매칭이 올바르게 작동하도록 합니다.

PUT /product_catalog
{
  "settings": {
    "analysis": {
      "analyzer": {
        "index_analyzer": {
          "type": "custom",
          "tokenizer": "standard",
          "filter": [
            "lowercase"
          ]
        },
        "search_analyzer": {
          "type": "custom",
          "tokenizer": "standard",
          "filter": [
            "lowercase",
            "english_stop"
          ]
        }
      },
      "filter": {
        "english_stop": {
          "type": "stop",
          "stopwords": "_english_"
        }
      }
    }
  },
  "mappings": {
    "properties": {
      "product_name": {
        "type": "text",
        "analyzer": "index_analyzer",
        "search_analyzer": "search_analyzer",
        "search_quote_analyzer": "index_analyzer"
      }
    }
  }
}

다음으로 인덱스에 샘플 문서를 추가합니다.

PUT /product_catalog/_doc/1
{
  "product_name": "The Smart Watch Pro"
}
PUT /product_catalog/_doc/2
{
  "product_name": "A Smart Watch Ultra"
}

인덱스에서 구문 "the smart watch"(따옴표로 묶음)를 검색합니다.

GET /product_catalog/_search
{
  "query": {
    "query_string": {
      "query": "\"the smart watch\"",
      "default_field": "product_name"
    }
  }
}

쿼리가 따옴표로 묶여 있으므로 구문 쿼리가 되며, 이는 다음 쿼리와 동일합니다.

GET /product_catalog/_search
{
  "query": {
    "match_phrase": {
      "product_name": "the smart watch"
    }
  }
}

구문 쿼리는 중단 단어를 보존하는 search_quote_analyzer를 사용합니다. 그 결과 쿼리 "the smart watch"는 해당 정확한 구문을 포함하는 문서만 매칭하므로 응답에는 첫 번째 문서만 포함됩니다.

{
  "took": 263,
  "timed_out": false,
  "_shards": {
    "total": 1,
    "successful": 1,
    "skipped": 0,
    "failed": 0
  },
  "hits": {
    "total": {
      "value": 1,
      "relation": "eq"
    },
    "max_score": 0.48081374,
    "hits": [
      {
        "_index": "product_catalog",
        "_id": "1",
        "_score": 0.48081374,
        "_source": {
          "product_name": "The Smart Watch Pro"
        }
      }
    ]
  }
}

이제 "the smart watch" 텍스트(따옴표 없이)를 검색합니다.

GET /product_catalog/_search
{
  "query": {
    "query_string": {
      "query": "the smart watch",
      "default_field": "product_name"
    }
  }
}

쿼리가 따옴표로 묶여 있지 않으므로 용어 수준(term-level) 쿼리가 되며, 이는 다음 쿼리와 동일합니다.

GET /product_catalog/_search
{
  "query": {
    "match": {
      "product_name": "the smart watch"
    }
  }
}

용어 수준 쿼리는 텍스트를 토큰화하고 중단 단어를 제거하는 search_analyzer를 사용합니다. 그 결과 쿼리는 [smart, watch] 토큰으로 분석되어 두 문서 모두와 매칭됩니다.

{
  "took": 38,
  "timed_out": false,
  "_shards": {
    "total": 1,
    "successful": 1,
    "skipped": 0,
    "failed": 0
  },
  "hits": {
    "total": {
      "value": 2,
      "relation": "eq"
    },
    "max_score": 0.16574687,
    "hits": [
      {
        "_index": "product_catalog",
        "_id": "1",
        "_score": 0.16574687,
        "_source": {
          "product_name": "The Smart Watch Pro"
        }
      },
      {
        "_index": "product_catalog",
        "_id": "2",
        "_score": 0.16574687,
        "_source": {
          "product_name": "A Smart Watch Ultra"
        }
      }
    ]
  }
}

관련 문서(Related documentation)

  • Text analysis
  • Search analyzers
  • Query string queries

더 알아보기 (Learn more)