term_vector 매핑 파라미터

term_vector 매핑 파라미터

term_vector 매핑 파라미터는 인덱싱 중에 개별 text 필드에 대한 용어 수준(term-level) 정보를 저장할지 여부를 제어해요. 이 정보에는 용어 빈도, 위치(position), 문자 오프셋(offset) 같은 세부 내용이 포함되며, 커스텀 스코어링이나 하이라이팅 같은 고급 기능에 사용될 수 있어요.

기본적으로 term_vector는 비활성화되어 있어요. 활성화하면 용어 벡터가 저장되며 _termvectors API로 검색할 수 있어요. term_vector를 활성화하면 인덱스 크기가 커지므로, 자세한 용어 수준 데이터가 필요할 때만 사용해야 해요.

출처: 문서

본문

구성 옵션

term_vector 파라미터는 다음 유효한 값을 지원해요:

  • no (기본값): 용어 벡터를 저장하지 않아요.
  • yes: 용어 빈도(특정 문서에서 용어가 나타난 횟수)와 기본 위치를 저장해요.
  • with_positions: 용어 위치를 저장해요. 필드에서 용어가 나타나는 순서를 의미해요.
  • with_offsets: 문자 오프셋을 저장해요. 필드 텍스트 내에서 용어의 정확한 시작/끝 문자 위치를 의미해요.
  • with_positions_offsets: 위치와 오프셋을 모두 저장해요.
  • with_positions_payloads: 용어 위치와 함께 payload를 저장해요. payload는 인덱싱 중에 개별 용어에 첨부할 수 있는 선택적 커스텀 메타데이터(예: 태그나 숫자 값) 조각이에요. payload는 커스텀 스코어링이나 태깅 같은 고급 시나리오에서 사용되지만, 설정하려면 특별한 analyzer가 필요해요.
  • with_positions_offsets_payloads: 모든 용어 벡터 데이터를 저장해요.

필드에서 term_vector 활성화하기

다음 요청은 위치와 오프셋을 포함해 용어 벡터를 저장하도록 구성된 content 필드가 있는 articles 인덱스를 생성해요:

PUT /articles
{
  "mappings": {
    "properties": {
      "content": {
        "type": "text",
        "term_vector": "with_positions_offsets"
      }
    }
  }
}

샘플 문서를 인덱스해요:

PUT /articles/_doc/1
{
  "content": "OpenSearch is an open-source search and analytics suite."
}

_termvectors API를 사용해 용어 수준 통계를 검색해요:

POST /articles/_termvectors/1
{
  "fields": ["content"],
  "term_statistics": true,
  "positions": true,
  "offsets": true
}

다음 응답에는 문서 ID 1의 content 필드에 대한 상세한 용어 수준 통계(용어 빈도, 문서 빈도, 토큰 위치, 문자 오프셋 등)가 포함돼요:

{
  "_index": "articles",
  "_id": "1",
  "_version": 1,
  "found": true,
  "took": 4,
  "term_vectors": {
    "content": {
      "field_statistics": {
        "sum_doc_freq": 9,
        "doc_count": 1,
        "sum_ttf": 9
      },
      "terms": {
        "an": {
          "doc_freq": 1,
          "ttf": 1,
          "term_freq": 1,
          "tokens": [
            {
              "position": 2,
              "start_offset": 14,
              "end_offset": 16
            }
          ]
        },
        "analytics": {
          "doc_freq": 1,
          "ttf": 1,
          "term_freq": 1,
          "tokens": [
            {
              "position": 7,
              "start_offset": 40,
              "end_offset": 49
            }
          ]
        },
        "and": {
          "doc_freq": 1,
          "ttf": 1,
          "term_freq": 1,
          "tokens": [
            {
              "position": 6,
              "start_offset": 36,
              "end_offset": 39
            }
          ]
        },
        "is": {
          "doc_freq": 1,
          "ttf": 1,
          "term_freq": 1,
          "tokens": [
            {
              "position": 1,
              "start_offset": 11,
              "end_offset": 13
            }
          ]
        },
        "open": {
          "doc_freq": 1,
          "ttf": 1,
          "term_freq": 1,
          "tokens": [
            {
              "position": 3,
              "start_offset": 17,
              "end_offset": 21
            }
          ]
        },
        "opensearch": {
          "doc_freq": 1,
          "ttf": 1,
          "term_freq": 1,
          "tokens": [
            {
              "position": 0,
              "start_offset": 0,
              "end_offset": 10
            }
          ]
        },
        "search": {
          "doc_freq": 1,
          "ttf": 1,
          "term_freq": 1,
          "tokens": [
            {
              "position": 5,
              "start_offset": 29,
              "end_offset": 35
            }
          ]
        },
        "source": {
          "doc_freq": 1,
          "ttf": 1,
          "term_freq": 1,
          "tokens": [
            {
              "position": 4,
              "start_offset": 22,
              "end_offset": 28
            }
          ]
        },
        "suite": {
          "doc_freq": 1,
          "ttf": 1,
          "term_freq": 1,
          "tokens": [
            {
              "position": 8,
              "start_offset": 50,
              "end_offset": 55
            }
          ]
        }
      }
    }
  }
}

term vector로 하이라이팅하기

필드에 저장된 용어 벡터를 사용해 "analytics"라는 용어를 검색하고 하이라이팅하려면 다음 명령을 사용해요:

POST /articles/_search
{
  "query": {
    "match": {
      "content": "analytics"
    }
  },
  "highlight": {
    "fields": {
      "content": {
        "type": "fvh"
      }
    }
  }
}

다음 응답은 content 필드에서 "analytics" 용어가 발견된 매칭 문서를 보여줘요. highlight 섹션에는 필드에 저장된 용어 벡터를 사용해 효율적이고 정확하게 하이라이팅된 매칭 용어가 <em> 태그로 감싸져 포함돼요:

{
  ...
  "hits": {
    "total": {
      "value": 1,
      "relation": "eq"
    },
    "max_score": 0.2876821,
    "hits": [
      {
        "_index": "articles",
        "_id": "1",
        "_score": 0.2876821,
        "_source": {
          "content": "OpenSearch is an open-source search and analytics suite."
        },
        "highlight": {
          "content": [
            "OpenSearch is an open-source search and analytics suite."
          ]
        }
      }
    ]
  }
}

더 알아보기 (Learn more)