term_vector 매핑 파라미터
term_vector 매핑 파라미터
term_vector 매핑 파라미터는 인덱싱 중에 개별 text 필드에 대한 용어 수준(term-level) 정보를 저장할지 여부를 제어해요. 이 정보에는 용어 빈도, 위치(position), 문자 오프셋(offset) 같은 세부 내용이 포함되며, 커스텀 스코어링이나 하이라이팅 같은 고급 기능에 사용될 수 있어요.
기본적으로 term_vector는 비활성화되어 있어요. 활성화하면 용어 벡터가 저장되며 _termvectors API로 검색할 수 있어요. term_vector를 활성화하면 인덱스 크기가 커지므로, 자세한 용어 수준 데이터가 필요할 때만 사용해야 해요.
출처: 문서
본문
구성 옵션
term_vector 파라미터는 다음 유효한 값을 지원해요:
- no (기본값): 용어 벡터를 저장하지 않아요.
- yes: 용어 빈도(특정 문서에서 용어가 나타난 횟수)와 기본 위치를 저장해요.
- with_positions: 용어 위치를 저장해요. 필드에서 용어가 나타나는 순서를 의미해요.
- with_offsets: 문자 오프셋을 저장해요. 필드 텍스트 내에서 용어의 정확한 시작/끝 문자 위치를 의미해요.
- with_positions_offsets: 위치와 오프셋을 모두 저장해요.
- with_positions_payloads: 용어 위치와 함께 payload를 저장해요. payload는 인덱싱 중에 개별 용어에 첨부할 수 있는 선택적 커스텀 메타데이터(예: 태그나 숫자 값) 조각이에요. payload는 커스텀 스코어링이나 태깅 같은 고급 시나리오에서 사용되지만, 설정하려면 특별한 analyzer가 필요해요.
- with_positions_offsets_payloads: 모든 용어 벡터 데이터를 저장해요.
필드에서 term_vector 활성화하기
다음 요청은 위치와 오프셋을 포함해 용어 벡터를 저장하도록 구성된 content 필드가 있는 articles 인덱스를 생성해요:
PUT /articles
{
"mappings": {
"properties": {
"content": {
"type": "text",
"term_vector": "with_positions_offsets"
}
}
}
}
샘플 문서를 인덱스해요:
PUT /articles/_doc/1
{
"content": "OpenSearch is an open-source search and analytics suite."
}
_termvectors API를 사용해 용어 수준 통계를 검색해요:
POST /articles/_termvectors/1
{
"fields": ["content"],
"term_statistics": true,
"positions": true,
"offsets": true
}
다음 응답에는 문서 ID 1의 content 필드에 대한 상세한 용어 수준 통계(용어 빈도, 문서 빈도, 토큰 위치, 문자 오프셋 등)가 포함돼요:
{
"_index": "articles",
"_id": "1",
"_version": 1,
"found": true,
"took": 4,
"term_vectors": {
"content": {
"field_statistics": {
"sum_doc_freq": 9,
"doc_count": 1,
"sum_ttf": 9
},
"terms": {
"an": {
"doc_freq": 1,
"ttf": 1,
"term_freq": 1,
"tokens": [
{
"position": 2,
"start_offset": 14,
"end_offset": 16
}
]
},
"analytics": {
"doc_freq": 1,
"ttf": 1,
"term_freq": 1,
"tokens": [
{
"position": 7,
"start_offset": 40,
"end_offset": 49
}
]
},
"and": {
"doc_freq": 1,
"ttf": 1,
"term_freq": 1,
"tokens": [
{
"position": 6,
"start_offset": 36,
"end_offset": 39
}
]
},
"is": {
"doc_freq": 1,
"ttf": 1,
"term_freq": 1,
"tokens": [
{
"position": 1,
"start_offset": 11,
"end_offset": 13
}
]
},
"open": {
"doc_freq": 1,
"ttf": 1,
"term_freq": 1,
"tokens": [
{
"position": 3,
"start_offset": 17,
"end_offset": 21
}
]
},
"opensearch": {
"doc_freq": 1,
"ttf": 1,
"term_freq": 1,
"tokens": [
{
"position": 0,
"start_offset": 0,
"end_offset": 10
}
]
},
"search": {
"doc_freq": 1,
"ttf": 1,
"term_freq": 1,
"tokens": [
{
"position": 5,
"start_offset": 29,
"end_offset": 35
}
]
},
"source": {
"doc_freq": 1,
"ttf": 1,
"term_freq": 1,
"tokens": [
{
"position": 4,
"start_offset": 22,
"end_offset": 28
}
]
},
"suite": {
"doc_freq": 1,
"ttf": 1,
"term_freq": 1,
"tokens": [
{
"position": 8,
"start_offset": 50,
"end_offset": 55
}
]
}
}
}
}
}
term vector로 하이라이팅하기
필드에 저장된 용어 벡터를 사용해 "analytics"라는 용어를 검색하고 하이라이팅하려면 다음 명령을 사용해요:
POST /articles/_search
{
"query": {
"match": {
"content": "analytics"
}
},
"highlight": {
"fields": {
"content": {
"type": "fvh"
}
}
}
}
다음 응답은 content 필드에서 "analytics" 용어가 발견된 매칭 문서를 보여줘요. highlight 섹션에는 필드에 저장된 용어 벡터를 사용해 효율적이고 정확하게 하이라이팅된 매칭 용어가 <em> 태그로 감싸져 포함돼요:
{
...
"hits": {
"total": {
"value": 1,
"relation": "eq"
},
"max_score": 0.2876821,
"hits": [
{
"_index": "articles",
"_id": "1",
"_score": 0.2876821,
"_source": {
"content": "OpenSearch is an open-source search and analytics suite."
},
"highlight": {
"content": [
"OpenSearch is an open-source search and analytics suite."
]
}
}
]
}
}