Combined fields 쿼리
Combined fields 쿼리
combined_fields 쿼리는 여러 text 필드를 하나의 통합 필드로 취급하며, BM25F 알고리즘을 사용해 일관된 관련성 점수를 제공해요. 필드별로 별도의 쿼리를 실행하는 multi_match의 cross_fields 타입과 달리, 모든 필드를 함께 처리하므로 성능과 점수 정확도가 더 좋습니다.
출처: 문서
본문
3.2 버전에서 도입되었어요.
combined_fields 쿼리는 여러 text 필드를 하나의 통합 필드로 취급하며, BM25F 알고리즘을 사용해 일관된 관련성 점수를 제공해요. 필드별로 별도의 쿼리를 실행하는 multi_match의 cross_fields 타입과 달리, combined_fields는 모든 필드를 함께 처리해 더 나은 성능과 더 정확한 점수를 얻어요. 모든 필드에 걸쳐 통합 점수를 유지하면서 서로 다른 필드 가중치를 적용할 수 있어요.
combined_fields 쿼리의 모든 필드는 text 필드여야 하며 동일한 텍스트 분석기를 사용해야 해요.
설정 (Setup)
예시를 따라 하려면 샘플 문서 몇 개를 인덱싱하세요.
POST /books/_bulk
{"index":{"_id":"1"}}
{"title":"Database Systems","description":"A comprehensive guide to database design and implementation"}
{"index":{"_id":"2"}}
{"title":"Introduction to Systems","description":"This book covers database architectures and distributed systems"}
예시 (Example)
다음 예시는 combined_fields 쿼리에서 필드 가중치를 보여줘요. 이 쿼리는 title과 description 필드에 걸쳐 "database systems"를 검색하며, title 필드에 description 필드보다 4배 많은 가중치를 줘요.
GET /books/_search
{
"query": {
"combined_fields": {
"query": "database systems",
"fields": ["title^4", "description"]
}
}
}
응답은 두 쿼리 용어가 모두 가중치가 높은 title 필드에 나타나므로 "Database Systems"가 "Introduction to Systems"보다 훨씬 높게 점수가 매겨지는 것을 보여줘요.
{
"took": 1,
"timed_out": false,
"_shards": {
"total": 1,
"successful": 1,
"skipped": 0,
"failed": 0
},
"hits": {
"total": {
"value": 2,
"relation": "eq"
},
"max_score": 0.2924412,
"hits": [
{
"_index": "books",
"_id": "1",
"_score": 0.2924412,
"_source": {
"title": "Database Systems",
"description": "A comprehensive guide to database design and implementation"
}
},
{
"_index": "books",
"_id": "2",
"_score": 0.2239699,
"_source": {
"title": "Introduction to Systems",
"description": "This book covers database architectures and distributed systems"
}
}
]
}
}
파라미터 (Parameters)
다음 표는 combined_fields 쿼리 파라미터를 나열해요.
| 파라미터 | 데이터 타입 | 설명 |
|---|---|---|
query |
String | 검색할 쿼리 문자열. 필수. |
fields |
문자열 배열 | 검색할 필드. ^ 구문을 사용한 필드 이름 패턴과 필드 가중치를 지원함(예: title^2). 필수. |
operator |
String | 쿼리 문자열을 해석하는 데 사용하는 Boolean 논리. 유효한 값은 OR(기본값)와 AND. 선택 사항. |
minimum_should_match |
String | 문서가 반환되기 위해 일치해야 하는 최소 용어 수. 절대값, 백분율, 또는 조합일 수 있음. 자세한 내용은 Minimum should match를 참조하세요. 선택 사항. |
boost |
Floating-point | 관련성 점수에 대한 이 필드의 가중치를 지정하는 부동소수점 값. 1.0보다 큰 값은 필드의 관련성을 높이고, 0.0과 1.0 사이의 값은 필드의 관련성을 낮춤. 기본값은 1.0. 선택 사항. |
_name |
String | 응답에서 쿼리를 식별하는 데 사용할 수 있는 쿼리 이름. 선택 사항. |
AND 연산자 사용하기 (Using the AND operator)
기본적으로 쿼리는 OR 연산자를 사용하므로 쿼리의 어떤 용어와라도 일치하는 문서가 반환돼요. AND로 바꾸면 모든 용어가 일치하도록 요구할 수 있어요.
GET /books/_search
{
"query": {
"combined_fields": {
"query": "introduction systems",
"fields": ["title", "description"],
"operator": "AND"
}
}
}
응답에는 두 용어("introduction"과 "systems")가 모두 결합된 필드에 나타나는 유일한 문서가 포함돼요.
{
"took": 1,
"timed_out": false,
"_shards": {
"total": 1,
"successful": 1,
"skipped": 0,
"failed": 0
},
"hits": {
"total": {
"value": 1,
"relation": "eq"
},
"max_score": 0.4214915,
"hits": [
{
"_index": "books",
"_id": "2",
"_score": 0.4214915,
"_source": {
"title": "Introduction to Systems",
"description": "This book covers database architectures and distributed systems"
}
}
]
}
}
minimum_should_match 사용하기 (Using minimum_should_match)
minimum_should_match 파라미터를 사용하면 문서가 반환되기 위해 일치해야 하는 최소 용어 수를 지정할 수 있어요. 예를 들어 다음 쿼리는 용어의 최소 75%가 일치하도록 요구해요.
GET /books/_search
{
"query": {
"combined_fields": {
"query": "comprehensive database architectures book",
"fields": ["title", "description"],
"minimum_should_match": "75%"
}
}
}
응답에는 4개의 쿼리 용어 중 최소 3개(75%)와 일치하는 유일한 문서가 포함돼요.
{
"took": 2,
"timed_out": false,
"_shards": {
"total": 1,
"successful": 1,
"skipped": 0,
"failed": 0
},
"hits": {
"total": {
"value": 1,
"relation": "eq"
},
"max_score": 0.6993829,
"hits": [
{
"_index": "books",
"_id": "2",
"_score": 0.6993829,
"_source": {
"title": "Introduction to Systems",
"description": "This book covers database architectures and distributed systems"
}
}
]
}
}
multi_match cross_fields와의 비교 (Comparison with multi_match cross_fields)
combined_fields 쿼리는 관련성 점수를 계산하는 방식에서 type: cross_fields가 있는 multi_match와 달라요. 두 쿼리 타입 모두 여러 필드에서 검색하지만, combined_fields는 BM25F 알고리즘을 사용해 점수 산정 목적으로 모든 필드를 단일 통합 필드로 취급해요. 즉 역문서 빈도(IDF)가 필드별이 아니라 모든 필드에 걸쳐 전역적으로 계산되고, 용어 빈도 정규화가 모든 필드의 결합 길이를 고려해요. 이 방식은 짧은 title 필드와 긴 body 필드처럼 길이가 매우 다른 필드가 있을 때 특히 유용해요. 관련성 계산에서 더 짧은 필드가 과도하게 가중되는 것을 방지하기 때문이에요. 또한 combined_fields는 용어 중심 매칭을 사용하는데, 쿼리 용어가 개별 필드 안에서 모두 일치할 필요 없이 필드의 어떤 조합으로든 충족될 수 있어요.
다음 예시는 같은 쿼리로 두 방식을 비교해요.
Combined fields 쿼리:
GET /books/_search
{
"query": {
"combined_fields": {
"query": "database systems",
"fields": ["title", "description"]
}
}
}
combined_fields 쿼리는 모든 필드에서 함께 용어 빈도를 계산해 더 정확한 BM25F 점수를 제공해요.
{
"took": 6,
"timed_out": false,
"_shards": {
"total": 1,
"successful": 1,
"skipped": 0,
"failed": 0
},
"hits": {
"total": {
"value": 2,
"relation": "eq"
},
"max_score": 0.20001775,
"hits": [
{
"_index": "books",
"_id": "1",
"_score": 0.20001775,
"_source": {
"title": "Database Systems",
"description": "A comprehensive guide to database design and implementation"
}
},
{
"_index": "books",
"_id": "2",
"_score": 0.19373488,
"_source": {
"title": "Introduction to Systems",
"description": "This book covers database architectures and distributed systems"
}
}
]
}
}
Multi_match cross_fields 쿼리:
GET /books/_search
{
"query": {
"multi_match": {
"query": "database systems",
"fields": ["title", "description"],
"type": "cross_fields"
}
}
}
cross_fields 방식은 필드별로 별도의 쿼리를 실행한 뒤 결과를 결합하므로 점수가 덜 정확해질 수 있어요.
{
"took": 3,
"timed_out": false,
"_shards": {
"total": 1,
"successful": 1,
"skipped": 0,
"failed": 0
},
"hits": {
"total": {
"value": 2,
"relation": "eq"
},
"max_score": 0.18051638,
"hits": [
{
"_index": "books",
"_id": "1",
"_score": 0.18051638,
"_source": {
"title": "Database Systems",
"description": "A comprehensive guide to database design and implementation"
}
},
{
"_index": "books",
"_id": "2",
"_score": 0.16574687,
"_source": {
"title": "Introduction to Systems",
"description": "This book covers database architectures and distributed systems"
}
}
]
}
}
combined_fields 쿼리가 더 높은 관련성 점수를 산출함을 주목하세요(최상위 결과가 0.18051638 대비 0.20001775).
제한 사항 (Limitations)
combined_fields 쿼리의 다음 제한 사항을 유의하세요.
- 모든 필드는 동일한 텍스트 분석기를 가져야 해요.
- 이 쿼리는 text 필드에서만 작동해요.
- 이 쿼리는
multi_match처럼 필드별 부스트를 지원하지 않아요. 대신 필드 가중치를 사용하세요.