상위 적중 집계
상위 적중 집계 (Top hits aggregation)
top_hits 집계는 각 집계 버킷 내에서 점수가 가장 높은 문서를 검색하는 다중 값(multi-value) 메트릭 집계예요. 버킷 집계 내부에서 사용해 그룹별 대표 문서 또는 상위 순위 문서를 반환해요.
terms 같은 버킷 집계와 결합하면 top_hits 집계는 지정된 속성으로 결과 집합을 그룹화하고 각 그룹에서 점수가 가장 높거나 가장 최근에 업데이트된 문서를 검색해요. 다음 시나리오에서 유용해요:
- 각 상품 카테고리의 최신 거래를 표시하기
- 각 제조업체에서 점수가 가장 높은 검색 결과 보여주기
- 모든 일치 항목을 반환하지 않고 그룹화된 결과에서 대표 문서 검색하기
출처: 문서
본문
파라미터
다음 표는 top_hits 집계가 받는 파라미터를 나열해요.
| 파라미터 | 데이터 타입 | 설명 |
|---|---|---|
from |
Integer | 가져올 첫 번째 결과로부터의 오프셋이에요. 기본값은 0이에요. |
size |
Integer | 버킷당 반환할 최대 상위 적중 수예요. 기본값은 3이에요. |
sort |
Object 또는 Array | 상위 적중 문서가 정렬되는 방식을 정의해요. 기본적으로 적중 문서는 메인 쿼리의 점수로 정렬돼요. |
지원되는 적중별 기능 (Supported per-hit features)
top_hits 집계는 표준 검색 적중(hits)을 반환하므로 다음 적중별 기능이 지원돼요:
- 하이라이팅 (Highlighting)
- explain
- 명명된 쿼리 (Named queries)
- 소스 필터링 (Source filtering)
- 저장된 필드 (Stored fields)
- 스크립트 필드 (Script fields)
- doc value 필드
- 버전 포함 (Include versions)
- 시퀀스 번호와 기본 용어 포함 (Include sequence numbers and primary terms)
예제: 결과를 카테고리별로 그룹화하기 (Grouping results by category)
다음 예제에서는 e-커머스 데이터셋의 주문을 terms 집계로 상품 카테고리별로 그룹화하고, top_hits 하위 집계가 각 카테고리에서 가장 최근 주문을 검색해요. 소스에는 order_date, taxful_total_price, customer_full_name 필드만 포함돼요:
GET /opensearch_dashboards_sample_data_ecommerce/_search
{
"size": 0,
"aggs": {
"top_categories": {
"terms": {
"field": "category.keyword",
"size": 3
},
"aggs": {
"most_recent_sales": {
"top_hits": {
"sort": [
{
"order_date": {
"order": "desc"
}
}
],
"_source": {
"includes": ["order_date", "taxful_total_price", "customer_full_name"]
},
"size": 1
}
}
}
}
}
}
예제 응답:
{
"took": 25,
"timed_out": false,
"_shards": {
"total": 1,
"successful": 1,
"skipped": 0,
"failed": 0
},
"hits": {
"total": {
"value": 4675,
"relation": "eq"
},
"max_score": null,
"hits": []
},
"aggregations": {
"top_categories": {
"doc_count_error_upper_bound": 0,
"sum_other_doc_count": 2346,
"buckets": [
{
"key": "Men's Clothing",
"doc_count": 2024,
"most_recent_sales": {
"hits": {
"total": {
"value": 2024,
"relation": "eq"
},
"max_score": null,
"hits": [
{
"_index": "opensearch_dashboards_sample_data_ecommerce",
"_id": "poN5u50BpPQaFxReh8Tz",
"_score": null,
"_source": {
"customer_full_name": "Youssef Jensen",
"order_date": "2026-05-09T23:45:36+00:00",
"taxful_total_price": 78.98
},
"sort": [
1778370336000
]
}
]
}
}
},
{
"key": "Women's Clothing",
"doc_count": 1903,
"most_recent_sales": {
"hits": {
"total": {
"value": 1903,
"relation": "eq"
},
"max_score": null,
"hits": [
{
"_index": "opensearch_dashboards_sample_data_ecommerce",
"_id": "6IN5u50BpPQaFxRehr5T",
"_score": null,
"_source": {
"customer_full_name": "Sonya Smith",
"order_date": "2026-05-09T23:31:12+00:00",
"taxful_total_price": 42.98
},
"sort": [
1778369472000
]
}
]
}
}
},
{
"key": "Women's Shoes",
"doc_count": 1136,
"most_recent_sales": {
"hits": {
"total": {
"value": 1136,
"relation": "eq"
},
"max_score": null,
"hits": [
{
"_index": "opensearch_dashboards_sample_data_ecommerce",
"_id": "3IN5u50BpPQaFxReh78O",
"_score": null,
"_source": {
"customer_full_name": "Brigitte Cross",
"order_date": "2026-05-09T23:22:34+00:00",
"taxful_total_price": 91.98
},
"sort": [
1778368954000
]
}
]
}
}
}
]
}
}
}
예제: 필드 접기 (Field collapsing)
필드 접기(Field collapsing) 또는 결과 그룹화(result grouping)는 결과 집합을 논리적 그룹으로 구성하고 각 그룹에서 최상위 문서를 반환해요. 그룹은 그룹 내 가장 높은 점수를 가진 문서의 관련성 순서로 정렬돼요.
버킷 집계 안에 top_hits 집계를 감싸서 필드 접기를 구현할 수 있어요. 다음 예제는 e-커머스 데이터셋에서 shirt와 일치하는 상품을 검색하고 결과를 제조업체별로 그룹화해요. max 집계가 제조업체별 최고 점수를 포착하고, terms 집계가 그 점수를 사용해 버킷을 관련성 순서로 정렬해요:
GET /opensearch_dashboards_sample_data_ecommerce/_search
{
"size": 0,
"query": {
"match": {
"products.product_name": "shirt"
}
},
"aggs": {
"top_manufacturers": {
"terms": {
"field": "manufacturer.keyword",
"size": 3,
"order": {
"top_score": "desc"
}
},
"aggs": {
"top_hits_per_manufacturer": {
"top_hits": {
"_source": {
"includes": ["products.product_name", "manufacturer"]
},
"size": 1
}
},
"top_score": {
"max": {
"script": {
"source": "_score"
}
}
}
}
}
}
}
max(또는 min) 집계가 필요해요. top_hits 집계는 terms 집계의 order 옵션에서 직접 사용할 수 없기 때문이에요.
예제 응답:
{
"took": 38,
"timed_out": false,
"_shards": {
"total": 1,
"successful": 1,
"skipped": 0,
"failed": 0
},
"hits": {
"total": {
"value": 1160,
"relation": "eq"
},
"max_score": null,
"hits": []
},
"aggregations": {
"top_manufacturers": {
"doc_count_error_upper_bound": -1,
"sum_other_doc_count": 953,
"buckets": [
{
"key": "Elitelligence",
"doc_count": 503,
"top_score": {
"value": 0.9982529878616333
},
"top_hits_per_manufacturer": {
"hits": {
"total": {
"value": 503,
"relation": "eq"
},
"max_score": 0.998253,
"hits": [
{
"_index": "opensearch_dashboards_sample_data_ecommerce",
"_id": "34N5u50BpPQaFxReicgP",
"_score": 0.998253,
"_source": {
"manufacturer": [
"Elitelligence",
"Low Tide Media"
],
"products": [
{
"product_name": "Shirt - white"
},
{
"product_name": "Shirt - white"
}
]
}
}
]
}
}
},
{
"key": "Low Tide Media",
"doc_count": 500,
"top_score": {
"value": 0.9982529878616333
},
"top_hits_per_manufacturer": {
"hits": {
"total": {
"value": 500,
"relation": "eq"
},
"max_score": 0.998253,
"hits": [
{
"_index": "opensearch_dashboards_sample_data_ecommerce",
"_id": "34N5u50BpPQaFxReicgP",
"_score": 0.998253,
"_source": {
"manufacturer": [
"Elitelligence",
"Low Tide Media"
],
"products": [
{
"product_name": "Shirt - white"
},
{
"product_name": "Shirt - white"
}
]
}
}
]
}
}
},
{
"key": "Oceanavigations",
"doc_count": 330,
"top_score": {
"value": 0.9561269283294678
},
"top_hits_per_manufacturer": {
"hits": {
"total": {
"value": 330,
"relation": "eq"
},
"max_score": 0.9561269,
"hits": [
{
"_index": "opensearch_dashboards_sample_data_ecommerce",
"_id": "VYN5u50BpPQaFxRehbyq",
"_score": 0.9561269,
"_source": {
"manufacturer": [
"Oceanavigations",
"Low Tide Media"
],
"products": [
{
"product_name": "Shirt - grey"
},
{
"product_name": "Vibrant Patterned Shirt"
}
]
}
}
]
}
}
}
]
}
}
}
예제: 중첩 객체와 top hits 집계 사용하기
top_hits 집계가 nested 또는 reverse_nested 집계에 감싸여 있으면 중첩 적중(nested hits)을 반환해요. 중첩 적중은 내부적으로 부모 문서와 동일한 문서 ID를 공유하는 별도의 Lucene 문서로 저장돼요. top_hits 집계는 nested 또는 reverse_nested 집계 컨텍스트에서 사용될 때 이러한 내부 문서를 표면화할 수 있어요.
각 중첩 적중에는 배열 필드를 식별하는 _nested 필드와 해당 배열 내 중첩 객체의 0부터 시작하는 오프셋이 응답에 포함돼요. 이 정보는 부모 문서 소스 내에서 원래 중첩 객체를 찾는 데 유용해요.
먼저 중첩 필드 타입을 가진 인덱스를 생성해요:
PUT /top-hits-products
{
"mappings": {
"properties": {
"tags": { "type": "keyword" },
"reviews": {
"type": "nested",
"properties": {
"reviewer": { "type": "keyword" },
"comment": { "type": "text" }
}
}
}
}
}
중첩된 reviews 필드를 포함한 문서를 추가해요:
PUT /top-hits-products/_doc/1?refresh=true
{
"tags": ["laptop", "electronics"],
"reviews": [
{"reviewer": "tech_guru", "comment": "This laptop has outstanding battery life"},
{"reviewer": "casual_user", "comment": "Great laptop for everyday tasks"},
{"reviewer": "power_user", "comment": "This laptop handles heavy workloads easily"}
]
}
다음 요청은 laptop 태그가 있는 상품을 검색하고, 중첩된 리뷰를 리뷰어별로 그룹화하며, 각 리뷰어의 최상위 리뷰를 검색해요:
GET /top-hits-products/_search
{
"query": {
"term": { "tags": "laptop" }
},
"aggs": {
"by_product": {
"nested": {
"path": "reviews"
},
"aggs": {
"by_reviewer": {
"terms": {
"field": "reviews.reviewer",
"size": 1
},
"aggs": {
"by_nested": {
"top_hits": {}
}
}
}
}
}
}
}
_nested 필드는 배열 필드(reviews)와 해당 배열 내 중첩 객체의 0부터 시작하는 위치(offset)를 식별해요:
예제 응답:
{
"took": 16,
"timed_out": false,
"_shards": {
"total": 1,
"successful": 1,
"skipped": 0,
"failed": 0
},
"hits": {
"total": {
"value": 1,
"relation": "eq"
},
"max_score": 1.0,
"hits": [
{
"_index": "top-hits-products",
"_id": "1",
"_score": 1.0,
"_source": {
"tags": [
"laptop",
"electronics"
],
"reviews": [
{
"reviewer": "tech_guru",
"comment": "This laptop has outstanding battery life"
},
{
"reviewer": "casual_user",
"comment": "Great laptop for everyday tasks"
},
{
"reviewer": "power_user",
"comment": "This laptop handles heavy workloads easily"
}
]
}
}
]
},
"aggregations": {
"by_product": {
"doc_count": 3,
"by_reviewer": {
"doc_count_error_upper_bound": 0,
"sum_other_doc_count": 2,
"buckets": [
{
"key": "casual_user",
"doc_count": 1,
"by_nested": {
"hits": {
"total": {
"value": 1,
"relation": "eq"
},
"max_score": 1.0,
"hits": [
{
"_index": "top-hits-products",
"_id": "1",
"_nested": {
"field": "reviews",
"offset": 1
},
"_score": 1.0,
"_source": {
"comment": "Great laptop for everyday tasks",
"reviewer": "casual_user"
}
}
]
}
}
}
]
}
}
}
}
중첩 적중에 대해 _source를 요청하면 부모 문서의 전체 소스 대신 중첩 객체의 소스만 반환돼요. 중첩 객체 수준에 정의된 저장된 필드는 top_hits가 nested 또는 reverse_nested 집계 안에 있을 때에도 접근할 수 있어요.
_nested 필드는 중첩 적중에만 포함돼요. 일반(비중첩) 적중에는 이 필드가 포함되지 않아요.
_nested 필드는 인덱스에서 _source가 비활성화되어 있을 때 원본 소스 내 중첩 객체를 찾는 참조로도 사용할 수 있어요.
여러 수준의 중첩 객체 타입을 포함하는 매핑의 경우 _nested 정보는 계층적일 수 있어요. 다음 스니펫은 그 자체로 nested_child_field의 두 번째 위치에 있는 nested_grand_child_field의 첫 번째 위치에 있는 중첩 적중을 보여줘요:
"hits": [
{
"_index": "my-index",
"_id": "1",
"_score": 1,
"_nested": {
"field": "nested_child_field",
"offset": 1,
"_nested": {
"field": "nested_grand_child_field",
"offset": 0
}
},
"_source": ...
}
]
예제: 일치하는 용어 하이라이팅 (Highlighting matched terms)
다음 예제는 상품 이름에서 shirt를 검색하고 highlight를 사용해 각 상위 적중 내 일치하는 용어를 <em> 태그로 감싸요:
GET /opensearch_dashboards_sample_data_ecommerce/_search
{
"size": 0,
"query": {
"match": {
"products.product_name": "shirt"
}
},
"aggs": {
"top_categories": {
"terms": {
"field": "category.keyword",
"size": 2
},
"aggs": {
"top_doc": {
"top_hits": {
"size": 1,
"_source": {
"includes": ["products.product_name"]
},
"highlight": {
"fields": {
"products.product_name": {}
}
}
}
}
}
}
}
}
예제 응답:
{
"took": 17,
"timed_out": false,
"_shards": {
"total": 1,
"successful": 1,
"skipped": 0,
"failed": 0
},
"hits": {
"total": {
"value": 1160,
"relation": "eq"
},
"max_score": null,
"hits": []
},
"aggregations": {
"top_categories": {
"doc_count_error_upper_bound": 0,
"sum_other_doc_count": 552,
"buckets": [
{
"key": "Men's Clothing",
"doc_count": 817,
"top_doc": {
"hits": {
"total": {
"value": 817,
"relation": "eq"
},
"max_score": 0.998253,
"hits": [
{
"_index": "opensearch_dashboards_sample_data_ecommerce",
"_id": "34N5u50BpPQaFxReicgP",
"_score": 0.998253,
"_source": {
"products": [
{
"product_name": "Shirt - white"
},
{
"product_name": "Shirt - white"
}
]
},
"highlight": {
"products.product_name": [
"Shirt - white",
"Shirt - white"
]
}
}
]
}
}
},
{
"key": "Women's Clothing",
"doc_count": 343,
"top_doc": {
"hits": {
"total": {
"value": 343,
"relation": "eq"
},
"max_score": 0.91741234,
"hits": [
{
"_index": "opensearch_dashboards_sample_data_ecommerce",
"_id": "54N5u50BpPQaFxReiccP",
"_score": 0.91741234,
"_source": {
"products": [
{
"product_name": "Shirt - white"
},
{
"product_name": "Shirt - light blue denim"
}
]
},
"highlight": {
"products.product_name": [
"Shirt - white",
"Shirt - light blue denim"
]
}
}
]
}
}
}
]
}
}
}
예제: 스크립트 필드 사용하기 (Using script fields)
다음 예제는 카테고리별 최고가 주문을 검색하고 script_fields 정의를 사용해 15% 할인을 계산해요:
GET /opensearch_dashboards_sample_data_ecommerce/_search
{
"size": 0,
"aggs": {
"top_categories": {
"terms": {
"field": "category.keyword",
"size": 2
},
"aggs": {
"top_doc": {
"top_hits": {
"size": 1,
"sort": [
{
"taxful_total_price": {
"order": "desc"
}
}
],
"_source": {
"includes": ["customer_full_name", "taxful_total_price"]
},
"script_fields": {
"price_with_tax_discount": {
"script": {
"source": "doc['taxful_total_price'].value * 0.85"
}
}
}
}
}
}
}
}
}
예제 응답:
{
"took": 51,
"timed_out": false,
"_shards": {
"total": 1,
"successful": 1,
"skipped": 0,
"failed": 0
},
"hits": {
"total": {
"value": 4675,
"relation": "eq"
},
"max_score": null,
"hits": []
},
"aggregations": {
"top_categories": {
"doc_count_error_upper_bound": 0,
"sum_other_doc_count": 3482,
"buckets": [
{
"key": "Men's Clothing",
"doc_count": 2024,
"top_doc": {
"hits": {
"total": {
"value": 2024,
"relation": "eq"
},
"max_score": null,
"hits": [
{
"_index": "opensearch_dashboards_sample_data_ecommerce",
"_id": "LoN5u50BpPQaFxRehr9T",
"_score": null,
"_source": {
"customer_full_name": "Wagdi Shaw",
"taxful_total_price": 2249.92
},
"fields": {
"price_with_tax_discount": [
1912.5
]
},
"sort": [
2250.0
]
}
]
}
}
},
{
"key": "Women's Clothing",
"doc_count": 1903,
"top_doc": {
"hits": {
"total": {
"value": 1903,
"relation": "eq"
},
"max_score": null,
"hits": [
{
"_index": "opensearch_dashboards_sample_data_ecommerce",
"_id": "z4N5u50BpPQaFxReicuj",
"_score": null,
"_source": {
"customer_full_name": "Elyssa Hart",
"taxful_total_price": 343.96
},
"fields": {
"price_with_tax_discount": [
292.4
]
},
"sort": [
344.0
]
}
]
}
}
}
]
}
}
}
응답 본문 필드 (Response body fields)
다음 표는 각 top_hits 집계 결과 내에 반환되는 응답 본문 필드를 나열해요.
| 필드 | 데이터 타입 | 설명 |
|---|---|---|
hits.total.value |
Integer | 버킷 내에서 집계와 일치하는 총 문서 수예요. |
hits.total.relation |
String | 총 개수가 정확한지(eq) 또는 하한인지(gte)를 나타내요. |
hits.max_score |
Float 또는 Null | 반환된 적중 중 가장 높은 관련성 점수예요. 적중이 _score가 아닌 다른 필드로 정렬되면 null이에요. |
hits.hits |
Array | 버킷의 상위 일치 문서 배열이에요. |
hits.hits._index |
String | 문서를 포함하는 인덱스예요. |
hits.hits._id |
String | 문서의 고유 식별자예요. |
hits.hits._score |
Float 또는 Null | 문서의 관련성 점수예요. _score가 아닌 다른 필드로 정렬하면 null이에요. |
hits.hits._source |
Object | 원본 문서 소스예요. 소스 필터링이 적용되면 요청한 필드만 반환돼요. |
hits.hits.sort |
Array | 이 적중을 정렬하는 데 사용된 정렬 값으로, 명시적 정렬이 지정된 경우에만 존재해요. |
hits.hits._nested |
Object | 중첩 적중에만 존재해요. field(중첩 배열 필드 이름)와 offset(배열 내 0부터 시작하는 위치)을 포함해요. |