Composite 집계

Composite 집계

composite 집계는 하나 이상의 문서 필드 또는 소스(source)를 기반으로 버킷을 만들어요. composite 집계는 각 개별 소스 값의 모든 조합에 대해 버킷을 만들어요. 기본적으로 하나 이상의 개별 필드에서 값이 누락된 조합은 결과에서 제외돼요.

각 소스는 다음 네 가지 유형의 집계 중 하나를 가져요:

  • terms 타입은 고유한(보통 String) 값으로 그룹화해요.
  • histogram 타입은 지정된 너비의 버킷으로 숫자를 그룹화해요.
  • date_histogram 타입은 지정된 너비의 날짜 또는 시간 범위로 그룹화해요.
  • geotile_grid 타입은 지정된 해상도의 그리드로 지리적 포인트를 그룹화해요.

composite 집계는 소스 키를 버킷으로 결합하는 방식으로 동작해요. 결과 버킷은 소스 간 및 소스 내에서 모두 정렬돼요:

  • 소스 간(Across): 버킷은 집계 요청에서 소스가 정렬된 순서대로 중첩돼요.
  • 소스 내(Within): 각 소스의 값 순서가 해당 소스의 버킷 순서를 결정해요. 정렬은 소스 타입에 따라 알파벳, 숫자, 날짜-시간, 또는 지리 타일 순서로 이루어져요.

마라톤 참가자 인덱스의 다음 필드들을 생각해보세요:

{... "city": "Albuquerque", "place": "Bronze" ...}
{... "city": "Boston",  ...}
{... "city": "Chicago", "place": "Bronze" ...}
{... "city": "Albuquerque", "place": "Gold" ...}
{... "city": "Chicago", "place": "Silver" ...}
{... "city": "Boston", "place": "Bronze" ...}
{... "city": "Chicago", "place": "Gold" ...}

요청에서 소스를 다음과 같이 지정한다고 가정해요:

    ...
    "sources": [
        { "marathon_city": { "terms": { "field": "city" }}},
        { "participant_medal": { "terms": { "field": "place" }}}
    ],
    ...

각 소스에 고유한 키 이름을 할당해야 해요.

결과 composite은 다음 버킷을 순서대로 포함해요:

{ "city": "Albuquerque", "place": "Bronze" }
{ "city": "Albuquerque", "place": "Gold" }
{ "city": "Boston", "place": "Bronze" }
{ "city": "Boston", "place": "Silver" }
{ "city": "Chicago", "place": "Bronze" }
{ "city": "Chicago", "place": "Gold" }
{ "city": "Chicago", "place": "Silver" }

city와 place 필드가 모두 알파벳순으로 정렬되어 있는 점을 주목하세요.

출처: 문서

본문

파라미터 (Parameters)

composite 집계는 다음 파라미터를 받아요.

파라미터 필수/선택 데이터 타입 설명
sources 필수 Array 소스 객체의 배열. 유효한 타입은 terms, histogram, date_histogram, geotile_grid예요.
size 선택 Numeric 결과로 반환할 composite 버킷 수. 기본값은 10이에요. Paginating composite results를 참고하세요.
after 선택 String 페이지 처리된 composite 버킷 표시를 재개할 위치를 지정하는 키. Paginating composite results를 참고하세요.
order 선택 String 각 소스에 대해 값을 오름차순 또는 내림차순으로 정렬할지 여부. 유효한 값은 asc와 desc예요. 기본값은 asc예요.
missing_bucket 선택 Boolean 각 소스에 대해 값이 없는 문서를 포함할지 여부. 기본값은 false예요. true로 설정하면 OpenSearch가 문서를 포함하고 필드 키에 null을 제공해요. null 값은 오름차순에서 가장 먼저 순위가 매겨져요.

집계별 파라미터는 해당 집계 문서를 참고하세요.

Terms

문자열 또는 Boolean 데이터를 집계할 때는 terms 집계를 사용해요. 자세한 내용은 Terms aggregations를 참고하세요.

terms 소스를 사용하면 모든 유형의 데이터에 대해 composite 버킷을 만들 수 있어요. 하지만 terms 소스는 모든 고유 값에 대해 버킷을 만들기 때문에, 숫자 데이터에는 보통 histogram 소스를 사용해요.

다음 예제 요청은 OpenSearch Dashboards 샘플 전자상거래 데이터에서 요일과 고객 성별에 대한 첫 4개의 composite 버킷을 반환해요:

GET opensearch_dashboards_sample_data_ecommerce/_search
{
  "size": 0,
  "aggs": {
    "composite_buckets": {
      "composite": {
        "sources": [
          { "day": { "terms": { "field": "day_of_week" }}},
          { "gender": { "terms": { "field": "customer_gender" }}}
        ],
        "size": 4
      }
    }
  }
}

이 예제의 데이터셋은 모든 버킷에 유효한 데이터를 포함하므로, 집계는 성별과 요일의 모든 조합에 대해 버킷을 생성해 총 14개의 버킷을 만들어요.

요청에서 size를 4로 지정했으므로 응답에는 처음 네 개의 composite 버킷이 포함돼요. 소스가 terms이므로 버킷은 소스 간 및 소스 내에서 모두 오름차순 알파벳순으로 정렬돼요:

{
  "took": 51,
  "timed_out": false,
  "_shards": {
    "total": 1,
    "successful": 1,
    "skipped": 0,
    "failed": 0
  },
  "hits": {
    "total": {
      "value": 4675,
      "relation": "eq"
    },
    "max_score": null,
    "hits": []
  },
  "aggregations": {
    "composite_buckets": {
      "after_key": {
        "day": "Monday",
        "gender": "MALE"
      },
      "buckets": [
        {
          "key": {
            "day": "Friday",
            "gender": "FEMALE"
          },
          "doc_count": 399
        },
        {
          "key": {
            "day": "Friday",
            "gender": "MALE"
          },
          "doc_count": 371
        },
        {
          "key": {
            "day": "Monday",
            "gender": "FEMALE"
          },
          "doc_count": 320
        },
        {
          "key": {
            "day": "Monday",
            "gender": "MALE"
          },
          "doc_count": 259
        }
      ]
    }
  }
}

응답에서 반환된 after_key를 사용하면 더 많은 결과를 볼 수 있어요. 다음 섹션의 예제를 참고하세요.

Histogram

숫자 데이터의 composite 집계를 만들 때는 histogram 소스를 사용해요. 자세한 내용은 Histogram aggregations를 참고하세요.

histogram 소스의 경우 각 composite 버킷 키에 사용되는 이름은 키의 히스토그램 간격에서 가장 낮은 값이에요. 각 소스 히스토그램 간격은 [lower_bound, lower_bound + interval) 범위의 값을 포함해요. 첫 번째 간격의 이름은 (오름차순 값 소스의 경우) 소스 필드에서 가장 낮은 값이에요.

다음 예제 요청은 OpenSearch Dashboards 샘플 전자상거래 데이터에서 수량과 기본 단가에 대해 각각 1과 50의 버킷 너비를 기준으로 첫 6개의 composite 버킷을 반환해요:

GET opensearch_dashboards_sample_data_ecommerce/_search
{
  "size": 0,
  "aggs": {
    "composite_buckets": {
      "composite": {
        "sources": [
          { "quantity": { "histogram": { "field": "products.quantity", "interval": 1 }}},
          { "unit_price": { "histogram": { "field": "products.base_unit_price", "interval": 50 }}}
        ],
        "size": 6
      }
    }
  }
}

집계는 두 histogram 소스에 대한 첫 6개의 버킷 키와 문서 수를 반환해요. terms 예제와 마찬가지로 버킷은 소스 필드 간 및 내에서 정렬돼요. 다만 이 경우 순서는 숫자 순서이며 각 히스토그램 너비의 포함적 하한(inclusive lower bound)을 기준으로 해요:

{
  "took": 11,
  "timed_out": false,
  "_shards": {
    "total": 1,
    "successful": 1,
    "skipped": 0,
    "failed": 0
  },
  "hits": {
    "total": {
      "value": 4675,
      "relation": "eq"
    },
    "max_score": null,
    "hits": []
  },
  "aggregations": {
    "composite_buckets": {
      "after_key": {
        "quantity": 2,
        "unit_price": 150
      },
      "buckets": [
        {
          "key": {
            "quantity": 1,
            "unit_price": 0
          },
          "doc_count": 17691
        },
        {
          "key": {
            "quantity": 1,
            "unit_price": 50
          },
          "doc_count": 5014
        },
        {
          "key": {
            "quantity": 1,
            "unit_price": 100
          },
          "doc_count": 482
        },
        {
          "key": {
            "quantity": 1,
            "unit_price": 150
          },
          "doc_count": 148
        },
        {
          "key": {
            "quantity": 1,
            "unit_price": 200
          },
          "doc_count": 32
        },
        {
          "key": {
            "quantity": 2,
            "unit_price": 150
          },
          "doc_count": 4
        }
      ]
    }
  }
}

각 필드의 버킷 키는 필드 간격의 하한이에요. 예를 들어 첫 composite 버킷의 unit_price 키는 0이에요.

다음 6개의 버킷을 검색하려면 응답의 after_key 객체를 after 파라미터에 다음과 같이 전달해요:

GET opensearch_dashboards_sample_data_ecommerce/_search
{
  "size": 0,
  "aggs": {
    "composite_buckets": {
      "composite": {
        "sources": [
          { "quantity": { "histogram": { "field": "products.quantity", "interval": 1 }}},
          { "unit_price": { "histogram": { "field": "products.base_unit_price", "interval": 50 }}}
        ],
        "size": 6,
        "after": {
            "quantity": 2,
            "unit_price": 150
        }
      }
    }
  }
}

이제 남은 버킷은 두 개뿐이에요:

{
  "took": 12,
  "timed_out": false,
  "_shards": {
    "total": 1,
    "successful": 1,
    "skipped": 0,
    "failed": 0
  },
  "hits": {
    "total": {
      "value": 4675,
      "relation": "eq"
    },
    "max_score": null,
    "hits": []
  },
  "aggregations": {
    "composite_buckets": {
      "after_key": {
        "quantity": 2,
        "unit_price": 500
      },
      "buckets": [
        {
          "key": {
            "quantity": 2,
            "unit_price": 200
          },
          "doc_count": 8
        },
        {
          "key": {
            "quantity": 2,
            "unit_price": 500
          },
          "doc_count": 4
        }
      ]
    }
  }
}

Date histogram

날짜 범위의 composite 집계를 만들려면 date_histogram 집계를 사용해요. 자세한 내용은 Date histogram aggregations를 참고하세요.

OpenSearch는 date_interval 버킷 키를 포함한 날짜를 Unix 시간에서 epoch 이후의 밀리초를 나타내는 long 정수로 표현해요. format 파라미터를 사용해 날짜 출력 형식을 지정할 수 있어요. 이렇게 해도 키 순서는 바뀌지 않아요.

OpenSearch는 날짜-시간을 UTC로 저장해요. time_zone 파라미터를 사용해 다른 시간대로 출력 결과를 표시할 수 있어요.

다음 예제 요청은 OpenSearch Dashboards 샘플 전자상거래 데이터에서 각 판매 제품이 생성된 연도와 판매된 날짜에 대해 각각 1년과 1일의 버킷 너비를 기준으로 첫 4개의 composite 버킷을 반환해요:

GET opensearch_dashboards_sample_data_ecommerce/_search
{
  "size": 0,
  "aggs": {
    "composite_buckets": {
      "composite": {
        "sources": [
          { "product_creation_date": { "date_histogram": { "field": "products.created_on", "calendar_interval": "1y", "format": "yyyy" }}},
          { "order_date": { "date_histogram": { "field": "order_date", "calendar_interval": "1d", "format": "yyyy-MM-dd" }}}
        ],
        "size": 4
      }
    }
  }
}

집계는 형식이 지정된 날짜 기반 버킷 키와 개수를 반환해요. date_interval composite 집계의 경우 필드 순서는 날짜 기준이에요:

{
  "took": 21,
  "timed_out": false,
  "_shards": {
    "total": 1,
    "successful": 1,
    "skipped": 0,
    "failed": 0
  },
  "hits": {
    "total": {
      "value": 4675,
      "relation": "eq"
    },
    "max_score": null,
    "hits": []
  },
  "aggregations": {
    "composite_buckets": {
      "after_key": {
        "product_creation_date": "2016",
        "order_date": "2025-02-23"
      },
      "buckets": [
        {
          "key": {
            "product_creation_date": "2016",
            "order_date": "2025-02-20"
          },
          "doc_count": 146
        },
        {
          "key": {
            "product_creation_date": "2016",
            "order_date": "2025-02-21"
          },
          "doc_count": 153
        },
        {
          "key": {
            "product_creation_date": "2016",
            "order_date": "2025-02-22"
          },
          "doc_count": 143
        },
        {
          "key": {
            "product_creation_date": "2016",
            "order_date": "2025-02-23"
          },
          "doc_count": 140
        }
      ]
    }
  }
}

Geotile grid

geo_point 값을 지도 타일을 나타내는 버킷으로 집계할 때는 geotile_grid 소스를 사용해요. 다른 composite 집계 소스와 마찬가지로 기본적으로 결과에는 데이터가 포함된 버킷만 담겨요. 자세한 내용은 Geotile grid aggregations를 참고하세요.

각 셀은 지도 타일에 해당해요. 셀 레이블은 {zoom}/{x}/{y} 형식을 사용해요.

다음 예제 요청은 정밀도 8에서 geoip.location 필드의 위치를 포함하는 첫 6개의 타일을 반환해요:

GET opensearch_dashboards_sample_data_ecommerce/_search
{
  "size": 0,
  "aggs": {
    "composite_buckets": {
      "composite": {
        "sources": [
          { "tile": { "geotile_grid": { "field": "geoip.location", "precision": 8 } } }
        ],
        "size": 6
      }
    }
  }
}

집계는 지정된 geo_tile과 포인트 수를 반환해요:

{
  "took": 3,
  "timed_out": false,
  "_shards": {
    "total": 1,
    "successful": 1,
    "skipped": 0,
    "failed": 0
  },
  "hits": {
    "total": {
      "value": 4675,
      "relation": "eq"
    },
    "max_score": null,
    "hits": []
  },
  "aggregations": {
    "composite_buckets": {
      "after_key": {
        "tile": "8/122/104"
      },
      "buckets": [
        {
          "key": {
            "tile": "8/43/102"
          },
          "doc_count": 310
        },
        {
          "key": {
            "tile": "8/75/96"
          },
          "doc_count": 896
        },
        {
          "key": {
            "tile": "8/75/124"
          },
          "doc_count": 178
        },
        {
          "key": {
            "tile": "8/122/104"
          },
          "doc_count": 408
        }
      ]
    }
  }
}

소스 결합 (Combining sources)

서로 다른 유형의 소스 두 개 이상을 결합할 수 있어요.

다음 예제 요청은 세 가지 서로 다른 소스 유형으로 구성된 버킷을 반환해요:

GET opensearch_dashboards_sample_data_ecommerce/_search
{
  "size": 0,
  "aggs": {
    "composite_buckets": {
      "composite": {
        "sources": [
          { "order_date": { "date_histogram": { "field": "order_date", "calendar_interval": "1M", "format": "yyyy-MM" }}},
          { "gender": { "terms": { "field": "customer_gender" }}},          
          { "unit_price": { "histogram": { "field": "products.base_unit_price", "interval": 200 }}}
        ],
        "size": 10
      }
    }
  }
}

집계는 혼합 타입의 composite 버킷과 문서 수를 반환해요:

{
  "took": 11,
  "timed_out": false,
  "_shards": {
    "total": 1,
    "successful": 1,
    "skipped": 0,
    "failed": 0
  },
  "hits": {
    "total": {
      "value": 4675,
      "relation": "eq"
    },
    "max_score": null,
    "hits": []
  },
  "aggregations": {
    "composite_buckets": {
      "after_key": {
        "order_date": "2025-03",
        "gender": "MALE",
        "unit_price": 200
      },
      "buckets": [
        {
          "key": {
            "order_date": "2025-02",
            "gender": "FEMALE",
            "unit_price": 0
          },
          "doc_count": 1517
        },
        {
          "key": {
            "order_date": "2025-02",
            "gender": "MALE",
            "unit_price": 0
          },
          "doc_count": 1369
        },
        {
          "key": {
            "order_date": "2025-02",
            "gender": "MALE",
            "unit_price": 200
          },
          "doc_count": 6
        },
        {
          "key": {
            "order_date": "2025-02",
            "gender": "MALE",
            "unit_price": 400
          },
          "doc_count": 1
        },
        {
          "key": {
            "order_date": "2025-03",
            "gender": "FEMALE",
            "unit_price": 0
          },
          "doc_count": 3656
        },
        {
          "key": {
            "order_date": "2025-03",
            "gender": "FEMALE",
            "unit_price": 200
          },
          "doc_count": 1
        },
        {
          "key": {
            "order_date": "2025-03",
            "gender": "MALE",
            "unit_price": 0
          },
          "doc_count": 3530
        },
        {
          "key": {
            "order_date": "2025-03",
            "gender": "MALE",
            "unit_price": 200
          },
          "doc_count": 7
        }
      ]
    }
  }
}

하위 집계 (Subaggregations)

composite 집계는 composite 버킷 안의 문서에 대한 정보를 드러내는 하위 집계와 결합할 때 가장 유용해요.

다음 예제 요청은 OpenSearch Dashboards 샘플 전자상거래 데이터에서 요일별 성별에 따른 평균 지출을 비교해요:

GET opensearch_dashboards_sample_data_ecommerce/_search
{
  "size": 0,
  "aggs": {
    "composite_buckets": {
      "composite": {
        "sources": [
          { "weekday": { "terms": { "field": "day_of_week" }}},
          { "gender": { "terms": { "field": "customer_gender" }}}          
        ],
        "size": 6
      },
      "aggs": {
        "avg_spend": {
          "avg": { "field": "taxful_total_price" }
        }
      }
    }
  }
}

집계는 처음 6개 버킷의 평균 taxful_total_price를 반환해요:

{
  "took": 30,
  "timed_out": false,
  "_shards": {
    "total": 1,
    "successful": 1,
    "skipped": 0,
    "failed": 0
  },
  "hits": {
    "total": {
      "value": 4675,
      "relation": "eq"
    },
    "max_score": null,
    "hits": []
  },
  "aggregations": {
    "composite_buckets": {
      "after_key": {
        "weekday": "Saturday",
        "gender": "MALE"
      },
      "buckets": [
        {
          "key": {
            "weekday": "Friday",
            "gender": "FEMALE"
          },
          "doc_count": 399,
          "avg_spend": {
            "value": 71.7733395989975
          }
        },
        {
          "key": {
            "weekday": "Friday",
            "gender": "MALE"
          },
          "doc_count": 371,
          "avg_spend": {
            "value": 79.72514108827494
          }
        },
        {
          "key": {
            "weekday": "Monday",
            "gender": "FEMALE"
          },
          "doc_count": 320,
          "avg_spend": {
            "value": 72.1588623046875
          }
        },
        {
          "key": {
            "weekday": "Monday",
            "gender": "MALE"
          },
          "doc_count": 259,
          "avg_spend": {
            "value": 86.1754946911197
          }
        },
        {
          "key": {
            "weekday": "Saturday",
            "gender": "FEMALE"
          },
          "doc_count": 365,
          "avg_spend": {
            "value": 73.53236301369863
          }
        },
        {
          "key": {
            "weekday": "Saturday",
            "gender": "MALE"
          },
          "doc_count": 371,
          "avg_spend": {
            "value": 72.78092360175202
          }
        }
      ]
    }
  }
}

composite 결과 페이지 처리 (Paginating composite results)

요청 결과가 size개보다 많은 버킷을 생성하면 size개만큼의 버킷이 반환돼요. 이 경우 결과에는 목록에서 다음 버킷의 키를 담은 after_key 객체가 포함돼요. 요청의 다음 size개 버킷을 검색하려면 after 파라미터에 after_key를 전달해 요청을 다시 보내요. 예제는 Histogram 섹션의 요청을 참고하세요.

페이지 처리된 응답을 이어갈 때는 마지막 버킷을 복사하지 말고 항상 after_key를 사용해요. 둘은 때때로 다를 수 있어요.

인덱스 정렬로 성능 개선 (Improving performance with index sorting)

대규모 데이터셋에서 composite 집계를 빠르게 하려면 집계 소스에서 사용한 것과 동일한 필드와 순서로 인덱스를 정렬할 수 있어요. index.sort.field와 index.sort.order가 composite 집계에 사용된 소스 필드와 순서와 일치하면 OpenSearch는 더 적은 메모리 사용으로 더 효율적으로 결과를 반환할 수 있어요. 인덱스 정렬은 색인 중에 약간의 오버헤드를 추가하지만, composite 집계의 쿼리 성능 개선 효과는 상당해요.

다음 예제 요청은 my-sorted-index 인덱스의 각 필드에 대해 정렬 필드와 정렬 순서를 설정해요:

PUT /my-sorted-index
{
  "settings": {
    "index": {
      "sort.field": ["customer_id", "timestamp"],
      "sort.order": ["asc", "desc"]
    }
  },
  "mappings": {
    "properties": {
      "customer_id": {
        "type": "keyword"
      },
      "timestamp": {
        "type": "date"
      },
      "price": {
        "type": "double"
      }
    }
  }
}

다음 요청은 my-sorted-index 인덱스에 composite 집계를 만들어요. 인덱스가 customer_id의 오름차순과 timestamp의 내림차순으로 정렬되어 있고 집계 소스가 해당 정렬 순서와 일치하므로, 이 쿼리는 더 빠르고 메모리 부담이 적게 실행돼요:

GET /my-sorted-index/_search
{
  "size": 0,
  "aggs": {
    "my_buckets": {
      "composite": {
        "size": 1000,
        "sources": [
          { "customer": { "terms": { "field": "customer_id", "order": "asc" } } },
          { "time": { "date_histogram": { "field": "timestamp", "calendar_interval": "1d", "order": "desc" } } }
        ]
      }
    }
  }
}

더 알아보기 (Learn more)