Rank feature 쿼리
Rank feature 쿼리
문서의 숫자 값(관련성 점수, 인기, 최신성 등)을 기반으로 문서 점수를 상향(boost)하려면 rank_feature 쿼리를 사용해요. 숫자 신호만을 다루므로 bool 같은 복합 쿼리에서 다른 쿼리와 함께 쓸 때 가장 효과적이에요.
출처: 문서
본문
rank_feature 쿼리는 문서의 숫자 값(관련성 점수, 인기, 최신성 등)을 기반으로 문서 점수를 상향(boost)해요. 숫자적 특성을 사용해 관련성 순위를 세밀하게 조정하고 싶을 때 이상적인 쿼리예요. 전문(full-text) 쿼리와 달리 rank_feature는 순수하게 숫자 신호 하나에만 집중해요. bool 같은 복합 쿼리에서 다른 쿼리와 함께 쓸 때 가장 효과적이에요. rank_feature 쿼리는 대상 필드가 rank_feature 필드 타입으로 매핑되어 있어야 해요. 그래야 빠르고 효율적인 상향을 위한 내부 최적화 점수 산정이 가능해요. 점수 영향은 필드 값과 선택적으로 사용하는 saturation, log, sigmoid 함수에 따라 달라져요. 이 함수들은 쿼리 시점에 동적으로 적용되어 최종 문서 점수를 계산해요. 문서 자체의 값을 바꾸거나 저장하지 않아요.
파라미터 (Parameters)
rank_feature 쿼리는 다음 파라미터를 지원해요.
Parameter Data type Required/Optional Description
field | 문자열(String) | 필수 | 문서 점수 산정에 기여하는 rank_feature 또는 rank_features 필드예요.
boost | 실수(Float) | 선택 | 점수에 적용되는 배수예요. 기본값은 1.0이에요. 0과 1 사이의 값은 점수를 줄이고, 1보다 큰 값은 점수를 키워요.
saturation | 객체(Object) | 선택 | 특성(feature) 값에 saturation 함수를 적용해요. 값이 클수록 상향도 커지지만 기준점(pivot)을 넘으면 완만해져요. 다른 함수를 제공하지 않을 때의 기본 함수예요. saturation, log, sigmoid 중 한 번에 하나만 사용할 수 있어요.
log | 객체(Object) | 선택 | 필드 값을 기반으로 하는 로그 점수 함수를 사용해요. 값의 범위가 클 때 가장 좋아요. saturation, log, sigmoid 중 한 번에 하나만 사용할 수 있어요.
sigmoid | 객체(Object) | 선택 | pivot과 exponent로 제어되는 시그모이드(S자형) 곡선을 점수 영향에 적용해요. saturation, log, sigmoid 중 한 번에 하나만 사용할 수 있어요.
positive_score_impact | 불리언(Boolean) | 선택 | false이면 낮은 값이 더 높은 점수를 받아요. 가격처럼 작을수록 좋은 특성에 유용해요. 매핑의 일부로 정의해요. 기본값은 true예요.
예제 (Example)
다음 예제는 문서 점수 산정에 영향을 주도록 rank_feature 필드를 정의하고 사용하는 방법을 보여줘요.
rank feature 필드가 있는 인덱스 생성
인기(popularity) 같은 신호를 나타내도록 rank_feature 필드가 있는 인덱스를 정의하세요:
PUT /products
{
"mappings": {
"properties": {
"title": { "type": "text" },
"popularity": { "type": "rank_feature" }
}
}
}
예제 문서 인덱싱
다양한 인기 값을 가진 샘플 상품을 추가하세요:
POST /products/_bulk
{ "index": { "_id": 1 } }
{ "title": "Wireless Earbuds", "popularity": 1 }
{ "index": { "_id": 2 } }
{ "title": "Bluetooth Speaker", "popularity": 10 }
{ "index": { "_id": 3 } }
{ "title": "Portable Charger", "popularity": 25 }
{ "index": { "_id": 4 } }
{ "title": "Smartwatch", "popularity": 50 }
{ "index": { "_id": 5 } }
{ "title": "Noise Cancelling Headphones", "popularity": 100 }
{ "index": { "_id": 6 } }
{ "title": "Gaming Laptop", "popularity": 250 }
{ "index": { "_id": 7 } }
{ "title": "4K Monitor", "popularity": 500 }
기본 rank feature 쿼리
rank_feature를 사용해 인기 점수를 기반으로 결과를 상향할 수 있어요:
POST /products/_search
{
"query": {
"rank_feature": {
"field": "popularity"
}
}
}
이 쿼리는 단독으로는 필터링을 수행하지 않아요. 대신 popularity 값에 따라 모든 문서에 점수를 매겨요. 값이 높을수록 점수도 높아져요:
{
...
"hits": {
"total": {
"value": 7,
"relation": "eq"
},
"max_score": 0.9252834,
"hits": [
{
"_index": "products",
"_id": "7",
"_score": 0.9252834,
"_source": {
"title": "4K Monitor",
"popularity": 500
}
},
{
"_index": "products",
"_id": "6",
"_score": 0.86095566,
"_source": {
"title": "Gaming Laptop",
"popularity": 250
}
},
{
"_index": "products",
"_id": "5",
"_score": 0.71237755,
"_source": {
"title": "Noise Cancelling Headphones",
"popularity": 100
}
},
{
"_index": "products",
"_id": "4",
"_score": 0.5532503,
"_source": {
"title": "Smartwatch",
"popularity": 50
}
},
{
"_index": "products",
"_id": "3",
"_score": 0.38240916,
"_source": {
"title": "Portable Charger",
"popularity": 25
}
},
{
"_index": "products",
"_id": "2",
"_score": 0.19851118,
"_source": {
"title": "Bluetooth Speaker",
"popularity": 10
}
},
{
"_index": "products",
"_id": "1",
"_score": 0.024169207,
"_source": {
"title": "Wireless Earbuds",
"popularity": 1
}
}
]
}
}
전문 검색과 결합
관련 결과를 필터링하고 인기도에 따라 상향하려면 다음 요청을 사용하세요. 이 쿼리는 “headphones”와 일치하는 모든 문서에 순위를 매기고 인기가 더 높은 문서를 상향해요:
POST /products/_search
{
"query": {
"bool": {
"must": {
"match": {
"title": "headphones"
}
},
"should": {
"rank_feature": {
"field": "popularity"
}
}
}
}
}
Boost 파라미터
boost 파라미터는 rank_feature 절의 점수 기여도를 조절할 수 있게 해줘요. 숫자 필드(예: popularity, freshness, relevance score)가 최종 문서 순위에 미치는 영향력을 제어하고 싶을 때, bool 같은 복합 쿼리에서 특히 유용해요. 다음 예제에서 bool 쿼리는 title에 “headphones”라는 용어가 있는 문서를 일치시키고, boost가 2.0인 rank_feature 절을 사용해 더 인기 있는 결과를 상향해요. 이로써 rank_feature 점수가 전체 문서 점수에 기여하는 비중이 두 배가 돼요:
POST /products/_search
{
"query": {
"bool": {
"must": {
"match": {
"title": "headphones"
}
},
"should": {
"rank_feature": {
"field": "popularity",
"boost": 2.0
}
}
}
}
}
점수 함수 구성
기본적으로 rank_feature 쿼리는 필드에서 도출된 pivot 값을 가진 saturation 함수를 사용해요. 함수를 saturation, log, sigmoid로 명시적으로 설정할 수 있어요.
Saturation 함수
saturation 함수는 rank_feature 쿼리에서 사용되는 기본 점수 산정 방식이에요. 특성 값이 클수록 문서에 더 높은 점수를 매기지만, 값이 지정된 pivot을 넘어가면 점수 증가가 점점 더 완만해져요. 매우 큰 값에 대해 체감 효과(diminishing returns)를 주고 싶을 때 유용해요. 예를 들어 인기를 상향하면서도 지나치게 큰 숫자를 과도하게 보상하지 않도록 할 때 말이에요. 점수 계산 공식은 value of the rank_feature field / (value of the rank_feature field + pivot)이에요. 생성되는 점수는 항상 0과 1 사이예요. pivot을 제공하지 않으면 인덱스에 있는 모든 rank_feature 값의 근사 기하 평균을 사용해요. 다음 예제는 pivot이 50인 saturation을 사용해요:
POST /products/_search
{
"query": {
"rank_feature": {
"field": "popularity",
"saturation": {
"pivot": 50
}
}
}
}
pivot은 점수 증가가 느려지는 지점을 정의해요. pivot보다 높은 값은 여전히 점수를 높이지만 체감 효과가 있어요. 반환된 hits에서 확인할 수 있어요:
{
...
"hits": {
"total": {
"value": 7,
"relation": "eq"
},
"max_score": 0.9090909,
"hits": [
{
"_index": "products",
"_id": "7",
"_score": 0.9090909,
"_source": {
"title": "4K Monitor",
"popularity": 500
}
},
{
"_index": "products",
"_id": "6",
"_score": 0.8333333,
"_source": {
"title": "Gaming Laptop",
"popularity": 250
}
},
{
"_index": "products",
"_id": "5",
"_score": 0.6666666,
"_source": {
"title": "Noise Cancelling Headphones",
"popularity": 100
}
},
{
"_index": "products",
"_id": "4",
"_score": 0.5,
"_source": {
"title": "Smartwatch",
"popularity": 50
}
},
{
"_index": "products",
"_id": "3",
"_score": 0.3333333,
"_source": {
"title": "Portable Charger",
"popularity": 25
}
},
{
"_index": "products",
"_id": "2",
"_score": 0.16666669,
"_source": {
"title": "Bluetooth Speaker",
"popularity": 10
}
},
{
"_index": "products",
"_id": "1",
"_score": 0.019607842,
"_source": {
"title": "Wireless Earbuds",
"popularity": 1
}
}
]
}
}
Log 함수
log 함수는 rank_feature 필드에 값의 범위가 넓을 때 유용해요. 점수에 로그 척도를 적용해 극도로 높은 값의 영향을 줄이고, 넓은 값 분포에서 점수 산정을 정규화해요. 낮은 값 사이의 작은 차이가 높은 값 사이의 큰 차이보다 더 큰 영향을 줘야 할 때 특히 좋아요. 점수는 log(scaling_factor + rank_feature field) 공식으로 계산돼요. 다음 예제는 scaling_factor가 2입니다.
POST /products/_search
{
"query": {
"rank_feature": {
"field": "popularity",
"log": {
"scaling_factor": 2
}
}
}
}
예제 데이터셋에서 popularity 필드는 1에서 500까지 범위예요. log 함수는 250, 500 같은 큰 값의 점수 기여도를 압축하면서도, 10이나 25 값을 가진 문서가 의미 있는 점수를 받도록 해줘요. 반면 saturation 함수를 적용하면 pivot보다 큰 문서는 빠르게 같은 최대 점수에 수렴해요:
{
...
"hits": {
"total": {
"value": 7,
"relation": "eq"
},
"max_score": 6.2186003,
"hits": [
{
"_index": "products",
"_id": "7",
"_score": 6.2186003,
"_source": {
"title": "4K Monitor",
"popularity": 500
}
},
{
"_index": "products",
"_id": "6",
"_score": 5.529429,
"_source": {
"title": "Gaming Laptop",
"popularity": 250
}
},
{
"_index": "products",
"_id": "5",
"_score": 4.624973,
"_source": {
"title": "Noise Cancelling Headphones",
"popularity": 100
}
},
{
"_index": "products",
"_id": "4",
"_score": 3.9512436,
"_source": {
"title": "Smartwatch",
"popularity": 50
}
},
{
"_index": "products",
"_id": "3",
"_score": 3.295837,
"_source": {
"title": "Portable Charger",
"popularity": 25
}
},
{
"_index": "products",
"_id": "2",
"_score": 2.4849067,
"_source": {
"title": "Bluetooth Speaker",
"popularity": 10
}
},
{
"_index": "products",
"_id": "1",
"_score": 1.0986123,
"_source": {
"title": "Wireless Earbuds",
"popularity": 1
}
}
]
}
}
Sigmoid 함수
sigmoid 함수는 부드러운 S자형 점수 곡선을 제공해요. 점수 영향의 가파른 정도와 중간점을 제어하고 싶을 때 특히 유용해요. 점수는 rank feature field value^exp / (rank feature field value^exp + pivot^exp) 공식으로 도출돼요. 다음 예제는 구성된 pivot과 exponent를 가진 sigmoid 함수를 사용해요. pivot은 점수가 0.5가 되는 값을 정의해요. exponent는 곡선이 얼마나 가파른지 제어해요. 값이 낮을수록 pivot 주변에서 더 급격한 전환이 일어나요:
POST /products/_search
{
"query": {
"rank_feature": {
"field": "popularity",
"sigmoid": {
"pivot": 50,
"exponent": 0.5
}
}
}
}
sigmoid 함수는 pivot 주변(이 예제에서는 50)에서 점수를 부드럽게 상향해요. pivot에 가까운 값에는 적당한 선호도를 주면서도 높고 낮은 양쪽 극단을 평평하게 만들어요:
{
...
"hits": {
"total": {
"value": 7,
"relation": "eq"
},
"max_score": 0.7597469,
"hits": [
{
"_index": "products",
"_id": "7",
"_score": 0.7597469,
"_source": {
"title": "4K Monitor",
"popularity": 500
}
},
{
"_index": "products",
"_id": "6",
"_score": 0.690983,
"_source": {
"title": "Gaming Laptop",
"popularity": 250
}
},
{
"_index": "products",
"_id": "5",
"_score": 0.58578646,
"_source": {
"title": "Noise Cancelling Headphones",
"popularity": 100
}
},
{
"_index": "products",
"_id": "4",
"_score": 0.5,
"_source": {
"title": "Smartwatch",
"popularity": 50
}
},
{
"_index": "products",
"_id": "3",
"_score": 0.41421357,
"_source": {
"title": "Portable Charger",
"popularity": 25
}
},
{
"_index": "products",
"_id": "2",
"_score": 0.309017,
"_source": {
"title": "Bluetooth Speaker",
"popularity": 10
}
},
{
"_index": "products",
"_id": "1",
"_score": 0.12389934,
"_source": {
"title": "Wireless Earbuds",
"popularity": 1
}
}
]
}
}
점수 영향 반전
기본적으로 값이 높을수록 점수가 높아져요. 값이 낮을수록 점수가 높아지길 원한다면(예: 더 낮은 가격이 더 관련성이 높은 경우) 인덱스 생성 시 positive_score_impact를 false로 설정하세요:
PUT /products_new
{
"mappings": {
"properties": {
"popularity": {
"type": "rank_feature",
"positive_score_impact": false
}
}
}
}
ExampleCreate an index with a rank feature field