Rank feature 쿼리

Rank feature 쿼리

문서의 숫자 값(관련성 점수, 인기, 최신성 등)을 기반으로 문서 점수를 상향(boost)하려면 rank_feature 쿼리를 사용해요. 숫자 신호만을 다루므로 bool 같은 복합 쿼리에서 다른 쿼리와 함께 쓸 때 가장 효과적이에요.

출처: 문서

본문

rank_feature 쿼리는 문서의 숫자 값(관련성 점수, 인기, 최신성 등)을 기반으로 문서 점수를 상향(boost)해요. 숫자적 특성을 사용해 관련성 순위를 세밀하게 조정하고 싶을 때 이상적인 쿼리예요. 전문(full-text) 쿼리와 달리 rank_feature는 순수하게 숫자 신호 하나에만 집중해요. bool 같은 복합 쿼리에서 다른 쿼리와 함께 쓸 때 가장 효과적이에요. rank_feature 쿼리는 대상 필드가 rank_feature 필드 타입으로 매핑되어 있어야 해요. 그래야 빠르고 효율적인 상향을 위한 내부 최적화 점수 산정이 가능해요. 점수 영향은 필드 값과 선택적으로 사용하는 saturation, log, sigmoid 함수에 따라 달라져요. 이 함수들은 쿼리 시점에 동적으로 적용되어 최종 문서 점수를 계산해요. 문서 자체의 값을 바꾸거나 저장하지 않아요.

파라미터 (Parameters)

rank_feature 쿼리는 다음 파라미터를 지원해요. Parameter Data type Required/Optional Description field | 문자열(String) | 필수 | 문서 점수 산정에 기여하는 rank_feature 또는 rank_features 필드예요. boost | 실수(Float) | 선택 | 점수에 적용되는 배수예요. 기본값은 1.0이에요. 0과 1 사이의 값은 점수를 줄이고, 1보다 큰 값은 점수를 키워요. saturation | 객체(Object) | 선택 | 특성(feature) 값에 saturation 함수를 적용해요. 값이 클수록 상향도 커지지만 기준점(pivot)을 넘으면 완만해져요. 다른 함수를 제공하지 않을 때의 기본 함수예요. saturation, log, sigmoid 중 한 번에 하나만 사용할 수 있어요. log | 객체(Object) | 선택 | 필드 값을 기반으로 하는 로그 점수 함수를 사용해요. 값의 범위가 클 때 가장 좋아요. saturation, log, sigmoid 중 한 번에 하나만 사용할 수 있어요. sigmoid | 객체(Object) | 선택 | pivot과 exponent로 제어되는 시그모이드(S자형) 곡선을 점수 영향에 적용해요. saturation, log, sigmoid 중 한 번에 하나만 사용할 수 있어요. positive_score_impact | 불리언(Boolean) | 선택 | false이면 낮은 값이 더 높은 점수를 받아요. 가격처럼 작을수록 좋은 특성에 유용해요. 매핑의 일부로 정의해요. 기본값은 true예요.

예제 (Example)

다음 예제는 문서 점수 산정에 영향을 주도록 rank_feature 필드를 정의하고 사용하는 방법을 보여줘요.

rank feature 필드가 있는 인덱스 생성

인기(popularity) 같은 신호를 나타내도록 rank_feature 필드가 있는 인덱스를 정의하세요:

PUT /products
{
  "mappings": {
    "properties": {
      "title": { "type": "text" },
      "popularity": { "type": "rank_feature" }
    }
  }
}

예제 문서 인덱싱

다양한 인기 값을 가진 샘플 상품을 추가하세요:

POST /products/_bulk
{ "index": { "_id": 1 } }
{ "title": "Wireless Earbuds", "popularity": 1 }
{ "index": { "_id": 2 } }
{ "title": "Bluetooth Speaker", "popularity": 10 }
{ "index": { "_id": 3 } }
{ "title": "Portable Charger", "popularity": 25 }
{ "index": { "_id": 4 } }
{ "title": "Smartwatch", "popularity": 50 }
{ "index": { "_id": 5 } }
{ "title": "Noise Cancelling Headphones", "popularity": 100 }
{ "index": { "_id": 6 } }
{ "title": "Gaming Laptop", "popularity": 250 }
{ "index": { "_id": 7 } }
{ "title": "4K Monitor", "popularity": 500 }

기본 rank feature 쿼리

rank_feature를 사용해 인기 점수를 기반으로 결과를 상향할 수 있어요:

POST /products/_search
{
  "query": {
    "rank_feature": {
      "field": "popularity"
    }
  }
}

이 쿼리는 단독으로는 필터링을 수행하지 않아요. 대신 popularity 값에 따라 모든 문서에 점수를 매겨요. 값이 높을수록 점수도 높아져요:

{
  ...
  "hits": {
    "total": {
      "value": 7,
      "relation": "eq"
    },
    "max_score": 0.9252834,
    "hits": [
      {
        "_index": "products",
        "_id": "7",
        "_score": 0.9252834,
        "_source": {
          "title": "4K Monitor",
          "popularity": 500
        }
      },
      {
        "_index": "products",
        "_id": "6",
        "_score": 0.86095566,
        "_source": {
          "title": "Gaming Laptop",
          "popularity": 250
        }
      },
      {
        "_index": "products",
        "_id": "5",
        "_score": 0.71237755,
        "_source": {
          "title": "Noise Cancelling Headphones",
          "popularity": 100
        }
      },
      {
        "_index": "products",
        "_id": "4",
        "_score": 0.5532503,
        "_source": {
          "title": "Smartwatch",
          "popularity": 50
        }
      },
      {
        "_index": "products",
        "_id": "3",
        "_score": 0.38240916,
        "_source": {
          "title": "Portable Charger",
          "popularity": 25
        }
      },
      {
        "_index": "products",
        "_id": "2",
        "_score": 0.19851118,
        "_source": {
          "title": "Bluetooth Speaker",
          "popularity": 10
        }
      },
      {
        "_index": "products",
        "_id": "1",
        "_score": 0.024169207,
        "_source": {
          "title": "Wireless Earbuds",
          "popularity": 1
        }
      }
    ]
  }
}

전문 검색과 결합

관련 결과를 필터링하고 인기도에 따라 상향하려면 다음 요청을 사용하세요. 이 쿼리는 “headphones”와 일치하는 모든 문서에 순위를 매기고 인기가 더 높은 문서를 상향해요:

POST /products/_search
{
  "query": {
    "bool": {
      "must": {
        "match": {
          "title": "headphones"
        }
      },
      "should": {
        "rank_feature": {
          "field": "popularity"
        }
      }
    }
  }
}

Boost 파라미터

boost 파라미터는 rank_feature 절의 점수 기여도를 조절할 수 있게 해줘요. 숫자 필드(예: popularity, freshness, relevance score)가 최종 문서 순위에 미치는 영향력을 제어하고 싶을 때, bool 같은 복합 쿼리에서 특히 유용해요. 다음 예제에서 bool 쿼리는 title에 “headphones”라는 용어가 있는 문서를 일치시키고, boost가 2.0인 rank_feature 절을 사용해 더 인기 있는 결과를 상향해요. 이로써 rank_feature 점수가 전체 문서 점수에 기여하는 비중이 두 배가 돼요:

POST /products/_search
{
  "query": {
    "bool": {
      "must": {
        "match": {
          "title": "headphones"
        }
      },
      "should": {
        "rank_feature": {
          "field": "popularity",
          "boost": 2.0
        }
      }
    }
  }
}

점수 함수 구성

기본적으로 rank_feature 쿼리는 필드에서 도출된 pivot 값을 가진 saturation 함수를 사용해요. 함수를 saturation, log, sigmoid로 명시적으로 설정할 수 있어요.

Saturation 함수

saturation 함수는 rank_feature 쿼리에서 사용되는 기본 점수 산정 방식이에요. 특성 값이 클수록 문서에 더 높은 점수를 매기지만, 값이 지정된 pivot을 넘어가면 점수 증가가 점점 더 완만해져요. 매우 큰 값에 대해 체감 효과(diminishing returns)를 주고 싶을 때 유용해요. 예를 들어 인기를 상향하면서도 지나치게 큰 숫자를 과도하게 보상하지 않도록 할 때 말이에요. 점수 계산 공식은 value of the rank_feature field / (value of the rank_feature field + pivot)이에요. 생성되는 점수는 항상 0과 1 사이예요. pivot을 제공하지 않으면 인덱스에 있는 모든 rank_feature 값의 근사 기하 평균을 사용해요. 다음 예제는 pivot이 50인 saturation을 사용해요:

POST /products/_search
{
  "query": {
    "rank_feature": {
      "field": "popularity",
      "saturation": {
        "pivot": 50
      }
    }
  }
}

pivot은 점수 증가가 느려지는 지점을 정의해요. pivot보다 높은 값은 여전히 점수를 높이지만 체감 효과가 있어요. 반환된 hits에서 확인할 수 있어요:

{
  ...
  "hits": {
    "total": {
      "value": 7,
      "relation": "eq"
    },
    "max_score": 0.9090909,
    "hits": [
      {
        "_index": "products",
        "_id": "7",
        "_score": 0.9090909,
        "_source": {
          "title": "4K Monitor",
          "popularity": 500
        }
      },
      {
        "_index": "products",
        "_id": "6",
        "_score": 0.8333333,
        "_source": {
          "title": "Gaming Laptop",
          "popularity": 250
        }
      },
      {
        "_index": "products",
        "_id": "5",
        "_score": 0.6666666,
        "_source": {
          "title": "Noise Cancelling Headphones",
          "popularity": 100
        }
      },
      {
        "_index": "products",
        "_id": "4",
        "_score": 0.5,
        "_source": {
          "title": "Smartwatch",
          "popularity": 50
        }
      },
      {
        "_index": "products",
        "_id": "3",
        "_score": 0.3333333,
        "_source": {
          "title": "Portable Charger",
          "popularity": 25
        }
      },
      {
        "_index": "products",
        "_id": "2",
        "_score": 0.16666669,
        "_source": {
          "title": "Bluetooth Speaker",
          "popularity": 10
        }
      },
      {
        "_index": "products",
        "_id": "1",
        "_score": 0.019607842,
        "_source": {
          "title": "Wireless Earbuds",
          "popularity": 1
        }
      }
    ]
  }
}

Log 함수

log 함수는 rank_feature 필드에 값의 범위가 넓을 때 유용해요. 점수에 로그 척도를 적용해 극도로 높은 값의 영향을 줄이고, 넓은 값 분포에서 점수 산정을 정규화해요. 낮은 값 사이의 작은 차이가 높은 값 사이의 큰 차이보다 더 큰 영향을 줘야 할 때 특히 좋아요. 점수는 log(scaling_factor + rank_feature field) 공식으로 계산돼요. 다음 예제는 scaling_factor가 2입니다.

POST /products/_search
{
  "query": {
    "rank_feature": {
      "field": "popularity",
      "log": {
        "scaling_factor": 2
      }
    }
  }
}

예제 데이터셋에서 popularity 필드는 1에서 500까지 범위예요. log 함수는 250, 500 같은 큰 값의 점수 기여도를 압축하면서도, 10이나 25 값을 가진 문서가 의미 있는 점수를 받도록 해줘요. 반면 saturation 함수를 적용하면 pivot보다 큰 문서는 빠르게 같은 최대 점수에 수렴해요:

{
  ...
  "hits": {
    "total": {
      "value": 7,
      "relation": "eq"
    },
    "max_score": 6.2186003,
    "hits": [
      {
        "_index": "products",
        "_id": "7",
        "_score": 6.2186003,
        "_source": {
          "title": "4K Monitor",
          "popularity": 500
        }
      },
      {
        "_index": "products",
        "_id": "6",
        "_score": 5.529429,
        "_source": {
          "title": "Gaming Laptop",
          "popularity": 250
        }
      },
      {
        "_index": "products",
        "_id": "5",
        "_score": 4.624973,
        "_source": {
          "title": "Noise Cancelling Headphones",
          "popularity": 100
        }
      },
      {
        "_index": "products",
        "_id": "4",
        "_score": 3.9512436,
        "_source": {
          "title": "Smartwatch",
          "popularity": 50
        }
      },
      {
        "_index": "products",
        "_id": "3",
        "_score": 3.295837,
        "_source": {
          "title": "Portable Charger",
          "popularity": 25
        }
      },
      {
        "_index": "products",
        "_id": "2",
        "_score": 2.4849067,
        "_source": {
          "title": "Bluetooth Speaker",
          "popularity": 10
        }
      },
      {
        "_index": "products",
        "_id": "1",
        "_score": 1.0986123,
        "_source": {
          "title": "Wireless Earbuds",
          "popularity": 1
        }
      }
    ]
  }
}

Sigmoid 함수

sigmoid 함수는 부드러운 S자형 점수 곡선을 제공해요. 점수 영향의 가파른 정도와 중간점을 제어하고 싶을 때 특히 유용해요. 점수는 rank feature field value^exp / (rank feature field value^exp + pivot^exp) 공식으로 도출돼요. 다음 예제는 구성된 pivot과 exponent를 가진 sigmoid 함수를 사용해요. pivot은 점수가 0.5가 되는 값을 정의해요. exponent는 곡선이 얼마나 가파른지 제어해요. 값이 낮을수록 pivot 주변에서 더 급격한 전환이 일어나요:

POST /products/_search
{
  "query": {
    "rank_feature": {
      "field": "popularity",
      "sigmoid": {
        "pivot": 50,
        "exponent": 0.5
      }
    }
  }
}

sigmoid 함수는 pivot 주변(이 예제에서는 50)에서 점수를 부드럽게 상향해요. pivot에 가까운 값에는 적당한 선호도를 주면서도 높고 낮은 양쪽 극단을 평평하게 만들어요:

{
  ...
  "hits": {
    "total": {
      "value": 7,
      "relation": "eq"
    },
    "max_score": 0.7597469,
    "hits": [
      {
        "_index": "products",
        "_id": "7",
        "_score": 0.7597469,
        "_source": {
          "title": "4K Monitor",
          "popularity": 500
        }
      },
      {
        "_index": "products",
        "_id": "6",
        "_score": 0.690983,
        "_source": {
          "title": "Gaming Laptop",
          "popularity": 250
        }
      },
      {
        "_index": "products",
        "_id": "5",
        "_score": 0.58578646,
        "_source": {
          "title": "Noise Cancelling Headphones",
          "popularity": 100
        }
      },
      {
        "_index": "products",
        "_id": "4",
        "_score": 0.5,
        "_source": {
          "title": "Smartwatch",
          "popularity": 50
        }
      },
      {
        "_index": "products",
        "_id": "3",
        "_score": 0.41421357,
        "_source": {
          "title": "Portable Charger",
          "popularity": 25
        }
      },
      {
        "_index": "products",
        "_id": "2",
        "_score": 0.309017,
        "_source": {
          "title": "Bluetooth Speaker",
          "popularity": 10
        }
      },
      {
        "_index": "products",
        "_id": "1",
        "_score": 0.12389934,
        "_source": {
          "title": "Wireless Earbuds",
          "popularity": 1
        }
      }
    ]
  }
}

점수 영향 반전

기본적으로 값이 높을수록 점수가 높아져요. 값이 낮을수록 점수가 높아지길 원한다면(예: 더 낮은 가격이 더 관련성이 높은 경우) 인덱스 생성 시 positive_score_impact를 false로 설정하세요:

PUT /products_new
{
  "mappings": {
    "properties": {
      "popularity": {
        "type": "rank_feature",
        "positive_score_impact": false
      }
    }
  }
}

ExampleCreate an index with a rank feature field

더 알아보기 (Learn more)