Rank 필드 타입
Rank 필드 타입
OpenSearch가 지원하는 모든 rank 필드 타입은 다음 표와 같아요.
출처: 문서
본문
| 필드 데이터 타입 | 설명 |
|---|---|
rank_feature |
문서의 관련성 점수를 높이거나 낮춰요. |
rank_features |
문서의 관련성 점수를 높이거나 낮춰요. 특징 목록이 희소(sparse)할 때 사용해요. |
rank feature과 rank features 필드는 rank feature 쿼리로만 쿼리할 수 있어요. 집계나 정렬은 지원하지 않아요.
Rank feature
rank feature 필드 타입은 양의 float 값을 사용해 rank_feature 쿼리에서 문서의 관련성 점수를 높이거나 낮춰요. 기본적으로 이 값은 관련성 점수를 높여요. 관련성 점수를 낮추려면 선택 사항인 positive_score_impact 파라미터를 false로 설정해요.
예시
rank feature 필드가 있는 매핑을 만들어 볼게요.
PUT chessplayers
{
"mappings": {
"properties": {
"name" : {
"type" : "text"
},
"rating": {
"type": "rank_feature"
},
"age": {
"type": "rank_feature",
"positive_score_impact": false
}
}
}
}
점수를 높이는 rank_feature 필드(rating)와 점수를 낮추는 rank_feature 필드(age)가 있는 문서 세 개를 색인해요.
PUT testindex1/_doc/1
{
"name" : "John Doe",
"rating" : 2554,
"age" : 75
}
PUT testindex1/_doc/2
{
"name" : "Kwaku Mensah",
"rating" : 2067,
"age": 10
}
PUT testindex1/_doc/3
{
"name" : "Nikki Wolf",
"rating" : 1864,
"age" : 22
}
Rank feature 쿼리
rank feature 쿼리를 사용해 플레이어를 rating, age, 또는 rating과 age 둘 다로 랭킹할 수 있어요. rating으로 랭킹하면 rating이 높은 플레이어일수록 관련성 점수가 높아요. age로 랭킹하면 어린 플레이어일수록 관련성 점수가 높아요.
rank feature 쿼리를 사용해 age와 rating을 기준으로 플레이어를 검색해 볼게요.
GET chessplayers/_search
{
"query": {
"bool": {
"should": [
{
"rank_feature": {
"field": "rating"
}
},
{
"rank_feature": {
"field": "age"
}
}
]
}
}
}
age와 rating 둘 다로 랭킹하면 어린 플레이어와 랭킹이 높은 플레이어가 더 좋은 점수를 받아요.
{
"took" : 2,
"timed_out" : false,
"_shards" : {
"total" : 1,
"successful" : 1,
"skipped" : 0,
"failed" : 0
},
"hits" : {
"total" : {
"value" : 3,
"relation" : "eq"
},
"max_score" : 1.2093145,
"hits" : [
{
"_index" : "chessplayers",
"_type" : "_doc",
"_id" : "2",
"_score" : 1.2093145,
"_source" : {
"name" : "Kwaku Mensah",
"rating" : 1967,
"age" : 10
}
},
{
"_index" : "chessplayers",
"_type" : "_doc",
"_id" : "3",
"_score" : 1.0150313,
"_source" : {
"name" : "Nikki Wolf",
"rating" : 1864,
"age" : 22
}
},
{
"_index" : "chessplayers",
"_type" : "_doc",
"_id" : "1",
"_score" : 0.8098284,
"_source" : {
"name" : "John Doe",
"rating" : 2554,
"age" : 75
}
}
]
}
}
Rank features
rank features 필드 타입은 rank feature 필드 타입과 비슷하지만, 희소한 특징 목록에 더 적합해요. rank features 필드는 나중에 rank_feature 쿼리에서 문서의 관련성 점수를 높이거나 낮추는 데 사용되는 숫자 특징 벡터를 색인할 수 있어요.
예시
rank features 필드가 있는 매핑을 만들어 볼게요.
PUT testindex1
{
"mappings": {
"properties": {
"correlations": {
"type": "rank_features"
}
}
}
}
rank features 필드가 있는 문서를 색인하려면 문자열 키와 양의 float 값을 가진 해시맵을 사용해요.
PUT testindex1/_doc/1
{
"correlations": {
"young kids" : 1,
"older kids" : 15,
"teens" : 25.9
}
}
PUT testindex1/_doc/2
{
"correlations": {
"teens": 10,
"adults": 95.7
}
}
rank feature 쿼리를 사용해 문서를 쿼리해요.
GET testindex1/_search
{
"query": {
"rank_feature": {
"field": "correlations.teens"
}
}
}
응답은 관련성 점수로 랭킹돼요.
{
"took" : 123,
"timed_out" : false,
"_shards" : {
"total" : 1,
"successful" : 1,
"skipped" : 0,
"failed" : 0
},
"hits" : {
"total" : {
"value" : 2,
"relation" : "eq"
},
"max_score" : 0.6258503,
"hits" : [
{
"_index" : "testindex1",
"_type" : "_doc",
"_id" : "1",
"_score" : 0.6258503,
"_source" : {
"correlations" : {
"young kids" : 1,
"older kids" : 15,
"teens" : 25.9
}
}
},
{
"_index" : "testindex1",
"_type" : "_doc",
"_id" : "2",
"_score" : 0.39263803,
"_source" : {
"correlations" : {
"teens" : 10,
"adults" : 95.7
}
}
}
]
}
}
rank feature과 rank features 필드는 정밀도를 위해 상위 9개의 유효 비트를 사용하므로 약 0.4%의 상대 오차가 발생해요. 값은 2⁻⁸ = 0.00390625의 상대 정밀도로 저장돼요.