Rank 필드 타입

Rank 필드 타입

OpenSearch가 지원하는 모든 rank 필드 타입은 다음 표와 같아요.

출처: 문서

본문

필드 데이터 타입 설명
rank_feature 문서의 관련성 점수를 높이거나 낮춰요.
rank_features 문서의 관련성 점수를 높이거나 낮춰요. 특징 목록이 희소(sparse)할 때 사용해요.

rank feature과 rank features 필드는 rank feature 쿼리로만 쿼리할 수 있어요. 집계나 정렬은 지원하지 않아요.

Rank feature

rank feature 필드 타입은 양의 float 값을 사용해 rank_feature 쿼리에서 문서의 관련성 점수를 높이거나 낮춰요. 기본적으로 이 값은 관련성 점수를 높여요. 관련성 점수를 낮추려면 선택 사항인 positive_score_impact 파라미터를 false로 설정해요.

예시

rank feature 필드가 있는 매핑을 만들어 볼게요.

PUT chessplayers
{
  "mappings": {
    "properties": {
      "name" : {
        "type" : "text"
      },
      "rating": {
        "type": "rank_feature" 
      },
      "age": {
        "type": "rank_feature",
        "positive_score_impact": false 
      }
    }
  }
}

점수를 높이는 rank_feature 필드(rating)와 점수를 낮추는 rank_feature 필드(age)가 있는 문서 세 개를 색인해요.

PUT testindex1/_doc/1
{
  "name" : "John Doe",
  "rating" : 2554,
  "age" : 75
}
PUT testindex1/_doc/2
{
  "name" : "Kwaku Mensah",
  "rating" : 2067,
  "age": 10
}
PUT testindex1/_doc/3
{
  "name" : "Nikki Wolf",
  "rating" : 1864,
  "age" : 22
}

Rank feature 쿼리

rank feature 쿼리를 사용해 플레이어를 rating, age, 또는 rating과 age 둘 다로 랭킹할 수 있어요. rating으로 랭킹하면 rating이 높은 플레이어일수록 관련성 점수가 높아요. age로 랭킹하면 어린 플레이어일수록 관련성 점수가 높아요.

rank feature 쿼리를 사용해 age와 rating을 기준으로 플레이어를 검색해 볼게요.

GET chessplayers/_search
{
  "query": {
    "bool": {
      "should": [
        {
          "rank_feature": {
            "field": "rating"
          }
        },
        {
          "rank_feature": {
            "field": "age"
          }
        }
      ]
    }
  }
}

age와 rating 둘 다로 랭킹하면 어린 플레이어와 랭킹이 높은 플레이어가 더 좋은 점수를 받아요.

{
  "took" : 2,
  "timed_out" : false,
  "_shards" : {
    "total" : 1,
    "successful" : 1,
    "skipped" : 0,
    "failed" : 0
  },
  "hits" : {
    "total" : {
      "value" : 3,
      "relation" : "eq"
    },
    "max_score" : 1.2093145,
    "hits" : [
      {
        "_index" : "chessplayers",
        "_type" : "_doc",
        "_id" : "2",
        "_score" : 1.2093145,
        "_source" : {
          "name" : "Kwaku Mensah",
          "rating" : 1967,
          "age" : 10
        }
      },
      {
        "_index" : "chessplayers",
        "_type" : "_doc",
        "_id" : "3",
        "_score" : 1.0150313,
        "_source" : {
          "name" : "Nikki Wolf",
          "rating" : 1864,
          "age" : 22
        }
      },
      {
        "_index" : "chessplayers",
        "_type" : "_doc",
        "_id" : "1",
        "_score" : 0.8098284,
        "_source" : {
          "name" : "John Doe",
          "rating" : 2554,
          "age" : 75
        }
      }
    ]
  }
}

Rank features

rank features 필드 타입은 rank feature 필드 타입과 비슷하지만, 희소한 특징 목록에 더 적합해요. rank features 필드는 나중에 rank_feature 쿼리에서 문서의 관련성 점수를 높이거나 낮추는 데 사용되는 숫자 특징 벡터를 색인할 수 있어요.

예시

rank features 필드가 있는 매핑을 만들어 볼게요.

PUT testindex1
{
  "mappings": {
    "properties": {
      "correlations": {
        "type": "rank_features" 
      }
    }
  }
}

rank features 필드가 있는 문서를 색인하려면 문자열 키와 양의 float 값을 가진 해시맵을 사용해요.

PUT testindex1/_doc/1
{
  "correlations": { 
    "young kids" : 1,
    "older kids" : 15,
    "teens" : 25.9
  }
}
PUT testindex1/_doc/2
{
  "correlations": {
    "teens": 10,
    "adults": 95.7
  }
}

rank feature 쿼리를 사용해 문서를 쿼리해요.

GET testindex1/_search
{
  "query": {
    "rank_feature": {
      "field": "correlations.teens"
    }
  }
}

응답은 관련성 점수로 랭킹돼요.

{
  "took" : 123,
  "timed_out" : false,
  "_shards" : {
    "total" : 1,
    "successful" : 1,
    "skipped" : 0,
    "failed" : 0
  },
  "hits" : {
    "total" : {
      "value" : 2,
      "relation" : "eq"
    },
    "max_score" : 0.6258503,
    "hits" : [
      {
        "_index" : "testindex1",
        "_type" : "_doc",
        "_id" : "1",
        "_score" : 0.6258503,
        "_source" : {
          "correlations" : {
            "young kids" : 1,
            "older kids" : 15,
            "teens" : 25.9
          }
        }
      },
      {
        "_index" : "testindex1",
        "_type" : "_doc",
        "_id" : "2",
        "_score" : 0.39263803,
        "_source" : {
          "correlations" : {
            "teens" : 10,
            "adults" : 95.7
          }
        }
      }
    ]
  }
}

rank feature과 rank features 필드는 정밀도를 위해 상위 9개의 유효 비트를 사용하므로 약 0.4%의 상대 오차가 발생해요. 값은 2⁻⁸ = 0.00390625의 상대 정밀도로 저장돼요.

더 알아보기 (Learn more)