Boost 매핑 파라미터

Boost 매핑 파라미터

boost 매핑 파라미터는 검색 쿼리 중 특정 필드의 관련성 점수(relevance score)를 높이거나 낮추는 데 사용돼요. 문서의 전체 관련성 점수를 계산할 때 특정 필드에 더 많거나 적은 가중치를 적용해요.

출처: 문서

본문

boost 매핑 파라미터는 검색 쿼리 중 특정 필드의 관련성 점수를 높이거나 낮추는 데 사용합니다. 문서의 전체 관련성 점수를 계산할 때 특정 필드에 더 많거나 적은 가중치를 적용할 수 있습니다.

boost 파라미터는 필드의 점수에 곱해지는 승수(multiplier)로 적용됩니다. 예를 들어 필드의 boost 값이 2이면 해당 필드의 점수 기여도는 두 배가 됩니다. 반대로 boost 값이 0.5이면 해당 필드의 점수 기여도는 절반이 됩니다.

boost 파라미터를 사용할 때는 작은 값(1.5 또는 2)으로 시작해 검색 결과에 미치는 영향을 테스트할 것을 권장합니다. 지나치게 높은 boost 값은 관련성 점수를 왜곡해 예상치 못하거나 바람직하지 않은 검색 결과를 초래할 수 있습니다.

boost 파라미터는 용어 수준(term-level) 쿼리에만 적용됩니다. 접두어(prefix), 범위(range), 퍼지(fuzzy) 용어 수준 쿼리에는 적용되지 않습니다.

인덱스 시점 부스팅과 쿼리 시점 부스팅(Index-time and query-time boosting)

필드 매핑에서 boost 값을 설정할 수 있지만(인덱스 시점 부스팅), 이는 권장되지 않습니다. 대신 여러 장점이 있는 쿼리 시점 부스팅(query-time boosting)을 사용하세요.

  • 유연성(Flexibility): 쿼리 시점 부스팅을 사용하면 문서를 재인덱싱하지 않고도 boost 값을 조정할 수 있습니다.
  • 정밀성(Precision): 인덱스 시점 부스팅은 1바이트만 사용하는 norms의 일부로 저장됩니다. 이는 필드 길이 정규화의 해상도를 낮출 수 있습니다.
  • 동적 제어(Dynamic control): 쿼리 시점 부스팅은 다양한 사용 사례에 대해 서로 다른 boost 값을 실험할 수 있는 능력을 제공합니다.

예제

boost 파라미터를 사용해 특정 필드에 더 많은 가중치를 부여할 수 있습니다. 예를 들어 title이 관련성의 더 강한 지표라면 description 필드보다 title 필드를 더 강하게 부스팅하면 결과가 개선될 수 있습니다.

이 예제에서 title 필드는 boost 2를 가지므로 description 필드(기본 boost 1)보다 관련성 점수에 두 배로 기여합니다.

인덱스 시점 부스팅(권장되지 않음)

부스팅된 필드로 인덱스를 만듭니다(데모 목적일 뿐입니다).

PUT /article_index
{
  "mappings": {
    "properties": {
      "title": {
        "type": "text",
        "boost": 2
      },
      "description": {
        "type": "text"
      }
    }
  }
}

인덱스에 샘플 문서를 추가합니다.

PUT /article_index/_doc/1
{
  "title": "Introduction to Machine Learning",
  "description": "This article covers basic algorithms and their applications in data science."
}
PUT /article_index/_doc/2
{
  "title": "Data Science Fundamentals",
  "description": "Learn about machine learning algorithms and statistical methods for analyzing data."
}

인덱스 시점 부스팅을 사용해 두 필드를 모두 검색합니다.

POST /article_index/_search
{
  "query": {
    "multi_match": {
      "query": "machine learning algorithms",
      "fields": ["title", "description"]
    }
  }
}

"machine learning"이 부스팅된 title 필드에 나타나므로 문서 1이 더 높은 점수를 받습니다. "machine learning"이 부스팅되지 않은 description 필드에 나타나므로 문서 2는 더 낮은 점수를 받습니다. 두 문서 모두 "algorithms"를 포함하므로 점수에 기여합니다.

{
  "took": 415,
  "timed_out": false,
  "_shards": {
    "total": 1,
    "successful": 1,
    "skipped": 0,
    "failed": 0
  },
  "hits": {
    "total": {
      "value": 2,
      "relation": "eq"
    },
    "max_score": 1.1906823,
    "hits": [
      {
        "_index": "article_index",
        "_id": "1",
        "_score": 1.1906823,
        "_source": {
          "title": "Introduction to Machine Learning",
          "description": "This article covers basic algorithms and their applications in data science."
        }
      },
      {
        "_index": "article_index",
        "_id": "2",
        "_score": 0.7130072,
        "_source": {
          "title": "Data Science Fundamentals",
          "description": "Learn about machine learning algorithms and statistical methods for analyzing data."
        }
      }
    ]
  }
}

쿼리 시점 부스팅(권장됨)

인덱스 시점 부스팅 대신 더 나은 제어와 유연성을 위해 쿼리 시점 부스팅을 사용하세요. 쿼리 시점 부스팅은 특별한 필드 매핑을 구성할 필요가 없습니다.

인덱스에 샘플 문서를 추가합니다.

PUT /article_index_2/_doc/1
{
  "title": "Introduction to Machine Learning",
  "description": "This article covers basic algorithms and their applications in data science."
}
PUT /article_index_2/_doc/2
{
  "title": "Data Science Fundamentals",
  "description": "Learn about machine learning algorithms and statistical methods for analyzing data."
}

먼저 부스팅 없이 title 필드를 검색합니다.

POST /article_index_2/_search
{
  "query": {
    "match": {
      "title": {
        "query": "machine learning algorithms"
      }
    }
  }
}

매칭되는 문서의 점수는 0.59입니다.

{
  "took": 13,
  "timed_out": false,
  "_shards": {
    "total": 1,
    "successful": 1,
    "skipped": 0,
    "failed": 0
  },
  "hits": {
    "total": {
      "value": 1,
      "relation": "eq"
    },
    "max_score": 0.59534115,
    "hits": [
      {
        "_index": "article_index_2",
        "_id": "1",
        "_score": 0.59534115,
        "_source": {
          "title": "Introduction to Machine Learning",
          "description": "This article covers basic algorithms and their applications in data science."
        }
      }
    ]
  }
}

다음으로 같은 필드를 부스팅하여 검색합니다.

POST /article_index_2/_search
{
  "query": {
    "match": {
      "title": {
        "query": "machine learning algorithms",
        "boost": 2
      }
    }
  }
}

문서 점수가 두 배가 됩니다.

{
  "took": 16,
  "timed_out": false,
  "_shards": {
    "total": 1,
    "successful": 1,
    "skipped": 0,
    "failed": 0
  },
  "hits": {
    "total": {
      "value": 1,
      "relation": "eq"
    },
    "max_score": 1.1906823,
    "hits": [
      {
        "_index": "article_index_2",
        "_id": "1",
        "_score": 1.1906823,
        "_source": {
          "title": "Introduction to Machine Learning",
          "description": "This article covers basic algorithms and their applications in data science."
        }
      }
    ]
  }
}

title과 description 필드를 모두 검색하려면 먼저 부스팅 없이 검색해 기준선(baseline)을 얻습니다.

POST /article_index_2/_search
{
  "query": {
    "multi_match": {
      "query": "machine learning algorithms",
      "fields": ["title", "description"]
    }
  }
}

문서 2가 문서 1보다 더 높은 점수를 받습니다.

{
  "took": 10,
  "timed_out": false,
  "_shards": {
    "total": 1,
    "successful": 1,
    "skipped": 0,
    "failed": 0
  },
  "hits": {
    "total": {
      "value": 2,
      "relation": "eq"
    },
    "max_score": 0.7130072,
    "hits": [
      {
        "_index": "article_index_2",
        "_id": "2",
        "_score": 0.7130072,
        "_source": {
          "title": "Data Science Fundamentals",
          "description": "Learn about machine learning algorithms and statistical methods for analyzing data."
        }
      },
      {
        "_index": "article_index_2",
        "_id": "1",
        "_score": 0.59534115,
        "_source": {
          "title": "Introduction to Machine Learning",
          "description": "This article covers basic algorithms and their applications in data science."
        }
      }
    ]
  }
}

인덱스 시점 부스팅과 쿼리 시점 부스팅을 비교하려면 쿼리 시점 부스팅으로 여러 필드를 검색합니다.

POST /article_index_2/_search
{
  "query": {
    "multi_match": {
      "query": "machine learning algorithms",
      "fields": ["title^2", "description"]
    }
  }
}

이 쿼리는 인덱스 시점 부스팅 쿼리와 같은 응답을 생성하며, 이제 문서 1이 문서 2보다 더 높은 점수를 받습니다.

{
  "took": 5,
  "timed_out": false,
  "_shards": {
    "total": 1,
    "successful": 1,
    "skipped": 0,
    "failed": 0
  },
  "hits": {
    "total": {
      "value": 2,
      "relation": "eq"
    },
    "max_score": 1.1906823,
    "hits": [
      {
        "_index": "article_index_2",
        "_id": "1",
        "_score": 1.1906823,
        "_source": {
          "title": "Introduction to Machine Learning",
          "description": "This article covers basic algorithms and their applications in data science."
        }
      },
      {
        "_index": "article_index_2",
        "_id": "2",
        "_score": 0.7130072,
        "_source": {
          "title": "Data Science Fundamentals",
          "description": "Learn about machine learning algorithms and statistical methods for analyzing data."
        }
      }
    ]
  }
}

더 알아보기 (Learn more)