Multi-match 쿼리
Multi-match 쿼리
multi_match 연산은 match 연산과 비슷하게 동작해요. multi_match 쿼리를 사용해 여러 필드를 검색할 수 있습니다. ^는 특정 필드를 "부스트"하며, 한 필드의 일치를 다른 필드의 일치보다 더 무겁게 가중하는 배수(multiplier) 역할을 해요.
출처: 문서
본문
multi-match 연산은 match 연산과 비슷하게 동작해요. multi_match 쿼리를 사용해 여러 필드를 검색할 수 있어요.
^는 특정 필드를 "부스트"해요. 부스트는 한 필드의 일치를 다른 필드의 일치보다 더 무겁게 가중하는 배수예요. 다음 예시에서 title 필드의 "wind" 일치는 plot 필드의 일치보다 _score에 4배 더 많은 영향을 줘요.
GET _search
{
"query": {
"multi_match": {
"query": "wind",
"fields": ["title^4", "plot"]
}
}
}
그 결과 The Wind Rises나 Gone with the Wind 같은 영화는 검색 결과 상위에, plot 요약에 "wind"가 있을 것으로 추정되는 Twister 같은 영화는 하위에 옵니다.
필드 이름에 와일드카드를 사용할 수 있어요. 예를 들어 다음 쿼리는 speaker 필드와 play_로 시작하는 모든 필드(예: play_name 또는 play_title)를 검색해요.
GET _search
{
"query": {
"multi_match": {
"query": "hamlet",
"fields": ["speaker", "play_*"]
}
}
}
fields 파라미터를 제공하지 않으면 multi_match 쿼리는 기본적으로 *인 index.query.default_field 설정에 지정된 필드를 검색해요. 기본 동작은 term-level 쿼리 대상이 되는 매핑의 모든 필드를 추출하고, 메타데이터 필드를 필터링하며, 추출된 모든 필드를 결합해 쿼리를 만드는 거예요.
쿼리의 최대 절 수는 기본값 1,024인 indices.query.bool.max_clause_count 설정으로 정의돼요.
Multi-match 쿼리 타입
OpenSearch는 내부적으로 실행 방식이 다른 다음 multi-match 쿼리 타입을 지원해요.
best_fields(기본값): 어떤 필드와도 일치하는 문서를 반환함. 가장 잘 일치하는 필드의_score를 사용함.most_fields: 어떤 필드와도 일치하는 문서를 반환함. 각 일치 필드의 결합 점수를 사용함.cross_fields: 모든 필드를 하나의 필드처럼 취급함. 같은 분석기를 가진 필드를 처리하고 어떤 필드에서든 단어를 매칭함.phrase: 각 필드에서match_phrase쿼리를 실행함. 가장 잘 일치하는 필드의_score를 사용함.phrase_prefix: 각 필드에서match_phrase_prefix쿼리를 실행함. 가장 잘 일치하는 필드의_score를 사용함.bool_prefix: 각 필드에서match_bool_prefix쿼리를 실행함. 각 일치 필드의 결합 점수를 사용함.
Best fields
개념을 지정하는 두 단어를 검색한다면, 두 단어가 서로 붙어 있는 결과가 더 높게 점수가 매겨지기를 원할 거예요.
예를 들어 다음 과학 논문을 포함하는 인덱스를 생각해 볼게요.
PUT /articles/_doc/1
{
"title": "Aurora borealis",
"description": "Northern lights, or aurora borealis, explained"
}
PUT /articles/_doc/2
{
"title": "Sun deprivation in the Northern countries",
"description": "Using fluorescent lights for therapy"
}
title 또는 description에서 northern lights를 포함한 논문을 검색할 수 있어요.
GET articles/_search
{
"query": {
"multi_match" : {
"query": "northern lights",
"type": "best_fields",
"fields": [ "title", "description" ],
"tie_breaker": 0.3
}
}
}
앞의 쿼리는 각 필드에 대한 match 쿼리가 있는 다음 dis_max 쿼리로 실행돼요.
GET /articles/_search
{
"query": {
"dis_max": {
"queries": [
{ "match": { "title": "northern lights" }},
{ "match": { "description": "northern lights" }}
],
"tie_breaker": 0.3
}
}
}
결과에는 두 문서가 모두 포함되지만, description 필드에 두 단어가 모두 있으므로 문서 1이 더 높게 점수가 매겨져요.
{
"took": 30,
"timed_out": false,
"_shards": {
"total": 1,
"successful": 1,
"skipped": 0,
"failed": 0
},
"hits": {
"total": {
"value": 2,
"relation": "eq"
},
"max_score": 0.84407747,
"hits": [
{
"_index": "articles",
"_id": "1",
"_score": 0.84407747,
"_source": {
"title": "Aurora borealis",
"description": "Northern lights, or aurora borealis, explained"
}
},
{
"_index": "articles",
"_id": "2",
"_score": 0.6322521,
"_source": {
"title": "Sun deprivation in the Northern countries",
"description": "Using fluorescent lights for therapy"
}
}
]
}
}
best_fields 쿼리는 가장 잘 일치하는 필드의 점수를 사용해요. tie_breaker를 지정하면 점수는 다음 알고리즘으로 계산돼요. 가장 잘 일치하는 필드의 점수를 가져와 다른 모든 일치 필드에 대해 (tie_breaker * _score)를 더해요.
Most fields
서로 다른 방식으로 분석된 동일한 텍스트를 포함하는 여러 필드에는 most_fields 쿼리를 사용하세요. 예를 들어 원래 필드는 표준 분석기로 분석된 텍스트를 포함하고, 다른 필드는 형태소 분석을 수행하는 english 분석기로 분석된 동일한 텍스트를 포함할 수 있어요.
PUT /articles
{
"mappings": {
"properties": {
"title": {
"type": "text",
"fields": {
"english": {
"type": "text",
"analyzer": "english"
}
}
}
}
}
}
articles 인덱스에 인덱싱된 다음 두 문서를 생각해 볼게요.
PUT /articles/_doc/1
{
"title": "Buttered toasts"
}
PUT /articles/_doc/2
{
"title": "Buttering a toast"
}
표준 분석기는 title Buttered toast를 [buttered, toasts]로, title Buttering a toast를 [buttering, a, toast]로 분석해요. 반면 english 분석기는 형태소 분석 때문에 두 title 모두에 대해 동일한 토큰 목록 [butter, toast]를 만들어요.
가능한 한 많은 문서를 반환하려면 most_fields 쿼리를 사용할 수 있어요.
GET /articles/_search
{
"query": {
"multi_match": {
"query": "buttered toast",
"fields": [
"title",
"title.english"
],
"type": "most_fields"
}
}
}
앞의 쿼리는 다음 Boolean 쿼리로 실행돼요.
GET articles/_search
{
"query": {
"bool": {
"should": [
{ "match": { "title": "buttered toasts" }},
{ "match": { "title.english": "buttered toasts" }}
]
}
}
}
관련성 점수를 계산하기 위해 문서의 모든 match 절에 대한 점수를 더한 다음 그 결과를 match 절의 수로 나눠요.
title.english 필드를 포함하면 형태소 분석된 토큰과 일치하는 두 번째 문서를 검색할 수 있어요.
{
"took": 9,
"timed_out": false,
"_shards": {
"total": 1,
"successful": 1,
"skipped": 0,
"failed": 0
},
"hits": {
"total": {
"value": 2,
"relation": "eq"
},
"max_score": 1.4418206,
"hits": [
{
"_index": "articles",
"_id": "1",
"_score": 1.4418206,
"_source": {
"title": "Buttered toasts"
}
},
{
"_index": "articles",
"_id": "2",
"_score": 0.09304003,
"_source": {
"title": "Buttering a toast"
}
}
]
}
}
첫 번째 문서는 title과 title.english 필드가 모두 일치하므로 더 높은 관련성 점수를 가져요.
Operator와 minimum should match
best_fields와 most_fields 쿼리는 필드 기준으로(필드당 하나씩) match 쿼리를 생성해요. 따라서 minimum_should_match와 operator 파라미터는 각 필드에 적용되며, 이는 일반적으로 원하는 동작이 아니에요.
예를 들어 다음 문서가 있는 customers 인덱스를 생각해 볼게요.
PUT customers/_doc/1
{
"first_name": "John",
"last_name": "Doe"
}
PUT customers/_doc/2
{
"first_name": "Jane",
"last_name": "Doe"
}
customers 인덱스에서 John Doe를 검색한다면 다음 쿼리를 구성할 수 있어요.
GET customers/_validate/query?explain
{
"query": {
"multi_match" : {
"query": "John Doe",
"type": "best_fields",
"fields": [ "first_name", "last_name" ],
"operator": "and"
}
}
}
이 쿼리에서 and 연산자의 의도는 John과 Doe를 모두 일치하는 문서를 찾는 거예요. 하지만 쿼리는 어떤 결과도 반환하지 않아요. Validate API를 실행해 쿼리가 어떻게 실행되는지 알 수 있어요.
GET customers/_validate/query?explain
{
"query": {
"multi_match" : {
"query": "John Doe",
"type": "best_fields",
"fields": [ "first_name", "last_name" ],
"operator": "and"
}
}
}
응답에서 쿼리가 John과 Doe 둘 다를 first_name 또는 last_name 필드에 일치시키려고 한다는 것을 볼 수 있어요.
{
"_shards": {
"total": 1,
"successful": 1,
"failed": 0
},
"valid": true,
"explanations": [
{
"index": "customers",
"valid": true,
"explanation": "((+first_name:john +first_name:doe) | (+last_name:john +last_name:doe))"
}
]
}
어떤 필드도 두 단어를 모두 포함하지 않으므로 결과가 반환되지 않아요.
필드 걸쳐 검색하는 더 나은 대안은 cross_fields 쿼리를 사용하는 거예요. 필드 중심인 best_fields와 most_fields 쿼리와 달리 cross_fields 쿼리는 용어 중심이에요.
Cross fields
cross_fields 쿼리를 사용해 여러 필드에 걸쳐 데이터를 검색하세요. 예를 들어 인덱스가 고객 데이터를 포함하면 고객의 이름과 성은 다른 필드에 있어요. 하지만 John Doe를 검색할 때는 John이 first_name 필드에, Doe가 last_name 필드에 있는 문서를 받고 싶을 거예요.
most_fields 쿼리는 다음 문제 때문에 이 경우에 작동하지 않아요.
operator와minimum_should_match파라미터가 용어 기준이 아니라 필드 기준으로 적용돼요.first_name과last_name필드의 용어 빈도가 예상치 못한 결과를 낳을 수 있어요. 예를 들어 누군가의 이름이 우연히Doe라면, 이 이름이 다른 문서에는 나타나지 않기 때문에 이 이름을 가진 문서가 더 나은 일치로 간주돼요.
cross_fields 쿼리는 쿼리 문자열을 개별 용어로 분석한 다음 필드들 중 어떤 필드에서든 각 용어를 마치 하나의 필드인 것처럼 검색해요.
다음은 John Doe에 대한 cross_fields 쿼리예요.
GET /customers/_search
{
"query": {
"multi_match" : {
"query": "John Doe",
"type": "cross_fields",
"fields": [ "first_name", "last_name" ],
"operator": "and"
}
}
}
응답에는 John과 Doe가 모두 존재하는 유일한 문서가 포함돼요.
{
"took": 19,
"timed_out": false,
"_shards": {
"total": 1,
"successful": 1,
"skipped": 0,
"failed": 0
},
"hits": {
"total": {
"value": 1,
"relation": "eq"
},
"max_score": 0.8754687,
"hits": [
{
"_index": "customers",
"_id": "1",
"_score": 0.8754687,
"_source": {
"first_name": "John",
"last_name": "Doe"
}
}
]
}
}
validate API 연산을 사용해 앞의 쿼리가 어떻게 실행되는지 통찰을 얻을 수 있어요.
GET /customers/_validate/query?explain
{
"query": {
"multi_match" : {
"query": "John Doe",
"type": "cross_fields",
"fields": [ "first_name", "last_name" ],
"operator": "and"
}
}
}
응답에서 쿼리가 적어도 하나의 필드에서 모든 용어를 검색하고 있음을 볼 수 있어요.
{
"_shards": {
"total": 1,
"successful": 1,
"failed": 0
},
"valid": true,
"explanations": [
{
"index": "customers",
"valid": true,
"explanation": "+blended(terms:[last_name:john, first_name:john]) +blended(terms:[last_name:doe, first_name:doe])"
}
]
}
따라서 모든 필드의 용어 빈도를 혼합(blend)하면 차이를 보정해 서로 다른 용어 빈도의 문제를 해결할 수 있어요.
cross_fields 쿼리는 보통 부스트 1의 짧은 문자열 필드에서만 유용해요. 그 외의 경우에는 부스트, 용어 빈도, 길이 정규화가 점수에 기여하는 방식 때문에 점수가 의미 있는 용어 통계 혼합을 만들어내지 못해요.
cross_fields 쿼리에는 fuzziness 파라미터가 지원되지 않아요.
분석 (Analysis)
cross_fields 쿼리는 같은 분석기를 가진 필드에서만 용어 중심 쿼리로 작동해요. 같은 분석기를 가진 필드들은 함께 그룹화되고 이 그룹들은 Boolean 쿼리로 결합돼요.
예를 들어 first_name과 last_name 필드가 기본 표준 분석기로 분석되고, 그 .edge 하위 필드들은 edge n-gram 분석기로 분석되는 인덱스를 생각해 볼게요.
응답
PUT customers
{
"settings": {
"analysis": {
"analyzer": {
"my_analyzer": {
"tokenizer": "my_tokenizer"
}
},
"tokenizer": {
"my_tokenizer": {
"type": "edge_ngram",
"min_gram": 2,
"max_gram": 10
}
}
}
},
"mappings": {
"properties": {
"first_name": {
"type": "text",
"fields": {
"edge": {
"type": "text",
"analyzer": "my_analyzer"
}
}
},
"last_name": {
"type": "text",
"fields": {
"edge": {
"type": "text",
"analyzer": "my_analyzer"
}
}
}
}
}
}
customers 인덱스에 문서 하나를 인덱싱해요.
PUT /customers/_doc/1
{
"first": "John",
"last": "Doe"
}
cross_fields 쿼리를 사용해 John Doe를 필드들에 걸쳐 검색할 수 있어요.
GET /customers/_search
{
"query": {
"multi_match" : {
"query": "John",
"type": "cross_fields",
"fields": [
"first_name", "first_name.edge",
"last_name", "last_name.edge"
]
}
}
}
쿼리가 어떻게 실행되는지 보려면 Validate API를 실행할 수 있어요.
GET /customers/_validate/query?explain
{
"query": {
"multi_match" : {
"query": "John",
"type": "cross_fields",
"fields": [
"first_name", "first_name.edge",
"last_name", "last_name.edge"
]
}
}
}
응답은 last_name과 first_name 필드가 함께 그룹화되고 단일 필드로 취급되는 것을 보여줘요. 마찬가지로 last_name.edge와 first_name.edge 필드가 함께 그룹화되어 단일 필드로 취급돼요.
{
"_shards": {
"total": 1,
"successful": 1,
"failed": 0
},
"valid": true,
"explanations": [
{
"index": "customers",
"valid": true,
"explanation": "(blended(terms:[last_name:john, first_name:john]) | (blended(terms:[last_name.edge:Jo, first_name.edge:Jo]) blended(terms:[last_name.edge:Joh, first_name.edge:Joh]) blended(terms:[last_name.edge:John, first_name.edge:John])))"
}
]
}
앞의 것처럼 여러 필드 그룹과 함께 operator 또는 minimum_should_match 파라미터를 사용하면 이전 섹션에서 설명한 문제가 생길 수 있어요. 이를 피하려면 앞의 쿼리를 Boolean 쿼리로 결합한 두 개의 cross_fields 하위 쿼리로 재작성하고 minimum_should_match를 하위 쿼리 중 하나에 적용할 수 있어요.
GET /customers/_search
{
"query": {
"bool": {
"should": [
{
"multi_match": {
"query": "John Doe",
"type": "cross_fields",
"fields": [
"first_name",
"last_name"
],
"minimum_should_match": "1"
}
},
{
"multi_match": {
"query": "John Doe",
"type": "cross_fields",
"fields": [
"first_name.edge",
"last_name.edge"
]
}
}
]
}
}
}
모든 필드에 대해 하나의 그룹을 만들려면 쿼리에 분석기를 지정하세요.
GET customers/_search
{
"query": {
"multi_match" : {
"query": "John Doe",
"type": "cross_fields",
"analyzer": "standard",
"fields": [ "first_name", "last_name", "*.edge" ]
}
}
}
앞의 쿼리에서 Validate API를 실행하면 쿼리가 어떻게 실행되는지 보여줘요.
{
"_shards": {
"total": 1,
"successful": 1,
"failed": 0
},
"valid": true,
"explanations": [
{
"index": "customers",
"valid": true,
"explanation": "blended(terms:[last_name.edge:john, last_name:john, first_name:john, first_name.edge:john]) blended(terms:[last_name.edge:doe, last_name:doe, first_name:doe, first_name.edge:doe])"
}
]
}
Phrase
phrase 쿼리는 best_fields 쿼리와 비슷하게 동작하지만 match 쿼리 대신 match_phrase 쿼리를 사용해요.
다음은 best_fields 섹션에서 설명한 인덱스에 대한 phrase 쿼리 예시예요.
GET articles/_search
{
"query": {
"multi_match" : {
"query": "northern lights",
"type": "phrase",
"fields": [ "title", "description" ]
}
}
}
앞의 쿼리는 각 필드에 대한 match_phrase 쿼리가 있는 다음 dis_max 쿼리로 실행돼요.
GET articles/_search
{
"query": {
"dis_max": {
"queries": [
{ "match_phrase": { "title": "northern lights" }},
{ "match_phrase": { "description": "northern lights" }}
]
}
}
}
기본적으로 phrase 쿼리는 용어가 같은 순서로 나타날 때만 텍스트를 매칭하므로 문서 1만 결과에 반환돼요.
응답
{
"took": 3,
"timed_out": false,
"_shards": {
"total": 1,
"successful": 1,
"skipped": 0,
"failed": 0
},
"hits": {
"total": {
"value": 1,
"relation": "eq"
},
"max_score": 0.84407747,
"hits": [
{
"_index": "articles",
"_id": "1",
"_score": 0.84407747,
"_source": {
"title": "Aurora borealis",
"description": "Northern lights, or aurora borealis, explained"
}
}
]
}
}
slop 파라미터를 사용해 쿼리 구문의 단어 사이에 다른 단어를 허용할 수 있어요. 예를 들어 다음 쿼리는 flourescent와 therapy 사이에 단어가 최대 두 개 있으면 텍스트를 일치로 받아들여요.
GET articles/_search
{
"query": {
"multi_match" : {
"query": "fluorescent therapy",
"type": "phrase",
"fields": [ "title", "description" ],
"slop": 2
}
}
}
응답에는 문서 2가 포함돼요.
응답
{
"took": 3,
"timed_out": false,
"_shards": {
"total": 1,
"successful": 1,
"skipped": 0,
"failed": 0
},
"hits": {
"total": {
"value": 1,
"relation": "eq"
},
"max_score": 0.7003825,
"hits": [
{
"_index": "articles",
"_id": "2",
"_score": 0.7003825,
"_source": {
"title": "Sun deprivation in the Northern countries",
"description": "Using fluorescent lights for therapy"
}
}
]
}
}
slop 값이 2보다 작으면 어떤 문서도 반환되지 않아요.
phrase 쿼리에는 fuzziness 파라미터가 지원되지 않아요.
Phrase prefix
phrase_prefix 쿼리는 phrase 쿼리와 비슷하게 동작하지만 match_phrase 쿼리 대신 match_phrase_prefix 쿼리를 사용해요.
다음은 best_fields 섹션에서 설명한 인덱스에 대한 phrase_prefix 쿼리 예시예요.
GET articles/_search
{
"query": {
"multi_match" : {
"query": "northern light",
"type": "phrase_prefix",
"fields": [ "title", "description" ]
}
}
}
앞의 쿼리는 각 필드에 대한 match_phrase_prefix 쿼리가 있는 다음 dis_max 쿼리로 실행돼요.
GET articles/_search
{
"query": {
"dis_max": {
"queries": [
{ "match_phrase_prefix": { "title": "northern light" }},
{ "match_phrase_prefix": { "description": "northern light" }}
]
}
}
}
slop 파라미터를 사용해 쿼리 구문의 단어 사이에 다른 단어를 허용할 수 있어요.
phrase_prefix 쿼리에는 fuzziness 파라미터가 지원되지 않아요.
Boolean prefix
bool_prefix 쿼리는 most_fields 쿼리와 비슷하게 문서에 점수를 매기지만 match 쿼리 대신 match_bool_prefix 쿼리를 사용해요.
다음은 best_fields 섹션에서 설명한 인덱스에 대한 bool_prefix 쿼리 예시예요.
GET articles/_search
{
"query": {
"multi_match" : {
"query": "li northern",
"type": "bool_prefix",
"fields": [ "title", "description" ]
}
}
}
앞의 쿼리는 각 필드에 대한 match_bool_prefix 쿼리가 있는 다음 dis_max 쿼리로 실행돼요.
GET articles/_search
{
"query": {
"dis_max": {
"queries": [
{ "match_bool_prefix": { "title": "li northern" }},
{ "match_bool_prefix": { "description": "li northern" }}
]
}
}
}
fuzziness, prefix_length, max_expansions, fuzzy_rewrite, fuzzy_transpositions 파라미터는 term 쿼리를 구성하는 데 사용되는 용어에 대해 지원되지만, 마지막 용어로 구성된 prefix 쿼리에는 효과가 없어요.
파라미터 (Parameters)
이 쿼리는 다음 파라미터를 받아들여요. query를 제외한 모든 파라미터는 선택 사항이에요.
| 파라미터 | 데이터 타입 | 설명 |
|---|---|---|
query |
String | 검색에 사용할 쿼리 문자열. 필수. |
auto_generate_synonyms_phrase_query |
Boolean | 다중 용어 동의어에 대해 match phrase 쿼리를 자동으로 생성할지 여부를 지정함. 예를 들어 ba,batting average를 동의어로 지정하고 ba를 검색하면 OpenSearch는 (이 옵션이 true이면) ba OR "batting average"를 검색하거나 (이 옵션이 false이면) ba OR (batting AND average)를 검색함. 기본값은 true. |
analyzer |
String | 쿼리 문자열 텍스트를 토큰화하는 데 사용되는 분석기. 기본값은 default_field에 대해 지정된 인덱스 시점 분석기. default_field에 분석기가 지정되지 않으면 분석기는 인덱스의 기본 분석기. index.query.default_field에 대한 자세한 내용은 Dynamic index-level index settings를 참조하세요. |
boost |
Floating-point | 주어진 배수로 절을 부스트함. 복합 쿼리에서 절에 가중치를 두는 데 유용함. [0, 1) 범위의 값은 관련성을 낮추고 1보다 큰 값은 관련성을 높임. 기본값은 1. |
fields |
문자열 배열 | 검색할 필드 목록. fields 파라미터를 제공하지 않으면 multi_match 쿼리는 기본적으로 *인 index.query.default_field 설정에 지정된 필드를 검색함. |
fuzziness |
String | 용어가 값과 일치하는지 결정할 때 한 단어를 다른 단어로 바꾸는 데 필요한 문자 편집 수(삽입, 삭제, 대체). 예를 들어 wined와 wind 사이의 거리는 1. 유효한 값은 음이 아닌 정수 또는 AUTO. 기본값 AUTO는 검색 용어의 길이에 따라 편집 거리를 동적으로 선택함. AUTO:[low],[high] 구문을 사용해 임계값을 사용자 지정할 수 있는데, 여기서 low와 high는 문자 길이 경계를 정의함. 생략하면 OpenSearch는 기본값으로 AUTO:3,6을 사용하며, 이는 다음 규칙을 적용함. - 0–2자의 용어: 정확히 일치 필요(편집 0). - 3–5자의 용어: 최대 1개 편집 허용. - 6자 이상의 용어: 최대 2개 편집 허용. 예를 들어 AUTO:4,7은 0–3자의 용어에서 정확히 일치를 요구하고, 4–6자의 용어에서 최대 1개 편집을 허용하며, 7자 이상의 용어에서 최대 2개 편집을 허용함. 대부분의 시나리오에서는 AUTO 사용을 권장함. phrase, phrase_prefix, cross_fields 쿼리에서는 지원되지 않음. |
fuzzy_rewrite |
String | OpenSearch가 쿼리를 어떻게 재작성할지 결정함. 유효한 값은 constant_score, scoring_boolean, constant_score_boolean, top_terms_N, top_terms_boost_N, top_terms_blended_freqs_N. fuzziness 파라미터가 0이 아니면 쿼리는 기본적으로 top_terms_blended_freqs_${max_expansions}의 fuzzy_rewrite 방식을 사용함. 기본값은 constant_score. |
fuzzy_transpositions |
Boolean | fuzzy_transpositions를 true(기본값)로 설정하면 fuzziness 옵션의 삽입, 삭제, 대체 연산에 인접 문자 교체가 추가됨. 예를 들어 fuzzy_transpositions가 true이면 wind와 wnid 사이의 거리는 1("n"과 "i"를 교체)이고, false이면 2("n"을 삭제, "n"을 삽입). fuzzy_transpositions가 false이면 rewind와 wnid는 wind로부터 같은 거리(2)를 가지며, 더 인간 중심적인 의견상 wnid가 명백한 오타임에도 그렇음. 대부분의 사용 사례에서 기본값이 좋은 선택임. |
lenient |
Boolean | lenient를 true로 설정하면 쿼리와 문서 필드 사이의 데이터 타입 불일치를 무시함. 예를 들어 "8.2" 쿼리 문자열은 float 타입 필드와 일치할 수 있음. 기본값은 false. |
max_expansions |
양의 정수 | 쿼리가 확장할 수 있는 최대 용어 수. 퍼지 쿼리는 fuzziness에 지정된 거리 안에 있는 일치 용어 수로 "확장"됨. 그런 다음 OpenSearch는 그 용어들을 매칭하려 함. 기본값은 50. |
minimum_should_match |
양 또는 음의 정수, 양 또는 음의 백분율, 조합 | 쿼리 문자열에 여러 검색 용어가 있고 or 연산자를 사용하는 경우, 문서가 일치로 간주되기 위해 일치해야 하는 용어 수. 예를 들어 minimum_should_match가 2이면 wind often rising은 The Wind Rises와 일치하지 않음. minimum_should_match가 1이면 일치함. 자세한 내용은 Minimum should match를 참조하세요. |
operator |
String | 쿼리 문자열에 여러 검색 용어가 있을 때 문서가 일치로 간주되기 위해 모든 용어가 일치해야 하는지(AND) 아니면 하나의 용어만 일치해도 되는지(OR). 유효한 값은 다음과 같음. - OR: 문자열 to be는 to OR be로 해석됨. - AND: 문자열 to be는 to AND be로 해석됨. 기본값은 OR. |
prefix_length |
음이 아닌 정수 | fuzziness에서 고려되지 않는 선행 문자 수. 기본값은 0. |
slop |
0(기본값) 또는 양의 정수 | 쿼리의 단어들이 순서가 뒤바뀌어도 여전히 일치로 간주될 수 있는 정도를 제어함. Lucene 문서에서: "쿼리 구문의 단어 사이에 허용되는 다른 단어의 수. 예를 들어 두 단어의 순서를 바꾸려면 두 번의 이동이 필요하므로(첫 번째 이동이 단어들을 서로 위에 놓음) 구문의 재정렬을 허용하려면 slop은 최소 2여야 함. 값 0은 정확히 일치를 요구함." phrase와 phrase_prefix 쿼리 타입에 대해 지원됨. |
tie_breaker |
Floating-point | 여러 쿼리 절과 일치하는 문서에 더 많은 가중치를 주는 데 사용하는 0과 1.0 사이의 계수. 자세한 내용은 The tie_breaker parameter를 참조하세요. |
type |
String | multi-match 쿼리 타입. 유효한 값은 best_fields, most_fields, cross_fields, phrase, phrase_prefix, bool_prefix. 기본값은 best_fields. |
zero_terms_query |
String | 어떤 경우에는 분석기가 쿼리 문자열에서 모든 용어를 제거함. 예를 들어 stop 분석기는 an but this 문자열에서 모든 용어를 제거함. 이런 경우 zero_terms_query는 문서를 하나도 매칭하지 않을지(none) 모든 문서를 매칭할지(all)를 지정함. 유효한 값은 none과 all. 기본값은 none. |
fuzziness 파라미터는 phrase, phrase_prefix, cross_fields 쿼리에서 지원되지 않아요.
slop 파라미터는 phrase와 phrase_prefix 쿼리에서만 지원돼요.
tie_breaker 파라미터
각 term-level 혼합(blended) 쿼리는 그룹의 어떤 필드가 반환한 최고 점수로 문서 점수를 계산해요. 모든 혼합 쿼리의 점수를 더해 최종 점수를 만들어요. tie_breaker 파라미터를 사용해 점수 계산 방식을 바꿀 수 있어요. tie_breaker 파라미터는 다음 값을 받아들여요.
0.0(best_fields, cross_fields, phrase, phrase_prefix 쿼리의 기본값): 그룹의 어떤 필드가 반환한 단일 최고 점수를 취함.1.0(most_fields와 bool_prefix 쿼리의 기본값): 그룹의 모든 필드 점수를 더함.- (0, 1) 범위의 부동소수점 값: 가장 잘 일치하는 필드의 단일 최고 점수를 취하고 다른 모든 일치 필드에 대해
(tie_breaker * _score)를 더함.