Linear Decay
Milvus 2.6.x와 호환돼요.
Linear decay는 검색 결과에서 절대적 0점에서 끝나는 직선형 감소를 만들어요. 다가오는 이벤트 카운트다운처럼 이벤트가 지나갈 때까지 관련성이 점차 사라지듯, linear decay는 항목이 이상적인 지점에서 멀어질수록 관련성이 예측 가능하고 꾸준히 감소하다가 완전히 사라지게 해요. 이 접근 방식은 일관된 감소율과 명확한 컷오프가 필요할 때 이상적이에요. 특정 경계 너머의 항목은 결과에서 완전히 제외된다는 점을 보장하거든요.
출처: Milvus 문서
본문
다른 decay 함수와 달리:
- Gaussian decay는 종 모양 곡선을 따라 점진적으로 접근하지만 결코 0에 도달하지 않아요.
- Exponential decay는 무한히 이어지는 최소 관련성의 긴 꼬리를 유지해요.
Linear decay는 유일하게 명확한 종점을 만들기 때문에, 자연스러운 경계나 마감이 있는 애플리케이션에서 특히 효과적이에요.
linear decay는 언제 사용하나요? (When to use linear decay)
Linear decay는 특히 다음 경우에 효과적이에요.
| 사용 사례 | 예시 | Linear가 잘 맞는 이유 |
|---|---|---|
| 이벤트 목록 (Event listings) | 콘서트 티켓 플랫폼 | 너무 먼 미래의 이벤트를 위한 명확한 컷오프를 만들어요 |
| 기간 한정 제공 (Limited-time offers) | 플래시 세일, 프로모션 | 만료됐거나 곧 만료될 제공이 나타나지 않게 해요 |
| 배달 반경 (Delivery radius) | 음식 배달, 택배 서비스 | 엄격한 지리적 경계를 강제해요 |
| 연령 제한 콘텐츠 (Age-restricted content) | 데이팅 플랫폼, 미디어 서비스 | 확고한 연령 임계값을 설정해요 |
다음 경우에 linear decay를 선택하세요.
- 애플리케이션에 자연스러운 경계, 마감, 혹은 임계값이 있을 때
- 특정 지점을 넘어선 항목이 결과에서 완전히 제외되어야 할 때
- 관련성 감소에 예측 가능하고 일관된 속도가 필요할 때
- 사용자가 관련 항목과 비관련 항목 사이의 명확한 구분을 봐야 할 때
꾸준한 감소 원칙 (Steady decline principle)
Linear decay는 정확히 0에 도달할 때까지 일정한 속도로 감소하는 직선 하강을 만들어요. 이 패턴은 카운트다운 타이머, 재고 소진, 마감 임박처럼 관련성이 명확한 만료 지점을 가진 일상적인 시나리오에서 많이 나타나요.
모든 시간 파라미터(origin, offset, scale)는 컬렉션 데이터와 같은 단위를 사용해야 해요. 컬렉션이 서로 다른 단위(밀리초, 마이크로초)로 타임스탬프를 저장한다면 모든 파라미터를 그에 맞게 조정해야 해요.
위 그래프는 티켓 플랫폼의 이벤트 목록에 linear decay가 어떻게 영향을 주는지 보여 줘요.
origin(현재 날짜): 관련성이 최대(1.0)인 현재 시점이에요.offset(1일): "즉시 이벤트 창"이에요. 다음 날 안에 일어나는 모든 이벤트는 전체 관련성 점수(1.0)를 유지해서, 아주 임박한 이벤트가 약간의 시간 차이 때문에 불이익을 받지 않게 해요.decay(0.5): scale 거리에서의 점수예요. 이 파라미터가 관련성 감소 속도를 제어해요.scale(10일): 관련성이 decay 값까지 떨어지는 기간이에요. 10일 뒤의 이벤트는 관련성 점수가 절반(0.5)으로 줄어들어요.
직선 곡선에서 보듯이, 약 16일을 넘어선 이벤트는 정확히 0의 관련성을 가지며 검색 결과에 전혀 나타나지 않아요. 이는 사용자가 정의된 시간 창 안의 관련 이벤트만 보도록 보장하는 명확한 경계를 만들어요.
이 동작은 일반적인 이벤트 기획 방식과 비슷해요. 임박한 이벤트가 가장 관련성이 높고, 다가올 몇 주의 이벤트는 중요도가 줄어들며, 너무 먼 미래의(혹은 이미 지나간) 이벤트는 아예 나타나지 않아야 해요.
공식 (Formula)
Linear decay 점수를 계산하는 수학 공식은 다음과 같아요.
S(doc) = max( (s - max(0, |fieldvalue_doc - origin| - offset)) / s, 0 )
여기서:
s = scale / (1.0 - decay)
이를 쉽게 풀어 설명하면:
- 필드 값이 origin에서 얼마나 떨어져 있는지 계산해요:
|fieldvalue_doc - origin| - offset이 있으면 빼되 0 아래로 내려가지 않게 해요:
max(0, distance - offset) - scale과 decay 값에서 파라미터
s를 구해요. - 조정된 거리를
s에서 빼고s로 나눠요. - 결과가 0 아래로 내려가지 않게 해요:
max(result, 0)
s 계산은 scale과 decay 파라미터를 점수가 0에 도달하는 지점으로 변환해요. 예를 들어 decay=0.5, scale=7이면 점수는 distance=14(scale 값의 두 배)에서 정확히 0에 도달해요.
linear decay 사용하기 (Use linear decay)
Linear decay는 Milvus의 표준 벡터 검색과 하이브리드 검색 연산 모두에 적용할 수 있어요. 아래는 이 기능을 구현하기 위한 핵심 코드 조각이에요.
Decay 함수를 사용하기 전에 먼저 decay 계산에 사용할 적절한 숫자 필드(타임스탬프, 거리 등)가 있는 컬렉션을 만들어야 해요. 컬렉션 설정, 스키마 정의, 데이터 삽입을 포함한 완전한 작업 예시는 Decay Ranker Tutorial을 참고하세요.
decay ranker 만들기 (Create a decay ranker)
숫자 필드(이 예시에서는 초 단위의 event_date)로 컬렉션을 설정한 뒤 linear decay ranker를 만들어요.
시간 단위 일관성: 시간 기반 decay를 사용할 때는 origin, scale, offset 파라미터가 컬렉션 데이터와 같은 시간 단위를 사용하는지 확인하세요. 컬렉션이 초 단위로 타임스탬프를 저장하면 모든 파라미터도 초를 사용하고, 밀리초라면 모두 밀리초를 사용해야 해요.
from pymilvus import Function, FunctionType
import time
# Calculate current time
current_time = int(time.time())
# Create a linear decay ranker for event listings
# Note: All time parameters must use the same unit as your collection data
ranker = Function(
name="event_relevance", # Function identifier
input_field_names=["event_date"], # Numeric field to use
function_type=FunctionType.RERANK, # Function type. Must be RERANK
params={
"reranker": "decay", # Specify decay reranker
"function": "linear", # Choose linear decay
"origin": current_time, # Current time (seconds, matching collection data)
"offset": 12 * 60 * 60, # 12 hour immediate events window (seconds)
"decay": 0.5, # Half score at scale distance
"scale": 7 * 24 * 60 * 60 # 7 days (in seconds, matching collection data)
}
)
import io.milvus.v2.service.vector.request.ranker.DecayRanker;
DecayRanker ranker = DecayRanker.builder()
.name("event_relevance")
.inputFieldNames(Collections.singletonList("event_date"))
.function("linear")
.origin(System.currentTimeMillis())
.offset(12 * 60 * 60)
.decay(0.5)
.scale(7 * 24 * 60 * 60)
.build();
import { FunctionType } from "@zilliz/milvus2-sdk-node";
const ranker = {
name: "event_relevance",
input_field_names: ["event_date"],
type: FunctionType.RERANK,
params: {
reranker: "decay",
function: "linear",
origin: new Date(2025, 1, 15).getTime(),
offset: 12 * 60 * 60,
decay: 0.5,
scale: 7 * 24 * 60 * 60,
},
};
표준 벡터 검색에 적용하기 (Apply to standard vector search)
Decay ranker를 정의한 뒤, 검색 연산 중 ranker 파라미터에 전달해 적용할 수 있어요.
# Apply decay ranker to vector search
result = milvus_client.search(
collection_name,
data=[your_query_vector], # Replace with your query vector
anns_field="dense", # Vector field to search
limit=10, # Number of results
output_fields=["title", "venue", "event_date"], # Fields to return
ranker=ranker, # Apply the decay ranker
consistency_level="Strong"
)
import io.milvus.v2.common.ConsistencyLevel;
import io.milvus.v2.service.vector.request.SearchReq;
import io.milvus.v2.service.vector.response.SearchResp;
import io.milvus.v2.service.vector.request.data.FloatVec;
SearchReq searchReq = SearchReq.builder()
.collectionName(COLLECTION_NAME)
.data(Collections.singletonList(new FloatVec(embedding)))
.annsField("dense")
.limit(10)
.outputFields(Arrays.asList("title", "venue", "event_date"))
.functionScore(FunctionScore.builder()
.addFunction(ranker)
.build())
.consistencyLevel(ConsistencyLevel.STRONG)
.build();
SearchResp searchResp = client.search(searchReq);
const result = await milvusClient.search({
collection_name: "collection_name",
data: [your_query_vector], // Replace with your query vector
anns_field: "dense",
limit: 10,
output_fields: ["title", "venue", "event_date"],
rerank: ranker,
consistency_level: "Strong",
});
더 알아보기 (Learn more)
- Decay Ranker Tutorial — 컬렉션 설정과 데이터 삽입을 포함한 완전한 예시
- Milvus 공식 문서 — 재순위화(reranking) 관련 자료