Milvus - 벡터 스토어
Milvus - 벡터 스토어 (Vector Store)
RAG용 벡터 스토어로 Milvus를 사용하는 방법을 알아봐요.
출처: 문서
본문
빠른 시작
세 가지가 필요해요:
- Milvus 인스턴스 (cloud 또는 자체 호스팅)
- 임베딩 모델 (쿼리를 벡터로 변환용)
- 벡터 필드가 있는 Milvus 컬렉션
사용법
기본 검색
from litellm import vector_stores
import os
# Set your credentials
os.environ["MILVUS_API_KEY"] = "your-milvus-api-key"
os.environ["MILVUS_API_BASE"] = "https://your-milvus-instance.milvus.io"
# Search the vector store
response = vector_stores.search(
vector_store_id="my-collection-name", # Your Milvus collection name
query="What is the capital of France?",
custom_llm_provider="milvus",
litellm_embedding_model="azure/text-embedding-3-large",
litellm_embedding_config={
"api_base": "your-embedding-endpoint",
"api_key": "your-embedding-api-key",
"api_version": "2025-09-01"
},
milvus_text_field="book_intro", # Field name that contains text content
api_key=os.getenv("MILVUS_API_KEY"),
)
print(response)
비동기 검색
from litellm import vector_stores
response = await vector_stores.asearch(
vector_store_id="my-collection-name",
query="What is the capital of France?",
custom_llm_provider="milvus",
litellm_embedding_model="azure/text-embedding-3-large",
litellm_embedding_config={
"api_base": "your-embedding-endpoint",
"api_key": "your-embedding-api-key",
"api_version": "2025-09-01"
},
milvus_text_field="book_intro",
api_key=os.getenv("MILVUS_API_KEY"),
)
print(response)
고급 옵션
from litellm import vector_stores
response = vector_stores.search(
vector_store_id="my-collection-name",
query="What is the capital of France?",
custom_llm_provider="milvus",
litellm_embedding_model="azure/text-embedding-3-large",
litellm_embedding_config={
"api_base": "your-embedding-endpoint",
"api_key": "your-embedding-api-key",
},
milvus_text_field="book_intro",
api_key=os.getenv("MILVUS_API_KEY"),
# Milvus-specific parameters
limit=10, # Number of results to return
offset=0, # Pagination offset
dbName="default", # Database name
annsField="book_intro_vector", # Vector field name
outputFields=["id", "book_intro", "title"], # Fields to return
filter='book_id > 0', # Metadata filter expression
searchParams={"metric_type": "L2", "params": {"nprobe": 10}}, # Search parameters
)
print(response)
Config 설정
config.yaml에 추가:
vector_store_registry:
- vector_store_name: "milvus-knowledgebase"
litellm_params:
vector_store_id: "my-collection-name"
custom_llm_provider: "milvus"
api_key: os.environ/MILVUS_API_KEY
api_base: https://your-milvus-instance.milvus.io
litellm_embedding_model: "azure/text-embedding-3-large"
litellm_embedding_config:
api_base: https://your-endpoint.cognitiveservices.azure.com/
api_key: os.environ/AZURE_API_KEY
api_version: "2025-09-01"
milvus_text_field: "book_intro"
# Optional Milvus parameters
annsField: "book_intro_vector"
limit: 10
Proxy 시작
litellm --config /path/to/config.yaml
API로 검색
curl -X POST 'http://0.0.0.0:4000/v1/vector_stores/my-collection-name/search' \
-H 'Content-Type: application/json' \
-H "Authorization: Bearer ***" \
-d '{
"query": "What is the capital of France?"
}'
필수 파라미터
| 파라미터 | 타입 | 설명 |
|---|---|---|
vector_store_id |
string | Milvus 컬렉션 이름 |
custom_llm_provider |
string | "milvus"로 설정 |
litellm_embedding_model |
string | 쿼리 임베딩 생성 모델 (예: "azure/text-embedding-3-large") |
litellm_embedding_config |
dict | 임베딩 모델 설정 (api_base, api_key, api_version) |
milvus_text_field |
string | 텍스트 콘텐츠를 담은 컬렉션의 필드 이름 |
api_key |
string | Milvus API 키 (또는 MILVUS_API_KEY env var) |
api_base |
string | Milvus API base URL (또는 MILVUS_API_BASE env var) |
선택 파라미터
| 파라미터 | 타입 | 설명 |
|---|---|---|
dbName |
string | 데이터베이스 이름 (기본: "default") |
annsField |
string | 검색할 벡터 필드 이름 (기본: "book_intro_vector") |
limit |
integer | 반환할 최대 결과 수 |
offset |
integer | 페이지네이션 오프셋 |
filter |
string | 메타데이터 필터링을 위한 필터 식 |
groupingField |
string | 결과를 그룹화할 필드 |
outputFields |
list | 결과에 반환할 필드 목록 |
searchParams |
dict | metric type과 검색 파라미터 같은 검색 파라미터 |
partitionNames |
list | 검색할 파티션 이름 목록 |
consistencyLevel |
string | 검색의 일관성 수준 |
지원 기능
| 기능 | 상태 | 비고 |
|---|---|---|
| Logging | ✅ 지원 | 전체 로깅 지원 |
| Guardrails | ❌ 아직 미지원 | 벡터 스토어용 가드레일은 현재 미지원 |
| Cost Tracking | ✅ 지원 | Milvus 검색 비용은 $0 |
| Unified API | ✅ 지원 | OpenAI 호환 /v1/vector_stores/search 엔드포인트로 호출 |
| Passthrough | ✅ 지원 | 네이티브 Milvus API 형식 사용 |
응답 형식
응답은 표준 LiteLLM 벡터 스토어 형식을 따르며:
{
"object": "vector_store.search_results.page",
"search_query": "What is the capital of France?",
"data": [
{
"score": 0.95,
"content": [
{
"text": "Paris is the capital of France...",
"type": "text"
}
],
"file_id": null,
"filename": null,
"attributes": {
"id": "123",
"title": "France Geography"
}
}
]
}
Passthrough API (네이티브 Milvus 형식)
개발자에게 Milvus 자격 증명을 주지 않고 네이티브 Milvus API 형식으로 벡터 스토어를 만들고 검색하게 하는 데 사용해요. proxy 전용이에요.
관리자 흐름
1. LiteLLM에 벡터 스토어 추가
model_list:
- model_name: embedding-model
litellm_params:
model: azure/text-embedding-3-large
api_base: https://your-endpoint.cognitiveservices.azure.com/
api_key: os.environ/AZURE_API_KEY
api_version: "2025-09-01"
vector_store_registry:
- vector_store_name: "milvus-store"
litellm_params:
vector_store_id: "can-be-anything" # vector store id can be anything for the purpose of passthrough api
custom_llm_provider: "milvus"
api_key: os.environ/MILVUS_API_KEY
api_base: https://your-milvus-instance.milvus.io
general_settings:
database_url: "postgresql://user:***@host:port/database"
master_key: "sk-"
2. Proxy 시작
litellm --config /path/to/config.yaml
# RUNNING on http://0.0.0.0:4000
3. 가상 인덱스 생성
curl -L -X POST 'http://0.0.0.0:4000/v1/indexes' \
-H 'Content-Type: application/json' \
-H "Authorization: Bearer ***" \
-d '{
"index_name": "dall-e-6",
"litellm_params": {
"vector_store_index": "real-collection-name",
"vector_store_name": "milvus-store"
}
}'
이것은 개발자가 벡터 스토어를 만들고 검색하는 데 쓸 수 있는 가상 인덱스예요.
4. 벡터 스토어 권한이 있는 키 만들기
curl -L -X POST 'http://0.0.0.0:4000/key/generate' \
-H 'Content-Type: application/json' \
-H "Authorization: Bearer ***" \
-d '{
"allowed_vector_store_indexes": [{"index_name": "dall-e-6", "index_permissions": ["write", "read"]}],
"models": ["embedding-model"]
}'
키에 가상 인덱스와 임베딩 모델 접근을 주세요.
예상 응답:
{
"key": "«redacted:sk-…»"
}
개발자 흐름
passthrough API를 쓰려면 간단한 REST 클라이언트가 필요해요. 이 milvus_rest_client.py 파일을 프로젝트에 복사하세요:
"""
Simple Milvus REST API v2 Client
Based on: https://milvus.io/api-reference/restful/v2.6.x/
"""
import requests
from typing import List, Dict, Any, Optional
class DataType:
"""Milvus data types"""
INT64 = "Int64"
FLOAT_VECTOR = "FloatVector"
VARCHAR = "VarChar"
BOOL = "Bool"
FLOAT = "Float"
class CollectionSchema:
"""Collection schema builder"""
def __init__(self):
self.fields = []
def add_field(
self,
field_name: str,
data_type: str,
is_primary: bool = False,
dim: Optional[int] = None,
description: str = "",
):
"""Add a field to the schema"""
field = {
"fieldName": field_name,
"dataType": data_type,
"isPrimary": is_primary,
"description": description,
}
if data_type == DataType.FLOAT_VECTOR and dim:
field["elementTypeParams"] = {"dim": str(dim)}
self.fields.append(field)
return self
def to_dict(self):
"""Convert schema to dict for API"""
return {"fields": self.fields}
class IndexParams:
"""Index parameters builder"""
def __init__(self):
self.indexes = []
def add_index(
self, field_name: str, metric_type: str = "L2", index_name: Optional[str] = None
):
"""Add an index"""
index = {
"fieldName": field_name,
"indexName": index_name or f"{field_name}_index",
"metricType": metric_type,
}
self.indexes.append(index)
return self
def to_list(self):
"""Convert to list for API"""
return self.indexes
class MilvusRESTClient:
"""
Simple Milvus REST API v2 Client
Reference: https://milvus.io/api-reference/restful/v2.6.x/
"""
def __init__(self, uri: str, token: str, db_name: str = "default"):
self.base_url = uri.rstrip("/")
self.token = token
self.db_name = db_name
self.headers = {
"Authorization": f"Bearer {token}",
"Content-Type": "application/json",
}
def _make_request(self, endpoint: str, data: Dict[str, Any]) -> Dict[str, Any]:
url = f"{self.base_url}{endpoint}"
if "dbName" not in data and self.db_name != "default":
data["dbName"] = self.db_name
try:
response = requests.post(url, json=data, headers=self.headers)
response.raise_for_status()
except requests.exceptions.HTTPError as e:
print(f"e.response.text: {e.response.content}")
raise e
result = response.json()
if result.get("code") != 0:
raise Exception(
f"Milvus API Error: {result.get('message', 'Unknown error')}"
)
return result
def has_collection(self, collection_name: str) -> bool:
try:
result = self._make_request(
"/v2/vectordb/collections/has", {"collectionName": collection_name}
)
return result.get("data", {}).get("has", False)
except Exception:
return False
def drop_collection(self, collection_name: str):
return self._make_request(
"/v2/vectordb/collections/drop", {"collectionName": collection_name}
)
def create_schema(self) -> CollectionSchema:
return CollectionSchema()
def prepare_index_params(self) -> IndexParams:
return IndexParams()
def create_collection(
self,
collection_name: str,
schema: CollectionSchema,
index_params: Optional[IndexParams] = None,
):
data = {"collectionName": collection_name, "schema": schema.to_dict()}
if index_params:
data["indexParams"] = index_params.to_list()
return self._make_request("/v2/vectordb/collections/create", data)
def describe_collection(self, collection_name: str) -> Dict[str, Any]:
result = self._make_request(
"/v2/vectordb/collections/describe", {"collectionName": collection_name}
)
return result.get("data", {})
def insert(
self,
collection_name: str,
data: List[Dict[str, Any]],
partition_name: Optional[str] = None,
):
payload = {"collectionName": collection_name, "data": data}
if partition_name:
payload["partitionName"] = partition_name
result = self._make_request("/v2/vectordb/entities/insert", payload)
return result.get("data", {})
def flush(self, collection_name: str):
return self._make_request(
"/v2/vectordb/collections/flush", {"collectionName": collection_name}
)
def search(
self,
collection_name: str,
data: List[List[float]],
anns_field: str,
limit: int = 10,
search_params: Optional[Dict[str, Any]] = None,
output_fields: Optional[List[str]] = None,
) -> List[List[Dict]]:
payload = {
"collectionName": collection_name,
"data": data,
"annsField": anns_field,
"limit": limit,
}
if search_params:
payload["searchParams"] = search_params
if output_fields:
payload["outputFields"] = output_fields
result = self._make_request("/v2/vectordb/entities/search", payload)
return result.get("data", [])
1. 스키마로 컬렉션 생성
참고: passthrough api에는 config의 milvus 제공사를 쓰는 /milvus 엔드포인트를 사용해요.
from milvus_rest_client import MilvusRESTClient, DataType # Use the client from above
import random
import time
# Configuration
uri = "http://0.0.0.0:4000/milvus" # IMPORTANT: Use the '/milvus' endpoint for passthrough
token = "«redacted:sk-…»"
collection_name = "dall-e-6" # Virtual index name
# Initialize client
milvus_client = MilvusRESTClient(uri=uri, token=token)
print(f"Connected to DB: {uri} successfully")
# Check if the collection exists and drop if it does
check_collection = milvus_client.has_collection(collection_name)
if check_collection:
milvus_client.drop_collection(collection_name)
print(f"Dropped the existing collection {collection_name} successfully")
# Define schema
dim = 64 # Vector dimension
print("Start to create the collection schema")
schema = milvus_client.create_schema()
schema.add_field(
"book_id", DataType.INT64, is_primary=True, description="customized primary id"
)
schema.add_field("word_count", DataType.INT64, description="word count")
schema.add_field(
"book_intro", DataType.FLOAT_VECTOR, dim=dim, description="book introduction"
)
# Prepare index parameters
print("Start to prepare index parameters with default AUTOINDEX")
index_params = milvus_client.prepare_index_params()
index_params.add_index("book_intro", metric_type="L2")
# Create collection
print(f"Start to create example collection: {collection_name}")
milvus_client.create_collection(
collection_name, schema=schema, index_params=index_params
)
collection_property = milvus_client.describe_collection(collection_name)
print("Collection details: %s" % collection_property)
2. 컬렉션에 데이터 삽입
# Insert data with customized ids
nb = 1000
insert_rounds = 2
start = 0 # first primary key id
total_rt = 0 # total response time for insert
print(
f"Start to insert {nb*insert_rounds} entities into example collection: {collection_name}"
)
for i in range(insert_rounds):
vector = [random.random() for _ in range(dim)]
rows = [
{"book_id": i, "word_count": random.randint(1, 100), "book_intro": vector}
for i in range(start, start + nb)
]
t0 = time.time()
milvus_client.insert(collection_name, rows)
ins_rt = time.time() - t0
start += nb
total_rt += ins_rt
print(f"Insert completed in {round(total_rt, 4)} seconds")
# Flush the collection
print("Start to flush")
start_flush = time.time()
milvus_client.flush(collection_name)
end_flush = time.time()
print(f"Flush completed in {round(end_flush - start_flush, 4)} seconds")
3. 컬렉션 검색
# Search configuration
nq = 3 # Number of query vectors
search_params = {"metric_type": "L2", "params": {"level": 2}}
limit = 2 # Number of results to return
# Perform searches
for i in range(5):
search_vectors = [[random.random() for _ in range(dim)] for _ in range(nq)]
t0 = time.time()
results = milvus_client.search(
collection_name,
data=search_vectors,
limit=limit,
search_params=search_params,
anns_field="book_intro",
)
t1 = time.time()
print(f"Search {i} results: {results}")
print(f"Search {i} latency: {round(t1-t0, 4)} seconds")
동작 원리
검색 시:
- LiteLLM이 지정한 임베딩 모델로 쿼리를 벡터로 변환
/v2/vectordb/entities/search엔드포인트로 벡터를 Milvus 인스턴스에 전송- Milvus가 벡터 유사도 검색으로 컬렉션에서 가장 유사한 문서를 찾음
- 결과가 거리 점수와 함께 반환
임베딩 모델은 LiteLLM이 지원하는 어떤 모델이든 가능해요 (Azure OpenAI, OpenAI, Bedrock 등).
더 알아보기 (Learn more)
- Milvus 공식 문서
- LiteLLM 벡터 스토어 레지스트리