Milvus - 벡터 스토어

Milvus - 벡터 스토어 (Vector Store)

RAG용 벡터 스토어로 Milvus를 사용하는 방법을 알아봐요.

출처: 문서

본문

빠른 시작

세 가지가 필요해요:

  • Milvus 인스턴스 (cloud 또는 자체 호스팅)
  • 임베딩 모델 (쿼리를 벡터로 변환용)
  • 벡터 필드가 있는 Milvus 컬렉션

사용법

기본 검색

from litellm import vector_stores
import os

# Set your credentials
os.environ["MILVUS_API_KEY"] = "your-milvus-api-key"
os.environ["MILVUS_API_BASE"] = "https://your-milvus-instance.milvus.io"

# Search the vector store
response = vector_stores.search(
    vector_store_id="my-collection-name",  # Your Milvus collection name
    query="What is the capital of France?",
    custom_llm_provider="milvus",
    litellm_embedding_model="azure/text-embedding-3-large",
    litellm_embedding_config={
        "api_base": "your-embedding-endpoint",
        "api_key": "your-embedding-api-key",
        "api_version": "2025-09-01"
    },
    milvus_text_field="book_intro",  # Field name that contains text content
    api_key=os.getenv("MILVUS_API_KEY"),
)

print(response)

비동기 검색

from litellm import vector_stores

response = await vector_stores.asearch(
    vector_store_id="my-collection-name",
    query="What is the capital of France?",
    custom_llm_provider="milvus",
    litellm_embedding_model="azure/text-embedding-3-large",
    litellm_embedding_config={
        "api_base": "your-embedding-endpoint",
        "api_key": "your-embedding-api-key",
        "api_version": "2025-09-01"
    },
    milvus_text_field="book_intro",
    api_key=os.getenv("MILVUS_API_KEY"),
)

print(response)

고급 옵션

from litellm import vector_stores

response = vector_stores.search(
    vector_store_id="my-collection-name",
    query="What is the capital of France?",
    custom_llm_provider="milvus",
    litellm_embedding_model="azure/text-embedding-3-large",
    litellm_embedding_config={
        "api_base": "your-embedding-endpoint",
        "api_key": "your-embedding-api-key",
    },
    milvus_text_field="book_intro",
    api_key=os.getenv("MILVUS_API_KEY"),
    # Milvus-specific parameters
    limit=10,  # Number of results to return
    offset=0,  # Pagination offset
    dbName="default",  # Database name
    annsField="book_intro_vector",  # Vector field name
    outputFields=["id", "book_intro", "title"],  # Fields to return
    filter='book_id > 0',  # Metadata filter expression
    searchParams={"metric_type": "L2", "params": {"nprobe": 10}},  # Search parameters
)

print(response)

Config 설정

config.yaml에 추가:

vector_store_registry:
  - vector_store_name: "milvus-knowledgebase"
    litellm_params:
        vector_store_id: "my-collection-name"
        custom_llm_provider: "milvus"
        api_key: os.environ/MILVUS_API_KEY
        api_base: https://your-milvus-instance.milvus.io
        litellm_embedding_model: "azure/text-embedding-3-large"
        litellm_embedding_config:
            api_base: https://your-endpoint.cognitiveservices.azure.com/
            api_key: os.environ/AZURE_API_KEY
            api_version: "2025-09-01"
        milvus_text_field: "book_intro"
        # Optional Milvus parameters
        annsField: "book_intro_vector"
        limit: 10

Proxy 시작

litellm --config /path/to/config.yaml

API로 검색

curl -X POST 'http://0.0.0.0:4000/v1/vector_stores/my-collection-name/search' \
-H 'Content-Type: application/json' \
-H "Authorization: Bearer ***" \
-d '{
  "query": "What is the capital of France?"
}'

필수 파라미터

파라미터 타입 설명
vector_store_id string Milvus 컬렉션 이름
custom_llm_provider string "milvus"로 설정
litellm_embedding_model string 쿼리 임베딩 생성 모델 (예: "azure/text-embedding-3-large")
litellm_embedding_config dict 임베딩 모델 설정 (api_base, api_key, api_version)
milvus_text_field string 텍스트 콘텐츠를 담은 컬렉션의 필드 이름
api_key string Milvus API 키 (또는 MILVUS_API_KEY env var)
api_base string Milvus API base URL (또는 MILVUS_API_BASE env var)

선택 파라미터

파라미터 타입 설명
dbName string 데이터베이스 이름 (기본: "default")
annsField string 검색할 벡터 필드 이름 (기본: "book_intro_vector")
limit integer 반환할 최대 결과 수
offset integer 페이지네이션 오프셋
filter string 메타데이터 필터링을 위한 필터 식
groupingField string 결과를 그룹화할 필드
outputFields list 결과에 반환할 필드 목록
searchParams dict metric type과 검색 파라미터 같은 검색 파라미터
partitionNames list 검색할 파티션 이름 목록
consistencyLevel string 검색의 일관성 수준

지원 기능

기능 상태 비고
Logging ✅ 지원 전체 로깅 지원
Guardrails ❌ 아직 미지원 벡터 스토어용 가드레일은 현재 미지원
Cost Tracking ✅ 지원 Milvus 검색 비용은 $0
Unified API ✅ 지원 OpenAI 호환 /v1/vector_stores/search 엔드포인트로 호출
Passthrough ✅ 지원 네이티브 Milvus API 형식 사용

응답 형식

응답은 표준 LiteLLM 벡터 스토어 형식을 따르며:

{
  "object": "vector_store.search_results.page",
  "search_query": "What is the capital of France?",
  "data": [
    {
      "score": 0.95,
      "content": [
        {
          "text": "Paris is the capital of France...",
          "type": "text"
        }
      ],
      "file_id": null,
      "filename": null,
      "attributes": {
        "id": "123",
        "title": "France Geography"
      }
    }
  ]
}

Passthrough API (네이티브 Milvus 형식)

개발자에게 Milvus 자격 증명을 주지 않고 네이티브 Milvus API 형식으로 벡터 스토어를 만들고 검색하게 하는 데 사용해요. proxy 전용이에요.

관리자 흐름

1. LiteLLM에 벡터 스토어 추가

model_list:
  - model_name: embedding-model
    litellm_params:
      model: azure/text-embedding-3-large
      api_base: https://your-endpoint.cognitiveservices.azure.com/
      api_key: os.environ/AZURE_API_KEY
      api_version: "2025-09-01"

vector_store_registry:
  - vector_store_name: "milvus-store"
    litellm_params:
      vector_store_id: "can-be-anything" # vector store id can be anything for the purpose of passthrough api
      custom_llm_provider: "milvus"
      api_key: os.environ/MILVUS_API_KEY
      api_base: https://your-milvus-instance.milvus.io

general_settings:
    database_url: "postgresql://user:***@host:port/database"
    master_key: "sk-"

2. Proxy 시작

litellm --config /path/to/config.yaml

# RUNNING on http://0.0.0.0:4000

3. 가상 인덱스 생성

curl -L -X POST 'http://0.0.0.0:4000/v1/indexes' \
-H 'Content-Type: application/json' \
-H "Authorization: Bearer ***" \
-d '{
    "index_name": "dall-e-6",
    "litellm_params": {
        "vector_store_index": "real-collection-name",
        "vector_store_name": "milvus-store"
    }
}'

이것은 개발자가 벡터 스토어를 만들고 검색하는 데 쓸 수 있는 가상 인덱스예요.

4. 벡터 스토어 권한이 있는 키 만들기

curl -L -X POST 'http://0.0.0.0:4000/key/generate' \
-H 'Content-Type: application/json' \
-H "Authorization: Bearer ***" \
-d '{
    "allowed_vector_store_indexes": [{"index_name": "dall-e-6", "index_permissions": ["write", "read"]}],
    "models": ["embedding-model"]
}'

키에 가상 인덱스와 임베딩 모델 접근을 주세요.

예상 응답:

{
    "key": "«redacted:sk-…»"
}

개발자 흐름

passthrough API를 쓰려면 간단한 REST 클라이언트가 필요해요. 이 milvus_rest_client.py 파일을 프로젝트에 복사하세요:

"""
Simple Milvus REST API v2 Client
Based on: https://milvus.io/api-reference/restful/v2.6.x/
"""

import requests
from typing import List, Dict, Any, Optional

class DataType:
    """Milvus data types"""

    INT64 = "Int64"
    FLOAT_VECTOR = "FloatVector"
    VARCHAR = "VarChar"
    BOOL = "Bool"
    FLOAT = "Float"

class CollectionSchema:
    """Collection schema builder"""

    def __init__(self):
        self.fields = []

    def add_field(
        self,
        field_name: str,
        data_type: str,
        is_primary: bool = False,
        dim: Optional[int] = None,
        description: str = "",
    ):
        """Add a field to the schema"""
        field = {
            "fieldName": field_name,
            "dataType": data_type,
            "isPrimary": is_primary,
            "description": description,
        }
        if data_type == DataType.FLOAT_VECTOR and dim:
            field["elementTypeParams"] = {"dim": str(dim)}
        self.fields.append(field)
        return self

    def to_dict(self):
        """Convert schema to dict for API"""
        return {"fields": self.fields}

class IndexParams:
    """Index parameters builder"""

    def __init__(self):
        self.indexes = []

    def add_index(
        self, field_name: str, metric_type: str = "L2", index_name: Optional[str] = None
    ):
        """Add an index"""
        index = {
            "fieldName": field_name,
            "indexName": index_name or f"{field_name}_index",
            "metricType": metric_type,
        }
        self.indexes.append(index)
        return self

    def to_list(self):
        """Convert to list for API"""
        return self.indexes

class MilvusRESTClient:
    """
    Simple Milvus REST API v2 Client

    Reference: https://milvus.io/api-reference/restful/v2.6.x/
    """

    def __init__(self, uri: str, token: str, db_name: str = "default"):
        self.base_url = uri.rstrip("/")
        self.token = token
        self.db_name = db_name
        self.headers = {
            "Authorization": f"Bearer {token}",
            "Content-Type": "application/json",
        }

    def _make_request(self, endpoint: str, data: Dict[str, Any]) -> Dict[str, Any]:
        url = f"{self.base_url}{endpoint}"
        if "dbName" not in data and self.db_name != "default":
            data["dbName"] = self.db_name
        try:
            response = requests.post(url, json=data, headers=self.headers)
            response.raise_for_status()
        except requests.exceptions.HTTPError as e:
            print(f"e.response.text: {e.response.content}")
            raise e
        result = response.json()
        if result.get("code") != 0:
            raise Exception(
                f"Milvus API Error: {result.get('message', 'Unknown error')}"
            )
        return result

    def has_collection(self, collection_name: str) -> bool:
        try:
            result = self._make_request(
                "/v2/vectordb/collections/has", {"collectionName": collection_name}
            )
            return result.get("data", {}).get("has", False)
        except Exception:
            return False

    def drop_collection(self, collection_name: str):
        return self._make_request(
            "/v2/vectordb/collections/drop", {"collectionName": collection_name}
        )

    def create_schema(self) -> CollectionSchema:
        return CollectionSchema()

    def prepare_index_params(self) -> IndexParams:
        return IndexParams()

    def create_collection(
        self,
        collection_name: str,
        schema: CollectionSchema,
        index_params: Optional[IndexParams] = None,
    ):
        data = {"collectionName": collection_name, "schema": schema.to_dict()}
        if index_params:
            data["indexParams"] = index_params.to_list()
        return self._make_request("/v2/vectordb/collections/create", data)

    def describe_collection(self, collection_name: str) -> Dict[str, Any]:
        result = self._make_request(
            "/v2/vectordb/collections/describe", {"collectionName": collection_name}
        )
        return result.get("data", {})

    def insert(
        self,
        collection_name: str,
        data: List[Dict[str, Any]],
        partition_name: Optional[str] = None,
    ):
        payload = {"collectionName": collection_name, "data": data}
        if partition_name:
            payload["partitionName"] = partition_name
        result = self._make_request("/v2/vectordb/entities/insert", payload)
        return result.get("data", {})

    def flush(self, collection_name: str):
        return self._make_request(
            "/v2/vectordb/collections/flush", {"collectionName": collection_name}
        )

    def search(
        self,
        collection_name: str,
        data: List[List[float]],
        anns_field: str,
        limit: int = 10,
        search_params: Optional[Dict[str, Any]] = None,
        output_fields: Optional[List[str]] = None,
    ) -> List[List[Dict]]:
        payload = {
            "collectionName": collection_name,
            "data": data,
            "annsField": anns_field,
            "limit": limit,
        }
        if search_params:
            payload["searchParams"] = search_params
        if output_fields:
            payload["outputFields"] = output_fields
        result = self._make_request("/v2/vectordb/entities/search", payload)
        return result.get("data", [])

1. 스키마로 컬렉션 생성

참고: passthrough api에는 config의 milvus 제공사를 쓰는 /milvus 엔드포인트를 사용해요.

from milvus_rest_client import MilvusRESTClient, DataType  # Use the client from above
import random
import time

# Configuration
uri = "http://0.0.0.0:4000/milvus"  # IMPORTANT: Use the '/milvus' endpoint for passthrough
token = "«redacted:sk-…»"
collection_name = "dall-e-6"  # Virtual index name

# Initialize client
milvus_client = MilvusRESTClient(uri=uri, token=token)
print(f"Connected to DB: {uri} successfully")

# Check if the collection exists and drop if it does
check_collection = milvus_client.has_collection(collection_name)
if check_collection:
    milvus_client.drop_collection(collection_name)
    print(f"Dropped the existing collection {collection_name} successfully")

# Define schema
dim = 64  # Vector dimension

print("Start to create the collection schema")
schema = milvus_client.create_schema()
schema.add_field(
    "book_id", DataType.INT64, is_primary=True, description="customized primary id"
)
schema.add_field("word_count", DataType.INT64, description="word count")
schema.add_field(
    "book_intro", DataType.FLOAT_VECTOR, dim=dim, description="book introduction"
)

# Prepare index parameters
print("Start to prepare index parameters with default AUTOINDEX")
index_params = milvus_client.prepare_index_params()
index_params.add_index("book_intro", metric_type="L2")

# Create collection
print(f"Start to create example collection: {collection_name}")
milvus_client.create_collection(
    collection_name, schema=schema, index_params=index_params
)
collection_property = milvus_client.describe_collection(collection_name)
print("Collection details: %s" % collection_property)

2. 컬렉션에 데이터 삽입

# Insert data with customized ids
nb = 1000
insert_rounds = 2
start = 0  # first primary key id
total_rt = 0  # total response time for insert

print(
    f"Start to insert {nb*insert_rounds} entities into example collection: {collection_name}"
)
for i in range(insert_rounds):
    vector = [random.random() for _ in range(dim)]
    rows = [
        {"book_id": i, "word_count": random.randint(1, 100), "book_intro": vector}
        for i in range(start, start + nb)
    ]
    t0 = time.time()
    milvus_client.insert(collection_name, rows)
    ins_rt = time.time() - t0
    start += nb
    total_rt += ins_rt
print(f"Insert completed in {round(total_rt, 4)} seconds")

# Flush the collection
print("Start to flush")
start_flush = time.time()
milvus_client.flush(collection_name)
end_flush = time.time()
print(f"Flush completed in {round(end_flush - start_flush, 4)} seconds")

3. 컬렉션 검색

# Search configuration
nq = 3  # Number of query vectors
search_params = {"metric_type": "L2", "params": {"level": 2}}
limit = 2  # Number of results to return

# Perform searches
for i in range(5):
    search_vectors = [[random.random() for _ in range(dim)] for _ in range(nq)]
    t0 = time.time()
    results = milvus_client.search(
        collection_name,
        data=search_vectors,
        limit=limit,
        search_params=search_params,
        anns_field="book_intro",
    )
    t1 = time.time()
    print(f"Search {i} results: {results}")
    print(f"Search {i} latency: {round(t1-t0, 4)} seconds")

동작 원리

검색 시:

  • LiteLLM이 지정한 임베딩 모델로 쿼리를 벡터로 변환
  • /v2/vectordb/entities/search 엔드포인트로 벡터를 Milvus 인스턴스에 전송
  • Milvus가 벡터 유사도 검색으로 컬렉션에서 가장 유사한 문서를 찾음
  • 결과가 거리 점수와 함께 반환

임베딩 모델은 LiteLLM이 지원하는 어떤 모델이든 가능해요 (Azure OpenAI, OpenAI, Bedrock 등).

더 알아보기 (Learn more)

  • Milvus 공식 문서
  • LiteLLM 벡터 스토어 레지스트리