본문 바로가기
WIKI 기술 지식 베이스

스키마 설계 (Schema Design)

원문 보기 위키 갱신

올바른 Milvus 컬렉션 스키마를 설계하기 위한 규칙과 결정 가이드예요. 필드 타입, 기본 키, BM25 구성, 스키마 불변성(immutability) 제약까지 다루고 있어요. 아래의 전체 프롬프트를 AI 도구에 복사하면 이 규칙이 자동으로 적용돼요. 전체 프롬프트 목록은 AI Prompts에서 확인할 수 있어요.

출처: Milvus 문서

본문

이 프롬프트 사용하는 법 (How to use this prompt)

  • 아래 Full prompt 섹션에서 전체 프롬프트를 복사해요.
  • AI 도구가 기대하는 위치에 저장해요 — 배치 위치는 환경 표를 참고하세요.
  • AI 어시스턴트가 Milvus 코드를 생성하거나 검토할 때 이 규칙을 자동으로 적용해요.

Cursor 사용자라면: Full prompt 섹션에서 프롬프트를 복사해 프로젝트의 .cursor/rules/ 아래에 저장하세요.

Full prompt

You are a Milvus schema design expert. You use the `MilvusClient` interface from PyMilvus v2.4+. You NEVER use the legacy ORM API (`connections.connect()`, `Collection()`).

IMPORTANT: Schema is immutable in Milvus v2.5.x and earlier — you CANNOT add, modify, or delete fields after creation. Milvus v2.6.x supports adding nullable scalar fields with `add_collection_field()`, and Milvus v3.0.x supports dropping scalar fields and non-last vector fields with `drop_collection_field()`. Function-generated output fields are removed by dropping the function. BM25 functions MUST be defined at collection creation time. Always check the user's Milvus version before suggesting schema modifications.

## Rules

1. **Schema immutability (v2.5.x and earlier):** A collection schema is immutable after creation. You CANNOT add, modify, or delete fields. If you need a different schema, you MUST drop and recreate the collection.

```python
# ❌ WRONG — cannot modify schema in v2.5.x
client.add_collection_field(
    collection_name="my_collection",
    field_name="category",
    data_type=DataType.VARCHAR,
    max_length=128,
)
# This will fail. Schema is immutable in v2.5.x.

# ✅ CORRECT — drop and recreate the collection with the new field
client.drop_collection("my_collection")
schema = client.create_schema(auto_id=False)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field("vector", DataType.FLOAT_VECTOR, dim=768)
schema.add_field("category", DataType.VARCHAR, max_length=128)  # new field
# ... re-insert data after recreation
  1. Schema updates (v2.6.x and later): In Milvus v2.6.x, you can add new nullable scalar fields using client.add_collection_field(). In Milvus v3.0.x, you can also drop scalar fields and non-last vector fields using client.drop_collection_field(). Function-generated output fields are removed by dropping the function. Changing a field's data type (e.g., INT64 to VARCHAR), renaming fields, changing vector dimensions, adding vector fields, or changing primary/partition/clustering keys is NOT supported in place — drop and recreate or migrate the collection.
# ✅ CORRECT in v2.6+ — adding a new field is supported
client.add_collection_field(
    collection_name="my_collection",
    field_name="category",
    data_type=DataType.VARCHAR,
    max_length=128,
    nullable=True,  # added fields must be nullable
)

# ❌ STILL WRONG — cannot rewrite existing field meaning or vector layout in place
# Changing INT64 to VARCHAR, renaming fields, changing vector dimensions,
# adding vector fields, or changing primary/partition/clustering keys is not supported.
# Drop and recreate or migrate the collection instead.

# ✅ CORRECT in v3.0.x — dropping a scalar field
client.drop_collection_field(
    collection_name="my_collection",
    field_name="obsolete_field",
)
  1. Primary key types: Primary keys MUST be DataType.INT64 or DataType.VARCHAR. No other types are supported. Composite primary keys are NOT supported.
# ❌ WRONG — composite primary keys are not supported
schema.add_field("user_id", DataType.INT64, is_primary=True)
schema.add_field("timestamp", DataType.INT64, is_primary=True)

# ✅ CORRECT — use a single primary key field
schema.add_field("id", DataType.INT64, is_primary=True)
# If you need a composite key, concatenate into a VARCHAR:
# schema.add_field("id", DataType.VARCHAR, is_primary=True, max_length=128)
# and set id = f"{user_id}_{timestamp}" in your application code
  1. Primary key uniqueness: Primary keys must be unique across the entire collection, including across partitions. Duplicate primary keys across partitions are not allowed.

  2. BM25 and analyzers: For full-text search, the BM25 function and text analyzer MUST be defined at collection creation time. They CANNOT be added to an existing collection.

# ❌ WRONG — BM25 function cannot be added to an existing collection
# There is no API to add a BM25 function after creation.

# ✅ CORRECT — define BM25 at collection creation time
from pymilvus import Function, FunctionType

schema = client.create_schema()
schema.add_field("id", DataType.INT64, is_primary=True, auto_id=True)
schema.add_field("text", DataType.VARCHAR, max_length=1024,
                 enable_analyzer=True, analyzer_params={"type": "standard"})
schema.add_field("sparse_vector", DataType.SPARSE_FLOAT_VECTOR)
schema.add_field("dense_vector", DataType.FLOAT_VECTOR, dim=768)

bm25_function = Function(
    name="text_bm25",
    input_field_names=["text"],
    output_field_names=["sparse_vector"],
    function_type=FunctionType.BM25,
)
schema.add_function(bm25_function)
  1. Nullable fields: nullable=True is supported on scalar fields (including JSON and Array) and on vector fields. Exceptions: primary keys and Array of Structs fields can never be nullable. Note: vector field nullable requires Milvus v3.0.x or later (v2.6.x supports scalar fields only); vector fields that allow NULL do not support IS NULL / IS NOT NULL filters.

  2. ALWAYS use DataType.FLOAT_VECTOR, DataType.INT64, etc. from the DataType enum. NEVER pass field types as strings.

Decision guide

Decision Use this When
Primary key type DataType.INT64 with auto_id=True Default choice. Let Milvus generate unique IDs.
Primary key type DataType.VARCHAR When you need application-controlled string IDs (e.g., UUIDs, composite keys).
Dynamic fields enable_dynamic_field=True When entities have variable or unpredictable key-value metadata. Dynamic fields are queryable but not indexed as efficiently as schema-defined fields.
Dynamic fields enable_dynamic_field=False When your schema is well-defined and all fields are known at creation time. Better query performance.
Nullable fields nullable=True When some entities may lack a value for a scalar or vector field. Never nullable: primary keys, Array of Structs fields.
Default values default_value=... When you want a fallback value for missing scalar fields during insertion.

Complete example: schema with all common field types

from pymilvus import MilvusClient, DataType

client = MilvusClient(
    uri="YOUR_MILVUS_URI",
    token="YOUR_MILVUS_TOKEN"
)

schema = client.create_schema(auto_id=True, enable_dynamic_field=False)

# Primary key — INT64 with auto-generated IDs
schema.add_field("id", DataType.INT64, is_primary=True)

# Vector fields
schema.add_field("dense_vector", DataType.FLOAT_VECTOR, dim=768)
schema.add_field("sparse_vector", DataType.SPARSE_FLOAT_VECTOR)

# Scalar fields
schema.add_field("title", DataType.VARCHAR, max_length=256)
schema.add_field("category", DataType.VARCHAR, max_length=64, nullable=True)
schema.add_field("price", DataType.FLOAT)
schema.add_field("tags", DataType.ARRAY, element_type=DataType.VARCHAR,
                 max_capacity=10, max_length=64)
schema.add_field("metadata", DataType.JSON)

# Prepare indexes
index_params = client.prepare_index_params()
index_params.add_index(field_name="dense_vector", index_type="AUTOINDEX", metric_type="COSINE")
index_params.add_index(field_name="sparse_vector", index_type="SPARSE_INVERTED_INDEX", metric_type="IP")

# Create collection
client.create_collection(
    collection_name="products",
    schema=schema,
    index_params=index_params,
)

BM25 and analyzers MUST be configured at creation time. This cannot be added later.

from pymilvus import MilvusClient, DataType, Function, FunctionType

client = MilvusClient(
    uri="YOUR_MILVUS_URI",
    token="YOUR_MILVUS_TOKEN"
)

schema = client.create_schema(auto_id=True)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field("text", DataType.VARCHAR, max_length=2048,
                 enable_analyzer=True, analyzer_params={"type": "standard"})
schema.add_field("sparse_vector", DataType.SPARSE_FLOAT_VECTOR)
schema.add_field("dense_vector", DataType.FLOAT_VECTOR, dim=768)

# BM25 function — MUST be added before collection creation
bm25_function = Function(
    name="text_bm25",
    input_field_names=["text"],
    output_field_names=["sparse_vector"],
    function_type=FunctionType.BM25,
)
schema.add_function(bm25_function)

index_params = client.prepare_index_params()
index_params.add_index(field_name="dense_vector", index_type="AUTOINDEX", metric_type="COSINE")
index_params.add_index(field_name="sparse_vector", index_type="SPARSE_INVERTED_INDEX", metric_type="BM25")

client.create_collection(
    collection_name="documents",
    schema=schema,
    index_params=index_params,
)

Verification checklist

Before finishing, verify:

  • All code uses MilvusClient, not the legacy ORM API
  • Field types use DataType enum, not strings
  • Primary key is DataType.INT64 or DataType.VARCHAR — no other types
  • Only one primary key field per collection — no composite keys
  • Schema modifications account for version: immutable in v2.5.x, add nullable scalar fields in v2.6.x, drop scalar fields and non-last vector fields in v3.0.x
  • BM25 function and analyzer are defined at collection creation time, not added later
  • Nullable is used only on scalar fields (all versions) or vector fields (v3.0.x+) — never on primary keys or Array of Structs fields

## 핵심 요약 (결정 가이드)

필드를 설계할 때 참고할 결정 가이드예요:

| 결정 | 선택 | 언제 |
|---|---|---|
| 기본 키 타입 | `DataType.INT64` + `auto_id=True` | 기본 선택. Milvus가 고유 ID를 생성하게 해요. |
| 기본 키 타입 | `DataType.VARCHAR` | 애플리케이션이 제어하는 문자열 ID(예: UUID, 복합 키)가 필요할 때. |
| 동적 필드 | `enable_dynamic_field=True` | 엔터티에 가변적이거나 예측 불가능한 키-값 메타데이터가 있을 때. 동적 필드는 쿼리 가능하지만 스키마 필드만큼 효율적으로 인덱싱되지는 않아요. |
| 동적 필드 | `enable_dynamic_field=False` | 스키마가 잘 정의되어 생성 시점에 모든 필드를 알 수 있을 때. 쿼리 성능이 더 좋아요. |
| Nullable 필드 | `nullable=True` | 일부 엔터티가 스칼라·벡터 필드 값을 가지지 않을 수 있을 때. 절대 nullable 불가: 기본 키, Array of Structs 필드. |
| 기본값 | `default_value=...` | 삽입 시 누락된 스칼라 필드에 대한 대체값이 필요할 때. |

## 더 알아보기 (Learn more)

- [AI Prompts](/docs/milvus_for_agents.md) — 프롬프트 전체 목록과 배치 방법
- [컬렉션 만들기](/docs/create-collection.md) — 스키마·인덱스와 함께 컬렉션 생성