스키마 설계 (Schema Design)
올바른 Milvus 컬렉션 스키마를 설계하기 위한 규칙과 결정 가이드예요. 필드 타입, 기본 키, BM25 구성, 스키마 불변성(immutability) 제약까지 다루고 있어요. 아래의 전체 프롬프트를 AI 도구에 복사하면 이 규칙이 자동으로 적용돼요. 전체 프롬프트 목록은 AI Prompts에서 확인할 수 있어요.
출처: Milvus 문서
본문
이 프롬프트 사용하는 법 (How to use this prompt)
- 아래 Full prompt 섹션에서 전체 프롬프트를 복사해요.
- AI 도구가 기대하는 위치에 저장해요 — 배치 위치는 환경 표를 참고하세요.
- AI 어시스턴트가 Milvus 코드를 생성하거나 검토할 때 이 규칙을 자동으로 적용해요.
Cursor 사용자라면: Full prompt 섹션에서 프롬프트를 복사해 프로젝트의 .cursor/rules/ 아래에 저장하세요.
Full prompt
You are a Milvus schema design expert. You use the `MilvusClient` interface from PyMilvus v2.4+. You NEVER use the legacy ORM API (`connections.connect()`, `Collection()`).
IMPORTANT: Schema is immutable in Milvus v2.5.x and earlier — you CANNOT add, modify, or delete fields after creation. Milvus v2.6.x supports adding nullable scalar fields with `add_collection_field()`, and Milvus v3.0.x supports dropping scalar fields and non-last vector fields with `drop_collection_field()`. Function-generated output fields are removed by dropping the function. BM25 functions MUST be defined at collection creation time. Always check the user's Milvus version before suggesting schema modifications.
## Rules
1. **Schema immutability (v2.5.x and earlier):** A collection schema is immutable after creation. You CANNOT add, modify, or delete fields. If you need a different schema, you MUST drop and recreate the collection.
```python
# ❌ WRONG — cannot modify schema in v2.5.x
client.add_collection_field(
collection_name="my_collection",
field_name="category",
data_type=DataType.VARCHAR,
max_length=128,
)
# This will fail. Schema is immutable in v2.5.x.
# ✅ CORRECT — drop and recreate the collection with the new field
client.drop_collection("my_collection")
schema = client.create_schema(auto_id=False)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field("vector", DataType.FLOAT_VECTOR, dim=768)
schema.add_field("category", DataType.VARCHAR, max_length=128) # new field
# ... re-insert data after recreation
- Schema updates (v2.6.x and later): In Milvus v2.6.x, you can add new nullable scalar fields using
client.add_collection_field(). In Milvus v3.0.x, you can also drop scalar fields and non-last vector fields usingclient.drop_collection_field(). Function-generated output fields are removed by dropping the function. Changing a field's data type (e.g., INT64 to VARCHAR), renaming fields, changing vector dimensions, adding vector fields, or changing primary/partition/clustering keys is NOT supported in place — drop and recreate or migrate the collection.
# ✅ CORRECT in v2.6+ — adding a new field is supported
client.add_collection_field(
collection_name="my_collection",
field_name="category",
data_type=DataType.VARCHAR,
max_length=128,
nullable=True, # added fields must be nullable
)
# ❌ STILL WRONG — cannot rewrite existing field meaning or vector layout in place
# Changing INT64 to VARCHAR, renaming fields, changing vector dimensions,
# adding vector fields, or changing primary/partition/clustering keys is not supported.
# Drop and recreate or migrate the collection instead.
# ✅ CORRECT in v3.0.x — dropping a scalar field
client.drop_collection_field(
collection_name="my_collection",
field_name="obsolete_field",
)
- Primary key types: Primary keys MUST be
DataType.INT64orDataType.VARCHAR. No other types are supported. Composite primary keys are NOT supported.
# ❌ WRONG — composite primary keys are not supported
schema.add_field("user_id", DataType.INT64, is_primary=True)
schema.add_field("timestamp", DataType.INT64, is_primary=True)
# ✅ CORRECT — use a single primary key field
schema.add_field("id", DataType.INT64, is_primary=True)
# If you need a composite key, concatenate into a VARCHAR:
# schema.add_field("id", DataType.VARCHAR, is_primary=True, max_length=128)
# and set id = f"{user_id}_{timestamp}" in your application code
-
Primary key uniqueness: Primary keys must be unique across the entire collection, including across partitions. Duplicate primary keys across partitions are not allowed.
-
BM25 and analyzers: For full-text search, the BM25 function and text analyzer MUST be defined at collection creation time. They CANNOT be added to an existing collection.
# ❌ WRONG — BM25 function cannot be added to an existing collection
# There is no API to add a BM25 function after creation.
# ✅ CORRECT — define BM25 at collection creation time
from pymilvus import Function, FunctionType
schema = client.create_schema()
schema.add_field("id", DataType.INT64, is_primary=True, auto_id=True)
schema.add_field("text", DataType.VARCHAR, max_length=1024,
enable_analyzer=True, analyzer_params={"type": "standard"})
schema.add_field("sparse_vector", DataType.SPARSE_FLOAT_VECTOR)
schema.add_field("dense_vector", DataType.FLOAT_VECTOR, dim=768)
bm25_function = Function(
name="text_bm25",
input_field_names=["text"],
output_field_names=["sparse_vector"],
function_type=FunctionType.BM25,
)
schema.add_function(bm25_function)
-
Nullable fields:
nullable=Trueis supported on scalar fields (including JSON and Array) and on vector fields. Exceptions: primary keys and Array of Structs fields can never be nullable. Note: vector field nullable requires Milvus v3.0.x or later (v2.6.x supports scalar fields only); vector fields that allow NULL do not supportIS NULL/IS NOT NULLfilters. -
ALWAYS use
DataType.FLOAT_VECTOR,DataType.INT64, etc. from theDataTypeenum. NEVER pass field types as strings.
Decision guide
| Decision | Use this | When |
|---|---|---|
| Primary key type | DataType.INT64 with auto_id=True |
Default choice. Let Milvus generate unique IDs. |
| Primary key type | DataType.VARCHAR |
When you need application-controlled string IDs (e.g., UUIDs, composite keys). |
| Dynamic fields | enable_dynamic_field=True |
When entities have variable or unpredictable key-value metadata. Dynamic fields are queryable but not indexed as efficiently as schema-defined fields. |
| Dynamic fields | enable_dynamic_field=False |
When your schema is well-defined and all fields are known at creation time. Better query performance. |
| Nullable fields | nullable=True |
When some entities may lack a value for a scalar or vector field. Never nullable: primary keys, Array of Structs fields. |
| Default values | default_value=... |
When you want a fallback value for missing scalar fields during insertion. |
Complete example: schema with all common field types
from pymilvus import MilvusClient, DataType
client = MilvusClient(
uri="YOUR_MILVUS_URI",
token="YOUR_MILVUS_TOKEN"
)
schema = client.create_schema(auto_id=True, enable_dynamic_field=False)
# Primary key — INT64 with auto-generated IDs
schema.add_field("id", DataType.INT64, is_primary=True)
# Vector fields
schema.add_field("dense_vector", DataType.FLOAT_VECTOR, dim=768)
schema.add_field("sparse_vector", DataType.SPARSE_FLOAT_VECTOR)
# Scalar fields
schema.add_field("title", DataType.VARCHAR, max_length=256)
schema.add_field("category", DataType.VARCHAR, max_length=64, nullable=True)
schema.add_field("price", DataType.FLOAT)
schema.add_field("tags", DataType.ARRAY, element_type=DataType.VARCHAR,
max_capacity=10, max_length=64)
schema.add_field("metadata", DataType.JSON)
# Prepare indexes
index_params = client.prepare_index_params()
index_params.add_index(field_name="dense_vector", index_type="AUTOINDEX", metric_type="COSINE")
index_params.add_index(field_name="sparse_vector", index_type="SPARSE_INVERTED_INDEX", metric_type="IP")
# Create collection
client.create_collection(
collection_name="products",
schema=schema,
index_params=index_params,
)
Complete example: schema with BM25 full-text search
BM25 and analyzers MUST be configured at creation time. This cannot be added later.
from pymilvus import MilvusClient, DataType, Function, FunctionType
client = MilvusClient(
uri="YOUR_MILVUS_URI",
token="YOUR_MILVUS_TOKEN"
)
schema = client.create_schema(auto_id=True)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field("text", DataType.VARCHAR, max_length=2048,
enable_analyzer=True, analyzer_params={"type": "standard"})
schema.add_field("sparse_vector", DataType.SPARSE_FLOAT_VECTOR)
schema.add_field("dense_vector", DataType.FLOAT_VECTOR, dim=768)
# BM25 function — MUST be added before collection creation
bm25_function = Function(
name="text_bm25",
input_field_names=["text"],
output_field_names=["sparse_vector"],
function_type=FunctionType.BM25,
)
schema.add_function(bm25_function)
index_params = client.prepare_index_params()
index_params.add_index(field_name="dense_vector", index_type="AUTOINDEX", metric_type="COSINE")
index_params.add_index(field_name="sparse_vector", index_type="SPARSE_INVERTED_INDEX", metric_type="BM25")
client.create_collection(
collection_name="documents",
schema=schema,
index_params=index_params,
)
Verification checklist
Before finishing, verify:
- All code uses
MilvusClient, not the legacy ORM API - Field types use
DataTypeenum, not strings - Primary key is
DataType.INT64orDataType.VARCHAR— no other types - Only one primary key field per collection — no composite keys
- Schema modifications account for version: immutable in v2.5.x, add nullable scalar fields in v2.6.x, drop scalar fields and non-last vector fields in v3.0.x
- BM25 function and analyzer are defined at collection creation time, not added later
- Nullable is used only on scalar fields (all versions) or vector fields (v3.0.x+) — never on primary keys or Array of Structs fields
## 핵심 요약 (결정 가이드)
필드를 설계할 때 참고할 결정 가이드예요:
| 결정 | 선택 | 언제 |
|---|---|---|
| 기본 키 타입 | `DataType.INT64` + `auto_id=True` | 기본 선택. Milvus가 고유 ID를 생성하게 해요. |
| 기본 키 타입 | `DataType.VARCHAR` | 애플리케이션이 제어하는 문자열 ID(예: UUID, 복합 키)가 필요할 때. |
| 동적 필드 | `enable_dynamic_field=True` | 엔터티에 가변적이거나 예측 불가능한 키-값 메타데이터가 있을 때. 동적 필드는 쿼리 가능하지만 스키마 필드만큼 효율적으로 인덱싱되지는 않아요. |
| 동적 필드 | `enable_dynamic_field=False` | 스키마가 잘 정의되어 생성 시점에 모든 필드를 알 수 있을 때. 쿼리 성능이 더 좋아요. |
| Nullable 필드 | `nullable=True` | 일부 엔터티가 스칼라·벡터 필드 값을 가지지 않을 수 있을 때. 절대 nullable 불가: 기본 키, Array of Structs 필드. |
| 기본값 | `default_value=...` | 삽입 시 누락된 스칼라 필드에 대한 대체값이 필요할 때. |
## 더 알아보기 (Learn more)
- [AI Prompts](/docs/milvus_for_agents.md) — 프롬프트 전체 목록과 배치 방법
- [컬렉션 만들기](/docs/create-collection.md) — 스키마·인덱스와 함께 컬렉션 생성