Array of Structs로 데이터 모델 설계하기 (Data Model Design with an Array of Structs)
Milvus 2.6.4+와 호환돼요.
현대 AI 애플리케이션, 특히 사물인터넷(IoT)과 자율주행 분야는 풍부하고 구조화된 이벤트를 바탕으로 추론하는 경우가 많아요. 예를 들어 타임스탬프와 벡터 임베딩을 가진 센서 판독값, 오류 코드와 오디오 조각을 가진 진단 로그, 위치·속도·장면 컨텍스트를 가진 주행 구간 등이 있죠. 이런 경우 데이터베이스가 중첩된 데이터(nested data)를 기본적으로 수집하고 검색할 수 있어야 해요.
사용자에게 원자적 구조 이벤트를 평면 데이터 모델로 변환하도록 요구하는 대신, Milvus는 Array of Structs를 도입했어요. 배열의 각 Struct는 스칼라와 벡터를 보유할 수 있어 의미적 무결성을 보존해요.
출처: Milvus 문서
본문
Array of Structs가 필요한 이유 (Why Array of Structs)
자율주행에서 멀티모달 검색에 이르는 현대 AI 애플리케이션은 점점 더 중첩되고 이질적인 데이터에 의존해요. 전통적인 평면 데이터 모델은 "주석이 달린 여러 청크를 가진 하나의 문서"나 "여러 관측된 기동을 가진 하나의 주행 장면" 같은 복잡한 관계를 표현하기 어려워요. 이때 Milvus의 Array of Structs 데이터 타입이 빛을 발해요.
Array of Structs를 사용하면 구조화된 요소의 정렬된 집합을 저장할 수 있어요. 각 Struct는 스칼라 필드와 벡터 임베딩의 자체 조합을 담아요. 따라서 다음에 이상적이에요.
- 계층적 데이터 (Hierarchical data): 여러 자식 레코드를 가진 부모 엔티티, 예를 들어 많은 텍스트 청크를 가진 책이나 많은 주석 프레임을 가진 비디오.
- 멀티모달 임베딩 (Multimodal embeddings): 각 Struct는 메타데이터와 함께 텍스트 임베딩과 이미지 임베딩 같은 여러 벡터를 보유할 수 있어요.
- 시간적 또는 순차적 데이터 (Temporal or sequential data): Array 필드의 Struct는 시계열이나 단계별 이벤트를 자연스럽게 표현해요.
JSON 블롭을 저장하거나 데이터를 여러 컬렉션으로 분할하는 전통적인 해결 방법과 달리, Array of Structs는 Milvus 내에서 기본적인 스키마 강제, 벡터 인덱싱, 효율적인 저장을 제공해요.
스키마 설계 가이드라인 (Schema design guidelines)
Data Model Design for Search에서 논의된 모든 가이드라인에 더해, 데이터 모델 설계에서 Array of Structs를 사용하기 전에 다음 사항도 고려해야 해요.
Struct 스키마 정의 (Define the Struct schema)
컬렉션에 Array 필드를 추가하기 전에 내부 Struct 스키마를 정의해요. struct의 각 필드는 명시적으로 타입이 지정되어야 해요. 스칼라(VARCHAR, INT, BOOLEAN 등) 또는 벡터(FLOAT_VECTOR)예요.
검색이나 표시에 사용할 필드만 포함해 Struct 스키마를 간결하게 유지하는 것이 좋아요. 사용하지 않는 메타데이터로 비대해지는 것을 피하세요.
최대 용량을 신중하게 설정 (Set the max capacity thoughtfully)
각 Array 필드에는 엔티티마다 Array 필드가 보유할 수 있는 최대 요소 수를 지정하는 속성이 있어요. 사용 사례의 상한을 기준으로 설정하세요. 예를 들어 문서당 텍스트 청크 1,000개, 주행 장면당 기동 100개 같은 식이에요.
지나치게 높은 값은 메모리를 낭비하므로, Array 필드의 최대 Struct 수를 결정하려면 몇 가지 계산을 해야 해요.
Struct의 벡터 필드 인덱싱 (Index vector fields in Structs)
인덱싱은 컬렉션의 벡터 필드와 Struct에 정의된 벡터 필드를 포함한 모든 벡터 필드에 필수예요. Struct 안의 벡터 필드에는 인덱스 타입으로 AUTOINDEX 또는 HNSW를, 메트릭 타입으로 MAX_SIM 계열을 사용해야 해요.
적용 가능한 모든 제한에 대한 자세한 내용은 the limits를 참고하세요.
실제 사례: 자율주행용 CoVLA 데이터셋 모델링 (A real-world example)
Turing Motors가 소개하고 WACV(응용 컴퓨터 비전 겨울 학회) 2025에서 채택된 CoVLA(포괄적 비전-언어-행동) 데이터셋은 자율주행 분야의 VLA(비전-언어-행동) 모델을 훈련하고 평가하기 위한 풍부한 기반을 제공해요. 보통 비디오 클립인 각 데이터 포인트는 원시 시각 입력뿐만 아니라 다음을 설명하는 구조화된 캡션도 포함해요.
- 자차(ego vehicle)의 행동 (예: "다가오는 교통에 양보하며 좌회전 병합"),
- 감지된 객체 (예: 선행 차량, 보행자, 신호등),
- 장면의 프레임 수준 캡션.
이 계층적이고 멀티모달인 성격 덕분에 Array of Structs 기능에 이상적인 후보예요. CoVLA 데이터셋에 대한 자세한 내용은 CoVLA Dataset Website를 참고하세요.
1단계: 데이터셋을 컬렉션 스키마로 매핑 (Step 1)
CoVLA 데이터셋은 총 80시간 이상의 영상을 담은 10,000개의 비디오 클립으로 구성된 대규모 멀티모달 주행 데이터셋이에요. 20Hz로 프레임을 샘플링하고 각 프레임에 상세한 자연어 캡션과 차량 상태 및 감지된 객체 좌표 정보를 주석으로 달아요.
데이터셋 구조는 다음과 같아요.
├── video_1 (VIDEO) # video.mp4
│ ├── video_id (INT)
│ ├── video_url (STRING)
│ ├── frames (ARRAY)
│ │ ├── frame_1 (STRUCT)
│ │ │ ├── caption (STRUCT) # captions.jsonl
│ │ │ │ ├── plain_caption (STRING)
│ │ │ │ ├── rich_caption (STRING)
│ │ │ │ ├── risk (STRING)
│ │ │ │ ├── risk_correct (BOOL)
│ │ │ │ ├── risk_yes_rate (FLOAT)
│ │ │ │ ├── weather (STRING)
│ │ │ │ ├── weather_rate (FLOAT)
│ │ │ │ ├── road (STRING)
│ │ │ │ ├── road_rate (FLOAT)
│ │ │ │ ├── is_tunnel (BOOL)
│ │ │ │ ├── is_tunnel_yes_rate (FLOAT)
│ │ │ │ ├── is_highway (BOOL)
│ │ │ │ ├── is_highway_yes_rate (FLOAT)
│ │ │ │ ├── has_pedestrain (BOOL)
│ │ │ │ ├── has_pedestrain_yes_rate (FLOAT)
│ │ │ │ ├── has_carrier_car (BOOL)
│ │ │ ├── traffic_light (STRUCT) # traffic_lights.jsonl
│ │ │ │ ├── index (INT)
│ │ │ │ ├── class (STRING)
│ │ │ │ ├── bbox (LIST<FLOAT>)
│ │ │ ├── front_car (STRUCT) # front_cars.jsonl
│ │ │ │ ├── has_lead (BOOL)
│ │ │ │ ├── lead_prob (FLOAT)
│ │ │ │ ├── lead_x (FLOAT)
│ │ │ │ ├── lead_y (FLOAT)
│ │ │ │ ├── lead_speed_kmh (FLOAT)
│ │ │ │ ├── lead_a (FLOAT)
│ │ ├── frame_2 (STRUCT)
│ │ ├── ... (STRUCT)
│ │ ├── frame_n (STRUCT)
├── video_2
├── ...
├── video_n
CoVLA 데이터셋의 구조가 매우 계층적이며 수집된 데이터를 여러 .jsonl 파일과 .mp4 형식의 비디오 클립으로 나눈다는 것을 알 수 있어요.
Milvus에서는 JSON 필드나 Array-of-Structs 필드를 사용해 컬렉션 스키마 내에 중첩 구조를 만들 수 있어요. 중첩 형식의 일부로 벡터 임베딩이 있을 때는 Array-of-Structs 필드만 지원돼요. 그러나 Array 안의 Struct는 그 자체로 추가 중첩 구조를 포함할 수 없어요. 필수 관계를 유지하면서 CoVLA 데이터셋을 저장하려면 불필요한 계층을 제거하고 데이터를 평탄화해 Milvus 컬렉션 스키마에 맞춰야 해요.
다음은 이 데이터셋을 스키마로 모델링하는 방법을 보여 주는 다이어그램이에요.
비디오 클립의 구조는 다음 필드로 구성돼요.
video_id는 INT64 타입의 정수를 받는 기본 키 역할을 해요.states는 현재 비디오의 각 프레임에서 자차의 상태를 담은 원시 JSON 본문이에요.captions는 Array of Structs이며 각 Struct는 다음 필드를 가져요.
frame_id는 현재 비디오 내 특정 프레임을 식별해요.
-
plain_caption은 날씨, 도로 상태 등의 주변 환경을 제외한 현재 프레임의 설명이고,plain_cap_vector는 그에 상응하는 벡터 임베딩이에요. -
rich_caption은 주변 환경을 포함한 현재 프레임의 설명이고,rich_cap_vector는 그에 상응하는 벡터 임베딩이에요. -
risk는 현재 프레임에서 자차가 직면한 위험에 대한 설명이고,risk_vector는 그에 상응하는 벡터 임베딩이에요. -
그 외 프레임의 다른 속성(예:
road,weather,is_tunnel,has_pedestrain등)이 있어요. -
traffic_lights는 현재 프레임에서 식별된 모든 신호등 신호를 담은 JSON 본문이에요. -
front_cars는 현재 프레임에서 식별된 모든 선행 차량을 담은 Array of Structs이기도 해요.
2단계: 스키마 초기화 (Step 2)
시작하려면 caption Struct, front_cars Struct, 컬렉션의 스키마를 초기화해야 해요.
Caption Struct의 스키마를 초기화해요.
client = MilvusClient("http://localhost:19530")
# create the schema for the caption struct
schema_for_caption = client.create_struct_field_schema()
schema_for_caption.add_field(
field_name="frame_id",
datatype=DataType.INT64,
description="ID of the frame to which the ego vehicle's behavior belongs"
)
schema_for_caption.add_field(
field_name="plain_caption",
datatype=DataType.VARCHAR,
max_length=1024,
description="plain description of the ego vehicle's behaviors"
)
schema_for_caption.add_field(
field_name="plain_cap_vector",
datatype=DataType.FLOAT_VECTOR,
dim=768,
description="vectors for the plain description of the ego vehicle's behaviors"
)
schema_for_caption.add_field(
field_name="rich_caption",
datatype=DataType.VARCHAR,
max_length=1024,
description="rich description of the ego vehicle's behaviors"
)
schema_for_caption.add_field(
field_name="rich_cap_vector",
datatype=DataType.FLOAT_VECTOR,
dim=768,
description="vectors for the rich description of the ego vehicle's behaviors"
)
schema_for_caption.add_field(
field_name="risk",
datatype=DataType.VARCHAR,
max_length=1024,
description="description of the ego vehicle's risks"
)
schema_for_caption.add_field(
field_name="risk_vector",
datatype=DataType.FLOAT_VECTOR,
dim=768,
description="vectors for the description of the ego vehicle's risks"
)
schema_for_caption.add_field(
field_name="risk_correct",
datatype=DataType.BOOL,
description="whether the risk assessment is correct"
)
schema_for_caption.add_field(
field_name="risk_yes_rate",
datatype=DataType.FLOAT,
description="probability/confidence of risk being present"
)
schema_for_caption.add_field(
field_name="weather",
datatype=DataType.VARCHAR,
max_length=50,
description="weather condition"
)
schema_for_caption.add_field(
field_name="weather_rate",
datatype=DataType.FLOAT,
description="probability/confidence of the weather condition"
)
schema_for_caption.add_field(
field_name="road",
datatype=DataType.VARCHAR,
max_length=50,
description="road type"
)
schema_for_caption.add_field(
field_name="road_rate",
datatype=DataType.FLOAT,
description="probability/confidence of the road type"
)
schema_for_caption.add_field(
field_name="is_tunnel",
datatype=DataType.BOOL,
description="whether the road is a tunnel"
)
schema_for_caption.add_field(
field_name="is_tunnel_yes_rate",
datatype=DataType.FLOAT,
description="probability/confidence of the road being a tunnel"
)
schema_for_caption.add_field(
field_name="is_highway",
datatype=DataType.BOOL,
description="whether the road is a highway"
)
schema_for_caption.add_field(
field_name="is_highway_yes_rate",
datatype=DataType.FLOAT,
description="probability/confidence of the road being a highway"
)
schema_for_caption.add_field(
field_name="has_pedestrian",
datatype=DataType.BOOL,
description="whether there is a pedestrian present"
)
schema_for_caption.add_field(
field_name="has_pedestrian_yes_rate",
datatype=DataType.FLOAT,
description="probability/confidence of pedestrian presence"
)
schema_for_caption.add_field(
field_name="has_carrier_car",
datatype=DataType.BOOL,
description="whether there is a carrier car present"
)
Front Car Struct의 스키마를 초기화해요.
전방 차량은 벡터 임베딩을 포함하지 않지만, 데이터 크기가 JSON 필드의 최대값을 초과하므로 여전히 Array of Structs로 포함해야 해요.
schema_for_front_car = client.create_struct_field_schema()
schema_for_front_car.add_field(
field_name="frame_id",
datatype=DataType.INT64,
description="ID of the frame to which the ego vehicle's behavior belongs"
)
schema_for_front_car.add_field(
field_name="has_lead",
datatype=DataType.BOOL,
description="whether there is a leading vehicle"
)
schema_for_front_car.add_field(
field_name="lead_prob",
datatype=DataType.FLOAT,
description="probability/confidence of the leading vehicle's presence"
)
schema_for_front_car.add_field(
field_name="lead_x",
datatype=DataType.FLOAT,
description="x position of the leading vehicle relative to the ego vehicle"
)
schema_for_front_car.add_field(
field_name="lead_y",
datatype=DataType.FLOAT,
description="y position of the leading vehicle relative to the ego vehicle"
)
schema_for_front_car.add_field(
field_name="lead_speed_kmh",
datatype=DataType.FLOAT,
description="speed of the leading vehicle in km/h"
)
schema_for_front_car.add_field(
field_name="lead_a",
datatype=DataType.FLOAT,
description="acceleration of the leading vehicle"
)
컬렉션의 스키마를 초기화해요.
schema = client.create_schema()
schema.add_field(
field_name="video_id",
datatype=DataType.VARCHAR,
description="primary key",
max_length=16,
is_primary=True,
auto_id=False
)
schema.add_field(
field_name="video_url",
datatype=DataType.VARCHAR,
max_length=512,
description="URL of the video"
)
schema.add_field(
field_name="captions",
datatype=DataType.ARRAY,
element_type=DataType.STRUCT,
struct_schema=schema_for_caption,
max_capacity=600,
description="captions for the current video"
)
schema.add_field(
field_name="traffic_lights",
datatype=DataType.JSON,
description="frame-specific traffic lights identified in the current video"
)
schema.add_field(
field_name="front_cars",
datatype=DataType.ARRAY,
element_type=DataType.STRUCT,
struct_schema=schema_for_front_car,
max_capacity=600,
description="frame-specific leading cars identified in the current video"
)
3단계: 인덱스 파라미터 설정 (Step 3)
모든 벡터 필드는 인덱싱되어야 해요. 요소 Struct의 벡터 필드를 인덱싱하려면 인덱스 타입으로 AUTOINDEX 또는 HNSW를, 메트릭 타입으로 MAX_SIM 계열을 사용해 임베딩 목록 간의 유사도를 측정해야 해요.
index_params = client.prepare_index_params()
index_params.add_index(
field_name="captions[plain_cap_vector]",
index_type="AUTOINDEX",
metric_type="MAX_SIM_COSINE",
index_name="captions_plain_cap_vector_idx", # mandatory for now
index_params={"M": 16, "efConstruction": 200}
)
index_params.add_index(
field_name="captions[rich_cap_vector]",
index_type="AUTOINDEX",
metric_type="MAX_SIM_COSINE",
index_name="captions_rich_cap_vector_idx", # mandatory for now
index_params={"M": 16, "efConstruction": 200}
)
index_params.add_index(
field_name="captions[risk_vector]",
index_type="AUTOINDEX",
metric_type="MAX_SIM_COSINE",
index_name="captions_risk_vector_idx", # mandatory for now
index_params={"M": 16, "efConstruction": 200}
)
이러한 필드 내 필터링을 가속화하려면 JSON 필드에 대해 JSON shredding을 활성화하는 것이 좋아요.
4단계: 컬렉션 생성 (Step 4)
스키마와 인덱스가 준비되면 다음과 같이 대상 컬렉션을 만들 수 있어요.
client.create_collection(
collection_name="covla_dataset",
schema=schema,
index_params=index_params
)
5단계: 데이터 삽입 (Step 5)
Turing Motors는 CoVLA 데이터셋을 원시 비디오 클립(.mp4), 상태(states.jsonl), 캡션(captions.jsonl), 신호등(traffic_lights.jsonl), 전방 차량(front_cars.jsonl) 등 여러 파일로 구성해요.
각 비디오 클립의 데이터 조각을 이 파일들에서 병합한 뒤 데이터를 삽입해야 해요. 다음은 특정 비디오 클립의 데이터 조각을 병합하는 스크립트예요.
import json
from openai import OpenAI
openai_client = OpenAI(
api_key='YOUR_OPENAI_API_KEY',
)
video_id = "0a0fc7a5db365174" # represent a single video with 600 frames
# get all front car records in the specified video clip
entries = []
front_cars = []
with open('data/front_car/{}.jsonl'.format(video_id), 'r') as f:
for line in f:
entries.append(json.loads(line))
for entry in entries:
for key, value in entry.items():
value['frame_id'] = int(key)
front_cars.append(value)
# get all traffic lights identified in the specified video clip
entries = []
traffic_lights = []
frame_id = 0
with open('data/traffic_lights/{}.jsonl'.format(video_id), 'r') as f:
for line in f:
entries.append(json.loads(line))
for entry in entries:
for key, value in entry.items():
if not value or (value['index'] == 1 and key != '0'):
frame_id+=1
if value:
value['frame_id'] = frame_id
traffic_lights.append(value)
else:
value_dict = {}
value_dict['frame_id'] = frame_id
traffic_lights.append(value_dict)
# get all captions generated in the video clip and convert them into vector embeddings
entries = []
captions = []
with open('data/captions/{}.jsonl'.format(video_id), 'r') as f:
for line in f:
entries.append(json.loads(line))
def get_embedding(text, model="embeddinggemma:latest"):
response = openai_client.embeddings.create(input=text, model=model)
return response.data[0].embedding
# Add embeddings to each entry
for entry in entries:
# Each entry is a dict with a single key (e.g., '0', '1', ...)
for key, value in entry.items():
value['frame_id'] = int(key) # Convert key to integer and assign to frame_id
if "plain_caption" in value and value["plain_caption"]:
value["plain_cap_vector"] = get_embedding(value["plain_caption"])
if "rich_caption" in value and value["rich_caption"]:
value["rich_cap_vector"] = get_embedding(value["rich_caption"])
if "risk" in value and value["risk"]:
value["risk_vector"] = get_embedding(value["risk"])
captions.append(value)
data = {
"video_id": video_id,
"video_url": "https://your-storage.com/{}".format(video_id),
"captions": captions,
"traffic_lights": traffic_lights,
"front_cars": front_cars
}
데이터를 적절히 처리했다면 다음과 같이 삽입할 수 있어요.
client.insert(
collection_name="covla_dataset",
data=[data]
)
# {'insert_count': 1, 'ids': ['0a0fc7a5db365174'], 'cost': 0}
더 알아보기 (Learn more)
- Data Model Design for Search — 스키마 설계 가이드
- the limits — Array of Structs 제한 사항
- CoVLA Dataset Website — CoVLA 데이터셋
- Milvus 공식 문서 — 스키마·벡터 검색 관련 자료