본문 바로가기
WIKI 기술 지식 베이스

Nullable 필드 (Nullable Fields)

원문 보기 위키 갱신

Milvus는 nullable 필드를 지원해요. nullable 필드는 필드 값이 누락되거나 명시적으로 NULL로 설정되는 것을 허용해요. Nullability는 스키마 수준에서 정의되며 데이터 수집, 인덱싱, 검색, 쿼리 작업 전반에 걸쳐 일관되게 적용돼요.

Nullable 필드는 다음과 같은 경우에 사용해요.

  • 외부 시스템에서 누락 값을 허용하는 데이터를 수집할 때.
  • 일부 메타데이터가 선택적이거나 데이터셋의 일부에서만 사용 가능할 때.
  • 벡터 임베딩이 비동기적으로 생성되어 나중에 삽입될 때.

출처: Milvus 문서

본문

제한 사항 (Limits)

  • NULL 값을 허용하는 벡터 필드는 IS NULL 또는 IS NOT NULL 필터 표현식을 지원하지 않아요. 벡터 필드 값이 NULL인지에 따라 엔티티를 명시적으로 필터링할 수 없어요.
  • Milvus 3.0.0부터 부모 StructArray 필드는 nullable일 수 있어요. 부모 StructArray 필드에 nullable=True를 설정하고 개별 서브필드에는 설정하지 마세요. NULL은 개별 Struct 요소가 아니라 전체 StructArray 필드에 적용되며, Milvus는 부모의 nullability를 내부적으로 서브필드에 전파해요. 기존 컬렉션에 추가된 StructArray 필드는 기존 엔티티가 새 필드에 대해 NULL을 반환할 수 있도록 nullable이어야 해요. 자세한 내용은 StructArray Limits를 참고하세요.
  • nullable 속성은 필드가 생성될 때 정의되며 나중에 수정할 수 없어요. 기존 필드에 대해 nullability를 활성화하거나 비활성화할 수 없어요.
  • nullable로 표시된 필드는 파티션 키로 사용할 수 없어요. 파티션 키 필드는 항상 유효하고 널이 아닌 값을 포함해야 해요. 자세한 내용은 Use Partition Key를 참고하세요.

nullable 필드란? (What is a nullable field?)

Milvus에서 필드가 NULL 값을 저장하도록 허용할지 여부는 nullable이라는 스키마 수준 필드 속성으로 제어돼요.

필드가 nullable=True로 정의되면 Milvus는 데이터 수집 중 필드 값이 누락되는 것을 허용해요. 실제로 Milvus는 다음 두 입력을 동등하게 취급하고 필드 값을 NULL로 저장해요.

  • 필드가 입력 엔티티에서 생략된 경우.
  • 필드가 명시적으로 NULL로 설정된 경우(예: Python의 None).

필드가 nullable로 정의되지 않으면(기본 동작) 모든 엔티티는 해당 필드에 유효한 값을 제공해야 해요. 필드를 생략하거나 NULL 값을 명시적으로 할당하면 삽입 또는 가져오기 작업이 실패해요.

nullable 속성은 컬렉션 스키마의 스칼라 필드와 벡터 필드 모두에서 지원돼요. Milvus 3.0.0부터 부모 StructArray 필드에서도 지원돼요. Struct 서브필드를 독립적으로 nullable로 구성하지 마세요. StructArray 부모에 nullability를 정의하면 Milvus가 그 설정을 내부적으로 서브필드에 전파해요.

Nullability는 필드 값이 누락될 수 있는지 여부를 결정할 뿐, 필드가 누락될 때 어떤 값이 사용되는지 정의하지는 않아요.

  • nullable 필드가 기본값 없이 구성된 경우 필드를 생략하면 NULL 값이 저장돼요.
  • 기본값이 구성된 경우 Milvus는 대신 기본값을 저장할 수 있어요. 자세한 내용은 Default Values를 참고하세요.

컬렉션 스키마에서 nullable 필드 정의 (Define a nullable field in the collection schema)

Nullable 필드를 사용하려면 컬렉션 스키마를 정의할 때 nullable 속성을 활성화해야 해요.

이 예시에서 컬렉션 스키마는 nullable=True인 embedding이라는 벡터 필드를 정의해요. 이렇게 하면 컬렉션의 엔티티가 데이터 수집 중 벡터 값을 생략하거나 명시적으로 NULL로 설정할 수 있어요.

from pymilvus import MilvusClient, DataType

client = MilvusClient(
    uri="http://localhost:19530",
    token="root:Milvus"
)

# Define schema fields
schema = client.create_schema()
schema.add_field("id", DataType.INT64, is_primary=True)  # Primary field
schema.add_field(
    field_name="embedding",
    datatype=DataType.FLOAT_VECTOR,
    dim=4,
    nullable=True,  # Enable the nullable attribute; defaults to False
)

client.create_collection(
    collection_name="my_collection",
    schema=schema,
)
import io.milvus.v2.client.ConnectConfig;
import io.milvus.v2.client.MilvusClientV2;
import io.milvus.v2.common.DataType;
import io.milvus.v2.service.collection.request.AddFieldReq;
import io.milvus.v2.service.collection.request.CreateCollectionReq;

MilvusClientV2 client = new MilvusClientV2(ConnectConfig.builder()
        .uri("http://localhost:19530")
        .token("root:Milvus")
        .build());

CreateCollectionReq.CollectionSchema schema = CreateCollectionReq.CollectionSchema.builder()
        .build();

schema.addField(AddFieldReq.builder()
        .fieldName("id")
        .dataType(DataType.Int64)
        .isPrimaryKey(true)
        .build());
schema.addField(AddFieldReq.builder()
        .fieldName("embedding")
        .dataType(DataType.FloatVector)
        .dimension(4)
        .isNullable(true)
        .build());

client.createCollection(CreateCollectionReq.builder()
        .collectionName("my_collection")
        .collectionSchema(schema)
        .build());
import { MilvusClient, DataType } from "@zilliz/milvus2-sdk-node";

const client = new MilvusClient({
  address: "http://localhost:19530",
  token: "root:Milvus",
});

await client.createCollection({
  collection_name: "my_collection",
  fields: [
    {
      name: "id",
      data_type: DataType.Int64,
      is_primary_key: true,
      autoID: false,
    },
    {
      name: "embedding",
      data_type: DataType.FloatVector,
      dim: 4,
      nullable: true,
    },
  ],
});
import (
    "context"
    "fmt"

    "github.com/milvus-io/milvus/client/v2/entity"
    "github.com/milvus-io/milvus/client/v2/milvusclient"
)

ctx, cancel := context.WithCancel(context.Background())
defer cancel()

client, err := milvusclient.New(ctx, &milvusclient.ClientConfig{
    Address: "localhost:19530",
})
if err != nil {
    fmt.Println(err.Error())
    // handle error
}
defer client.Close(ctx)

schema := entity.NewSchema()
schema.WithField(entity.NewField().
    WithName("id").
    WithDataType(entity.FieldTypeInt64).
    WithIsPrimaryKey(true),
).WithField(entity.NewField().
    WithName("embedding").
    WithDataType(entity.FieldTypeFloatVector).
    WithDim(4).
    WithNullable(true),
)

err = client.CreateCollection(ctx,
    milvusclient.NewCreateCollectionOption("my_collection", schema))
if err != nil {
    fmt.Println(err.Error())
    // handle error
}
export TOKEN="root:Milvus"
export CLUSTER_ENDPOINT="http://localhost:19530"

export pkField='{
  "fieldName": "id",
  "dataType": "Int64",
  "isPrimary": true
}'

export embeddingField='{
  "fieldName": "embedding",
  "dataType": "FloatVector",
  "typeParams": {"dim": "4"},
  "nullable": true
}'

curl --request POST \
  --url "${CLUSTER_ENDPOINT}/v2/vectordb/collections/create" \
  --header "Authorization: Bearer ${TOKEN}" \
  --header "Content-Type: application/json" \
  --header "Request-Timeout: 10" \
  -d "{
    \"collectionName\": \"my_collection\",
    \"schema\": {
      \"fields\": [
        $pkField,
        $embeddingField
      ]
    }
  }"

이 스키마에서:

  • embedding 필드가 명시적으로 nullable로 표시되어 있어요.
  • 엔티티는 삽입 중 embedding 필드를 생략하거나 NULL 값을 할당할 수 있어요.
  • NULL 값을 허용할지 여부는 컬렉션 생성 시점에 고정돼요.

명확성을 위해 다음 예시들은 nullable 벡터 필드(embedding)에 초점을 맞춰요. nullable 스칼라 필드를 정의하는 것은 선택 사항이며 이 가이드의 나머지 부분을 따르는 데 필요하지 않아요.

선택 사항: nullable 스칼라 필드 정의 (Optional: Define a nullable scalar field)

스칼라 필드도 동일한 nullable 속성을 사용해 nullable로 정의할 수 있으며 수집 중 동일한 규칙을 따르요. 예를 들어:

schema.add_field(
    field_name="age",
    datatype=DataType.INT64,
    nullable=True,
)
schema.addField(AddFieldReq.builder()
        .fieldName("age")
        .dataType(DataType.Int64)
        .isNullable(true)
        .build());
// Add to the fields array when calling createCollection:
// { name: "age", data_type: DataType.Int64, nullable: true },
schema.WithField(entity.NewField().
    WithName("age").
    WithDataType(entity.FieldTypeInt64).
    WithNullable(true),
)
# Add another field object to the schema "fields" array, for example:
# { "fieldName": "age", "dataType": "Int64", "nullable": true }

누락 또는 NULL 값에 대한 삽입 동작 (Insert behavior with missing or NULL values)

필드가 컬렉션 스키마에서 nullable로 정의되면 Milvus는 데이터 수집 중 필드 값을 생략하거나 명시적으로 NULL로 설정하는 것을 허용해요.

아래 예시는 컬렉션 스키마에서 nullable 필드 정의에서 만든 컬렉션에 세 엔티티를 삽입하면서 서로 다른 경우를 보여 줘요.

data = [
    {
        "id": 1,
        "embedding": [0.1, 0.2, 0.3, 0.4],
    },
    {
        "id": 2,
        "embedding": None,  # Explicitly set to NULL
    },
    {
        "id": 3,  # Field omitted → stored as NULL
    },
]

client.insert(
    collection_name="my_collection",
    data=data,
)
import com.google.gson.Gson;
import com.google.gson.JsonNull;
import com.google.gson.JsonObject;
import io.milvus.v2.service.vector.request.InsertReq;

import java.util.Arrays;
import java.util.List;

Gson gson = new Gson();

JsonObject row1 = new JsonObject();
row1.addProperty("id", 1);
row1.add("embedding", gson.toJsonTree(Arrays.asList(0.1f, 0.2f, 0.3f, 0.4f)));

JsonObject row2 = new JsonObject();
row2.addProperty("id", 2);
row2.add("embedding", JsonNull.INSTANCE); // Explicitly set to NULL

JsonObject row3 = new JsonObject();
row3.addProperty("id", 3); // Field omitted; stored as NULL

List<JsonObject> data = Arrays.asList(row1, row2, row3);

client.insert(InsertReq.builder()
        .collectionName("my_collection")
        .data(data)
        .build());
const data = [
  { id: 1, embedding: [0.1, 0.2, 0.3, 0.4] },
  { id: 2, embedding: null },
  { id: 3 },
];

await client.insert({
  collection_name: "my_collection",
  data: data,
});
import (
    "context"
    "fmt"

    "github.com/milvus-io/milvus/client/v2/milvusclient"
)

// Assumes `client` is the Milvus client from the Go schema example above.
ctx := context.Background()

rows := []any{
    map[string]any{"id": int64(1), "embedding": []float32{0.1, 0.2, 0.3, 0.4}},
    map[string]any{"id": int64(2), "embedding": nil},
    map[string]any{"id": int64(3)},
}

_, err := client.Insert(ctx, milvusclient.NewRowBasedInsertOption("my_collection", rows...))
if err != nil {
    fmt.Println(err.Error())
}
curl --request POST \
  --url "${CLUSTER_ENDPOINT}/v2/vectordb/entities/insert" \
  --header "Authorization: Bearer ${TOKEN}" \
  --header "Content-Type: application/json" \
  --header "Request-Timeout: 10" \
  -d '{
    "collectionName": "my_collection",
    "data": [
      {"id": 1, "embedding": [0.1, 0.2, 0.3, 0.4]},
      {"id": 2, "embedding": null},
      {"id": 3}
    ]
  }'

이 예시에서:

  • 엔티티 id = 1은 유효한 벡터 값을 제공해요.
  • 엔티티 id = 2는 embedding 필드에 NULL 값을 명시적으로 할당해요.
  • 엔티티 id = 3은 embedding 필드를 완전히 생략하며, Milvus는 이를 NULL로 저장해요.

nullable 필드의 인덱스 동작 (Index behavior on nullable fields)

데이터를 삽입한 후 평소처럼 nullable 필드에 인덱스를 구축할 수 있어요. 핵심 차이는 인덱스 구축 중 Milvus가 NULL 값을 처리하는 방식이에요.

  • 널이 아닌 값을 가진 엔티티만 인덱스에 추가돼요.
  • NULL 값을 가진 엔티티는 건너뛰며 인덱스 구축에 참여하지 않아요.

nullable 벡터 필드의 경우 이는 유효한 벡터를 가진 엔티티만 벡터 유사도 검색이 가능해진다는 뜻이에요.

# Set index parameters
index_params = client.prepare_index_params()
index_params.add_index(
    field_name="embedding",
    index_type="AUTOINDEX",
    metric_type="COSINE",
)

# Create index
client.create_index(
    collection_name="my_collection",
    index_params=index_params,
)

# Load collection for future search operations
client.load_collection(collection_name="my_collection")
import io.milvus.v2.common.IndexParam;
import io.milvus.v2.service.collection.request.LoadCollectionReq;
import io.milvus.v2.service.index.request.CreateIndexReq;

import java.util.Collections;

IndexParam indexParam = IndexParam.builder()
        .fieldName("embedding")
        .indexName("embedding_index")
        .indexType(IndexParam.IndexType.AUTOINDEX)
        .metricType(IndexParam.MetricType.COSINE)
        .build();

client.createIndex(CreateIndexReq.builder()
        .collectionName("my_collection")
        .indexParams(Collections.singletonList(indexParam))
        .build());

client.loadCollection(LoadCollectionReq.builder()
        .collectionName("my_collection")
        .build());
await client.createIndex({
  collection_name: "my_collection",
  field_name: "embedding",
  index_name: "embedding_idx",
  index_type: "AUTOINDEX",
  metric_type: "COSINE",
});

await client.loadCollection({
  collection_name: "my_collection",
});
import (
    "context"
    "fmt"

    "github.com/milvus-io/milvus/client/v2/entity"
    "github.com/milvus-io/milvus/client/v2/index"
    "github.com/milvus-io/milvus/client/v2/milvusclient"
)

// Assumes `client` is the Milvus client from the Go schema example above.
ctx := context.Background()

indexOption := milvusclient.NewCreateIndexOption("my_collection", "embedding",
    index.NewAutoIndex(entity.COSINE))

_, err := client.CreateIndex(ctx, indexOption)
if err != nil {
    fmt.Println(err.Error())
}

_, err = client.LoadCollection(ctx, milvusclient.NewLoadCollectionOption("my_collection"))
if err != nil {
    fmt.Println(err.Error())
}
curl --request POST \
  --url "${CLUSTER_ENDPOINT}/v2/vectordb/indexes/create" \
  --header "Authorization: Bearer ${TOKEN}" \
  --header "Content-Type: application/json" \
  --header "Request-Timeout: 10" \
  -d '{
    "collectionName": "my_collection",
    "indexParams": [
      {
        "fieldName": "embedding",
        "metricType": "COSINE",
        "indexType": "AUTOINDEX"
      }
    ]
  }'

curl --request POST \
  --url "${CLUSTER_ENDPOINT}/v2/vectordb/collections/load" \
  --header "Authorization: Bearer ${TOKEN}" \
  --header "Content-Type: application/json" \
  --header "Request-Timeout: 10" \
  -d '{"collectionName": "my_collection"}'

이 시점에서:

  • 유효한 embedding 값을 가진 엔티티는 인덱싱되어 검색할 준비가 돼요.
  • embedding이 NULL인 엔티티는 컬렉션에 남아 있지만 벡터 인덱스에는 포함되지 않아요.

nullable 필드의 검색 동작 (Search behavior with nullable fields)

nullable 필드에서 검색 연산을 수행하면 Milvus는 검색에 사용된 필드에 널이 아닌 값을 가진 엔티티만 평가해요. 벡터 필드가 NULL인 엔티티는 자동으로 건너뛰어져요.

이 예시의 embedding 같은 nullable 벡터 필드의 경우:

  • 유효한 벡터 값을 가진 엔티티만 평가되고 순위가 매겨져요.
  • NULL 벡터를 가진 엔티티는 오류를 발생시키지 않아요.
  • 유효한 벡터 수가 요청된 topK(limit)보다 작으면 Milvus는 limit보다 적은 결과를 반환할 수 있어요.

다음 예시는 nullable 벡터 필드 embedding에 대해 벡터 검색을 수행해요.

res = client.search(
    collection_name="my_collection",
    data=[[0.1, 0.2, 0.3, 0.4]],
    anns_field="embedding",
    limit=3,
    search_params={"metric_type": "COSINE"},
    output_fields=["embedding"],
)

print(res)
import io.milvus.v2.service.vector.request.SearchReq;
import io.milvus.v2.service.vector.request.data.FloatVec;
import io.milvus.v2.service.vector.response.SearchResp;

import java.util.Arrays;
import java.util.Collections;

SearchResp res = client.search(SearchReq.builder()
        .collectionName("my_collection")
        .data(Collections.singletonList(new FloatVec(Arrays.asList(0.1f, 0.2f, 0.3f, 0.4f))))
        .annsField("embedding")
        .limit(3)
        .outputFields(Collections.singletonList("embedding"))
        .build());

System.out.println(res);
const res = await client.search({
  collection_name: "my_collection",
  data: [[0.1, 0.2, 0.3, 0.4]],
  anns_field: "embedding",
  limit: 3,
  search_params: { metric_type: "COSINE" },
  output_fields: ["embedding"],
});

console.log(res);
import (
    "context"
    "fmt"

    "github.com/milvus-io/milvus/client/v2/entity"
    "github.com/milvus-io/milvus/client/v2/milvusclient"
)

// Assumes `client` is the Milvus client from the Go schema example above.
ctx := context.Background()

query := []float32{0.1, 0.2, 0.3, 0.4}
resultSets, err := client.Search(ctx, milvusclient.NewSearchOption(
    "my_collection",
    3,
    []entity.Vector{entity.FloatVector(query)},
).WithANNSField("embedding").
    WithOutputFields("embedding"))
if err != nil {
    fmt.Println(err.Error())
}
fmt.Println(resultSets)
curl --request POST \
  --url "${CLUSTER_ENDPOINT}/v2/vectordb/entities/search" \
  --header "Authorization: Bearer ${TOKEN}" \
  --header "Content-Type: application/json" \
  --header "Request-Timeout: 10" \
  -d '{
    "collectionName": "my_collection",
    "data": [[0.1, 0.2, 0.3, 0.4]],
    "annsField": "embedding",
    "limit": 3,
    "searchParams": {"metricType": "COSINE"},
    "outputFields": ["embedding"]
  }'

이 검색에서:

  • 널이 아닌 embedding 값을 가진 엔티티만 후보로 간주돼요.
  • embedding 값이 NULL인 엔티티는 평가에서 제외돼요.
  • 반환되는 결과 수는 컬렉션에 유효한 벡터가 몇 개 존재하는지에 따라 달라져요.

쿼리 및 필터링 함의 (Query and filtering implications)

이전 예시들은 벡터 필드에 초점을 맞췄어요. 이 섹션에서는 스칼라 필터 표현식에서 NULL 값이 어떻게 동작하는지 설명해요.

스칼라 필드는 nullable=True로 정의할 수 있으며 벡터 필드와 같은 수집 규칙을 따르요. 그러나 필터 표현식에서 NULL 스칼라 값은 항상 false로 평가돼요.

예를 들어 nullable 스칼라 필드 age가 주어지면 다음 필터는 age가 18보다 큰 엔티티를 선택해요.

expr = "age > 18"
String filter = "age > 18";
const expr = "age > 18";
filter := "age > 18"
# Use in query/search filter parameter, for example:
# "filter": "age > 18"

age가 NULL인 엔티티는 NULL 값이 필터 조건을 충족하지 않으므로 결과에서 제외돼요.

마찬가지로 동등 검사는 NULL 값과 일치하지 않아요. 예를 들어:

expr = 'status == "active"'
String filter = "status == \"active\"";
const expr = 'status == "active"';
filter := `status == "active"`
# "filter": "status == \"active\""

status가 NULL인 엔티티는 결과에서 제외돼요.

nullable 필드와 기본값 (Nullable fields and default values)

필드에 nullable과 default_value가 모두 구성된 경우, 삽입 중 NULL 입력이나 누락된 필드 값을 Milvus가 어떻게 처리하는지는 다음 규칙에 따라 결정돼요.

Nullable 활성화 기본값 사용자 입력 (NULL 또는 생략) 결과
예 예 (널 아님) NULL 또는 생략 기본값 사용
예 아니요 NULL 또는 생략 NULL로 저장
아니요 예 (널 아님) NULL 또는 생략 기본값 사용
아니요 아니요 NULL 또는 생략 오류 발생
아니요 예 (NULL 기본값) NULL 또는 생략 오류 발생

핵심 요점:

  • 필드에 널이 아닌 기본값이 있으면 nullable이 활성화되어 있는지와 관계없이 그 값이 사용돼요.
  • nullable=True인데 기본값이 설정되지 않았으면 필드가 NULL을 저장해요.
  • nullable=False이고 기본값이 설정되지 않았으면 삽입이 오류로 실패해요.
  • nullable이 아닌 필드에 NULL 기본값을 설정하는 것은 유효하지 않으며 오류를 발생시켜요.

기본값에 대한 전체 예시와 API 사용법은 Default Values를 참고하세요.

더 알아보기 (Learn more)