고급 비디오 검색: Twelve Labs와 Milvus로 시맨틱 검색 구현하기 (Advanced Video Search: Leveraging Twelve Labs and Milvus for Semantic Retrieval)
소개 (Introduction)
Twelve Labs Embed API와 Milvus를 사용해 시맨틱 비디오 검색을 구현하는 이 포괄적인 튜토리얼에 오신 것을 환영해요. 이 가이드에서는 Twelve Labs의 고급 멀티모달 임베딩과 Milvus의 효율적인 벡터 데이터베이스의 힘을 활용해 강력한 비디오 검색 솔루션을 만드는 방법을 탐구해요. 이 기술들을 통합하면 개발자는 비디오 콘텐츠 분석에서 새로운 가능성을 열 수 있어요. 콘텐츠 기반 비디오 검색, 추천 시스템, 비디오 데이터의 미묘한 차이를 이해하는 정교한 검색 엔진 같은 애플리케이션을 만들 수 있죠.
이 튜토리얼은 개발 환경 설정부터 기능적인 시맨틱 비디오 검색 애플리케이션 구현까지 전체 과정을 안내해요. 비디오에서 멀티모달 임베딩 생성, Milvus에 효율적으로 저장, 유사도 검색 수행으로 관련 콘텐츠 검색 같은 핵심 개념을 다룰 거예요. 비디오 분석 플랫폼을 만들든, 콘텐츠 발견 도구를 만들든, 기존 애플리케이션에 비디오 검색 기능을 추가하든, 이 가이드는 Twelve Labs와 Milvus의 결합된 강점을 활용하는 지식과 실용적인 단계를 제공해요.
출처: Milvus 문서
본문
사전 요구 사항 (Prerequisites)
시작하기 전에 다음 항목이 있는지 확인해요.
- Twelve Labs API 키 (없다면 https://api.twelvelabs.io에서 가입).
- 시스템에 설치된 Python 3.7 이상.
개발 환경 설정 (Setting Up the Development Environment)
프로젝트용 새 디렉토리를 만들고 이동해요.
mkdir video-search-tutorial
cd video-search-tutorial
가상 환경을 설정해요(선택 사항이지만 권장).
python -m venv venv
source venv/bin/activate # On Windows, use `venv\Scripts\activate`
필요한 Python 라이브러리를 설치해요.
pip install twelvelabs pymilvus
프로젝트용 새 Python 파일을 만들어요.
touch video_search.py
이 video_search.py 파일이 이 튜토리얼에서 사용할 주요 스크립트예요. 다음으로 보안을 위해 Twelve Labs API 키를 환경 변수로 설정해요.
export TWELVE_LABS_API_KEY='your_api_key_here'
Milvus에 연결하기 (Connecting to Milvus)
Milvus와 연결하려면 MilvusClient 클래스를 사용해요. 이 접근 방식은 연결 과정을 단순화하며, 튜토리얼에 완벽한 로컬 파일 기반 Milvus 인스턴스로 작업할 수 있게 해줘요.
from pymilvus import MilvusClient
# Initialize the Milvus client
milvus_client = MilvusClient("milvus_twelvelabs_demo.db")
print("Successfully connected to Milvus")
이 코드는 모든 데이터를 milvus_twelvelabs_demo.db라는 파일에 저장할 새 Milvus 클라이언트 인스턴스를 만들어요. 이 파일 기반 접근 방식은 개발과 테스트 목적에 이상적이에요.
비디오 임베딩용 Milvus 컬렉션 만들기 (Creating a Milvus Collection for Video Embeddings)
이제 Milvus에 연결됐으니 비디오 임베딩과 관련 메타데이터를 저장할 컬렉션을 만들어요. 컬렉션 스키마를 정의하고 아직 존재하지 않으면 컬렉션을 만들어요.
# Initialize the collection name
collection_name = "twelvelabs_demo_collection"
# Check if the collection already exists and drop it if it does
if milvus_client.has_collection(collection_name=collection_name):
milvus_client.drop_collection(collection_name=collection_name)
# Create the collection
milvus_client.create_collection(
collection_name=collection_name,
dimension=1024 # The dimension of the Twelve Labs embeddings
)
print(f"Collection '{collection_name}' created successfully")
이 코드에서는 컬렉션이 이미 존재하는지 확인하고 존재하면 삭제해요. 이렇게 하면 깨끗한 상태로 시작할 수 있어요. Twelve Labs 임베딩의 출력 차원과 일치하는 1024 차원으로 컬렉션을 만들어요.
Twelve Labs Embed API로 임베딩 생성하기 (Generating Embeddings with Twelve Labs Embed API)
Twelve Labs Embed API를 사용해 비디오 임베딩을 생성하려면 Twelve Labs Python SDK를 사용해요. 이 과정은 임베딩 작업을 만들고 완료를 기다린 뒤 결과를 검색하는 것을 포함해요. 구현 방법은 다음과 같아요.
먼저 Twelve Labs SDK가 설치되어 있고 필요한 모듈을 가져왔는지 확인해요.
from twelvelabs import TwelveLabs
from twelvelabs.models.embed import EmbeddingsTask
import os
# Retrieve the API key from environment variables
TWELVE_LABS_API_KEY = os.getenv('TWELVE_LABS_API_KEY')
Twelve Labs 클라이언트를 초기화해요.
twelvelabs_client = TwelveLabs(api_key=TWELVE_LABS_API_KEY)
주어진 비디오 URL에 대한 임베딩을 생성하는 함수를 만들어요.
def generate_embedding(video_url):
"""
Generate embeddings for a given video URL using the Twelve Labs API.
This function creates an embedding task for the specified video URL using
the Marengo-retrieval-2.6 engine. It monitors the task progress and waits
for completion. Once done, it retrieves the task result and extracts the
embeddings along with their associated metadata.
Args:
video_url (str): The URL of the video to generate embeddings for.
Returns:
tuple: A tuple containing two elements:
1. list: A list of dictionaries, where each dictionary contains:
- 'embedding': The embedding vector as a list of floats.
- 'start_offset_sec': The start time of the segment in seconds.
- 'end_offset_sec': The end time of the segment in seconds.
- 'embedding_scope': The scope of the embedding (e.g., 'shot', 'scene').
2. EmbeddingsTaskResult: The complete task result object from Twelve Labs API.
Raises:
Any exceptions raised by the Twelve Labs API during task creation,
execution, or retrieval.
"""
# Create an embedding task
task = twelvelabs_client.embed.task.create(
engine_name="Marengo-retrieval-2.6",
video_url=video_url
)
print(f"Created task: id={task.id} engine_name={task.engine_name} status={task.status}")
# Define a callback function to monitor task progress
def on_task_update(task: EmbeddingsTask):
print(f" Status={task.status}")
# Wait for the task to complete
status = task.wait_for_done(
sleep_interval=2,
callback=on_task_update
)
print(f"Embedding done: {status}")
# Retrieve the task result
task_result = twelvelabs_client.embed.task.retrieve(task.id)
# Extract and return the embeddings
embeddings = []
for v in task_result.video_embeddings:
embeddings.append({
'embedding': v.embedding.float,
'start_offset_sec': v.start_offset_sec,
'end_offset_sec': v.end_offset_sec,
'embedding_scope': v.embedding_scope
})
return embeddings, task_result
함수를 사용해 비디오 임베딩을 생성해요.
# Example usage
video_url = "https://example.com/your-video.mp4"
# Generate embeddings for the video
embeddings, task_result = generate_embedding(video_url)
print(f"Generated {len(embeddings)} embeddings for the video")
for i, emb in enumerate(embeddings):
print(f"Embedding {i+1}:")
print(f" Scope: {emb['embedding_scope']}")
print(f" Time range: {emb['start_offset_sec']} - {emb['end_offset_sec']} seconds")
print(f" Embedding vector (first 5 values): {emb['embedding'][:5]}")
print()
이 구현을 사용하면 Twelve Labs Embed API를 통해 어떤 비디오 URL이든 임베딩을 생성할 수 있어요. generate_embedding 함수는 작업 생성부터 결과 검색까지 전체 과정을 처리해요. 각 임베딩 벡터와 메타데이터(시간 범위와 범위)를 담은 딕셔너리 목록을 반환해요. 프로덕션 환경에서는 네트워크 문제나 API 한도 같은 잠재적 오류를 처리하는 것을 기억하세요. 사용 사례에 따라 재시도나 더 견고한 오류 처리를 구현할 수도 있어요.
Milvus에 임베딩 삽입하기 (Inserting Embeddings into Milvus)
Twelve Labs Embed API로 임베딩을 생성한 후 다음 단계는 이 임베딩과 메타데이터를 Milvus 컬렉션에 삽입하는 것이에요. 이 과정을 통해 비디오 임베딩을 저장하고 인덱싱해 나중에 효율적인 유사도 검색을 할 수 있어요.
Milvus에 임베딩을 삽입하는 방법은 다음과 같아요.
def insert_embeddings(milvus_client, collection_name, task_result, video_url):
"""
Insert embeddings into the Milvus collection.
Args:
milvus_client: The Milvus client instance.
collection_name (str): The name of the Milvus collection to insert into.
task_result (EmbeddingsTaskResult): The task result containing video embeddings.
video_url (str): The URL of the video associated with the embeddings.
Returns:
MutationResult: The result of the insert operation.
This function takes the video embeddings from the task result and inserts them
into the specified Milvus collection. Each embedding is stored with additional
metadata including its scope, start and end times, and the associated video URL.
"""
data = []
for i, v in enumerate(task_result.video_embeddings):
data.append({
"id": i,
"vector": v.embedding.float,
"embedding_scope": v.embedding_scope,
"start_offset_sec": v.start_offset_sec,
"end_offset_sec": v.end_offset_sec,
"video_url": video_url
})
insert_result = milvus_client.insert(collection_name=collection_name, data=data)
print(f"Inserted {len(data)} embeddings into Milvus")
return insert_result
# Usage example
video_url = "https://example.com/your-video.mp4"
# Assuming this function exists from previous step
embeddings, task_result = generate_embedding(video_url)
# Insert embeddings into the Milvus collection
insert_result = insert_embeddings(milvus_client, collection_name, task_result, video_url)
print(insert_result)
이 함수는 임베딩 벡터, 시간 범위, 소스 비디오 URL 같은 모든 관련 메타데이터를 포함해 삽입 데이터를 준비해요. 그런 다음 Milvus 클라이언트를 사용해 이 데이터를 지정된 컬렉션에 삽입해요.
유사도 검색 수행하기 (Performing Similarity Search)
임베딩이 Milvus에 저장되면 쿼리 벡터를 기준으로 가장 관련성 높은 비디오 세그먼트를 찾기 위해 유사도 검색을 수행할 수 있어요. 이 기능을 구현하는 방법은 다음과 같아요.
def perform_similarity_search(milvus_client, collection_name, query_vector, limit=5):
"""
Perform a similarity search on the Milvus collection.
Args:
milvus_client: The Milvus client instance.
collection_name (str): The name of the Milvus collection to search in.
query_vector (list): The query vector to search for similar embeddings.
limit (int, optional): The maximum number of results to return. Defaults to 5.
Returns:
list: A list of search results, where each result is a dictionary containing
the matched entity's metadata and similarity score.
This function searches the specified Milvus collection for embeddings similar to
the given query vector. It returns the top matching results, including metadata
such as the embedding scope, time range, and associated video URL for each match.
"""
search_results = milvus_client.search(
collection_name=collection_name,
data=[query_vector],
limit=limit,
output_fields=["embedding_scope", "start_offset_sec", "end_offset_sec", "video_url"]
)
return search_results
# define the query vector
# We use the embedding inserted previously as an example. In practice, you can replace it with any video embedding you want to query.
query_vector = task_result.video_embeddings[0].embedding.float
# Perform a similarity search on the Milvus collection
search_results = perform_similarity_search(milvus_client, collection_name, query_vector)
print("Search Results:")
for i, result in enumerate(search_results[0]):
print(f"Result {i+1}:")
print(f" Video URL: {result['entity']['video_url']}")
print(f" Time Range: {result['entity']['start_offset_sec']} - {result['entity']['end_offset_sec']} seconds")
print(f" Similarity Score: {result['entity']['distance']}")
print()
이 구현은 다음을 수행해요.
- 쿼리 벡터를 받아 Milvus 컬렉션에서 유사한 임베딩을 검색하는
perform_similarity_search함수를 정의해요. - Milvus 클라이언트의
search메서드를 사용해 가장 유사한 벡터를 찾아요. - 일치하는 비디오 세그먼트의 메타데이터를 포함해 검색할 출력 필드를 지정해요.
- 쿼리 비디오로 이 함수를 사용하는 예시를 제공해요. 먼저 임베딩을 생성한 뒤 검색에 사용해요.
- 관련 메타데이터와 유사도 점수를 포함한 검색 결과를 출력해요.
이 함수들을 구현하면 Milvus에 비디오 임베딩을 저장하고 유사도 검색을 수행하는 완전한 워크플로우를 만들었어요. 이 설정은 Twelve Labs Embed API가 생성한 멀티모달 임베딩을 기반으로 유사한 비디오 콘텐츠를 효율적으로 검색할 수 있게 해줘요.
성능 최적화 (Optimizing Performance)
이 앱을 한 단계 더 끌어올려 볼까요! 대규모 비디오 컬렉션을 다룰 때 성능이 핵심이에요. 최적화를 위해 임베딩 생성과 Milvus 삽입에 배치 처리를 구현해야 해요. 이렇게 하면 여러 비디오를 동시에 처리해 전체 처리 시간을 크게 줄일 수 있어요. 또한 Milvus의 파티셔닝 기능을 활용해 데이터를 더 효율적으로 구성할 수 있어요. 아마 비디오 카테고리나 기간별로요. 이렇게 하면 관련 파티션만 검색할 수 있어 쿼리 속도를 높일 수 있어요.
또 다른 최적화 팁은 자주 접근하는 임베딩이나 검색 결과에 캐싱 메커니즘을 사용하는 것이에요. 이렇게 하면 인기 있는 쿼리의 응답 시간을 극적으로 개선할 수 있어요. 특정 데이터셋과 쿼리 패턴에 맞게 Milvus 인덱스 파라미터를 미세 조정하는 것도 잊지 마세요. 여기서 약간만 조정해도 검색 성능을 크게 높일 수 있어요.
고급 기능 (Advanced Features)
이제 앱을 돋보이게 할 멋진 기능을 추가해요! 텍스트 쿼리와 비디오 쿼리를 결합한 하이브리드 검색을 구현할 수 있어요. 사실 Twelve Labs Embed API는 텍스트 쿼리용 텍스트 임베딩도 생성할 수 있어요. 사용자가 텍스트 설명과 샘플 비디오 클립을 모두 입력할 수 있다고 상상해 보세요. 두 가지 모두에 대한 임베딩을 생성하고 Milvus에서 가중 검색을 수행해요. 그러면 매우 정밀한 결과를 얻을 수 있어요.
또 다른 멋진 추가 기능은 비디오 내 시간적 검색이에요. 긴 비디오를 각각 자체 임베딩을 가진 작은 세그먼트로 나눌 수 있어요. 이렇게 하면 사용자가 전체 클립뿐 아니라 비디오 내 특정 시점도 찾을 수 있어요. 그리고 기본적인 비디오 분석은 어떨까요? 임베딩을 사용해 유사한 비디오 세그먼트를 클러스터링하고, 트렌드를 감지하고, 대규모 비디오 컬렉션에서 이상치를 식별할 수도 있어요.
오류 처리와 로깅 (Error Handling and Logging)
문제가 발생할 수 있고, 그럴 때 준비된 상태여야 해요. 견고한 오류 처리를 구현하는 것이 중요해요. API 호출과 데이터베이스 작업을 try-except 블록으로 감싸 실패 시 사용자에게 정보를 제공하는 오류 메시지를 주어야 해요. 네트워크 관련 문제에는 지수 백오프가 있는 재시도 구현이 일시적인 결함을 우아하게 처리하는 데 도움이 돼요.
로깅은 디버깅과 모니터링을 위한 최고의 친구예요. Python의 logging 모듈을 사용해 애플리케이션 전반에 걸쳐 중요한 이벤트, 오류, 성능 메트릭을 추적해야 해요. 서로 다른 로그 레벨을 설정해요. 개발용 DEBUG, 일반 작업용 INFO, 중요 문제용 ERROR. 그리고 파일 크기를 관리하기 위해 로그 회전(log rotation)을 구현하는 것도 잊지 마세요. 적절한 로깅이 준비되면 문제를 빠르게 식별하고 해결할 수 있어, 비디오 검색 앱이 확장될 때도 원활하게 실행되도록 보장할 수 있어요.
결론 (Conclusion)
축하해요! 이제 Twelve Labs Embed API와 Milvus를 사용해 강력한 시맨틱 비디오 검색 애플리케이션을 구축했어요. 이 통합을 통해 전례 없는 정확도와 효율로 비디오 콘텐츠를 처리, 저장, 검색할 수 있어요. 멀티모달 임베딩을 활용해 비디오 데이터의 미묘한 차이를 이해하는 시스템을 만들었고, 콘텐츠 발견, 추천 시스템, 고급 비디오 분석을 위한 흥미로운 가능성을 열었어요.
애플리케이션을 계속 개발하고 개선하면서, Twelve Labs의 고급 임베딩 생성과 Milvus의 확장 가능한 벡터 저장의 조합이 훨씬 더 복잡한 비디오 이해 문제를 해결하기 위한 견고한 기반을 제공한다는 것을 기억하세요. 논의한 고급 기능을 실험하고 비디오 검색과 분석에서 가능한 경계를 넓혀 보기를 권장해요.
더 알아보기 (Learn more)
- bulk_insert — 배치 임베딩 삽입
- partition_key — 파티셔닝 기능
- index-vector-fields — 인덱스 파라미터 미세 조정
- Twelve Labs Embed API — Twelve Labs 문서
- Milvus 공식 문서 — 벡터 데이터베이스 관련 자료