본문 바로가기
WIKI 기술 지식 베이스

Milvus와 DeepSeek로 RAG 구축하기 (Build RAG with Milvus and DeepSeek)

원문 보기 위키 갱신

DeepSeek는 고성능 언어 모델로 개발자가 AI 애플리케이션을 만들고 확장할 수 있게 해줘요. 효율적인 추론, 유연한 API, 그리고 견고한 추론·검색 작업을 위한 고급 Mixture-of-Experts(MoE) 아키텍처를 제공해요.

이 튜토리얼에서는 Milvus와 DeepSeek로 RAG(Retrieval-Augmented Generation, 검색 증강 생성) 파이프라인을 구축하는 방법을 보여 드릴게요.

출처: Milvus 문서

본문

준비 (Preparation)

의존성과 환경 (Dependencies and Environment)

! pip install --upgrade pymilvus[model] milvus-lite openai requests tqdm

Google Colab을 사용한다면, 방금 설치한 의존성을 활성화하기 위해 런타임을 재시작해야 할 수 있어요. (화면 상단의 "Runtime" 메뉴를 클릭하고 드롭다운에서 "Restart session"을 선택하세요.)

DeepSeek는 OpenAI 스타일 API를 제공해요. 공식 웹사이트에 로그인해 api key DEEPSEEK_API_KEY를 환경 변수로 준비하세요.

import os

os.environ["DEEPSEEK_API_KEY"] = "***********"

데이터 준비 (Prepare the data)

간단한 RAG 파이프라인에 좋은 데이터 소스인 Milvus Documentation 2.4.x의 FAQ 페이지를 RAG의 비공개 지식으로 사용해요.

zip 파일을 다운로드하고 문서를 milvus_docs 폴더에 압축 해제해요.

! wget https://github.com/milvus-io/milvus-docs/releases/download/v2.4.6-preview/milvus_docs_2.4.x_en.zip
! unzip -q milvus_docs_2.4.x_en.zip -d milvus_docs

milvus_docs/en/faq 폴더에서 모든 마크다운 파일을 불러와요. 각 문서에 대해 파일 내용을 단순히 "# "로 분리하면, 마크다운 파일의 각 주요 부분 내용을 대략적으로 나눌 수 있어요.

from glob import glob

text_lines = []

for file_path in glob("milvus_docs/en/faq/*.md", recursive=True):
    with open(file_path, "r") as file:
        file_text = file.read()

    text_lines += file_text.split("# ")

LLM과 임베딩 모델 준비 (Prepare the LLM and Embedding Model)

DeepSeek는 OpenAI 스타일 API를 제공하므로, 약간만 조정해서 같은 API로 LLM을 호출할 수 있어요.

from openai import OpenAI

deepseek_client = OpenAI(
    api_key=os.environ["DEEPSEEK_API_KEY"],
    base_url="https://api.deepseek.com",
)

milvus_model로 텍스트 임베딩을 생성할 임베딩 모델을 정의해요. 예시로 DefaultEmbeddingFunction 모델(사전 훈련된 가벼운 임베딩 모델)을 사용해요.

from pymilvus import model as milvus_model

embedding_model = milvus_model.DefaultEmbeddingFunction()

테스트 임베딩을 생성하고 차원과 처음 몇 개 요소를 출력해요.

test_embedding = embedding_model.encode_queries(["This is a test"])[0]
embedding_dim = len(test_embedding)
print(embedding_dim)
print(test_embedding[:10])
768
[-0.04836066  0.07163023 -0.01130064 -0.03789345 -0.03320649 -0.01318448
 -0.03041712 -0.02269499 -0.02317863 -0.00426028]

Milvus에 데이터 로드하기 (Load data into Milvus)

컬렉션 만들기 (Create the Collection)

from pymilvus import MilvusClient

milvus_client = MilvusClient(uri="./milvus_demo.db")

collection_name = "my_rag_collection"

MilvusClient의 인수에 대해 설명할게요:

  • uri를 로컬 파일(예: ./milvus.db)로 설정하는 것이 가장 편리한 방법이에요. Milvus Lite를 자동으로 사용해 모든 데이터를 이 파일에 저장해요.
  • 데이터가 대규모라면 docker나 kubernetes에 더 성능 좋은 Milvus 서버를 설정할 수 있어요. 이 경우 서버 uri(예: http://localhost:19530)를 uri로 사용하세요.
  • Milvus의 완전 관리형 클라우드 서비스인 Zilliz Cloud를 사용하려면, Zilliz Cloud의 Public Endpoint와 Api key에 해당하는 uri와 token을 조정하세요.

컬렉션이 이미 존재하는지 확인하고 존재하면 제거해요.

if milvus_client.has_collection(collection_name):
    milvus_client.drop_collection(collection_name)

지정한 파라미터로 새 컬렉션을 만들어요.

어떤 필드 정보도 지정하지 않으면 Milvus가 기본 키용 id 필드와 벡터 데이터 저장용 vector 필드를 자동으로 만들어요. 예약된 JSON 필드는 스키마에 정의되지 않은 필드와 그 값을 저장하는 데 사용돼요.

milvus_client.create_collection(
    collection_name=collection_name,
    dimension=embedding_dim,
    metric_type="IP",  # Inner product distance
    consistency_level="Bounded",  # Supported values are ("Strong", "Session", "Bounded", "Eventually"). See https://milvus.io/docs/tune_consistency.md#Consistency-Level for more details.
)

데이터 삽입하기 (Insert data)

텍스트 줄을 순회하며 임베딩을 만들고, 데이터를 Milvus에 삽입해요.

여기 text라는 새 필드가 있는데, 이는 컬렉션 스키마에 정의되지 않은 필드예요. 이 필드는 예약된 JSON 동적 필드에 자동으로 추가되며, 높은 수준에서 일반 필드처럼 취급할 수 있어요.

from tqdm import tqdm

data = []

doc_embeddings = embedding_model.encode_documents(text_lines)

for i, line in enumerate(tqdm(text_lines, desc="Creating embeddings")):
    data.append({"id": i, "vector": doc_embeddings[i], "text": line})

milvus_client.insert(collection_name=collection_name, data=data)
Creating embeddings:   0%|          | 0/72 [00:00<?, ?it/s]huggingface/tokenizers: The current process just got forked, after parallelism has already been used. Disabling parallelism to avoid deadlocks...
To disable this warning, you can either:
    - Avoid using `tokenizers` before the fork if possible
    - Explicitly set the environment variable TOKENIZERS_PARALLELISM=(true | false)
Creating embeddings: 100%|██████████| 72/72 [00:00<00:00, 246522.36it/s]

{'insert_count': 72, 'ids': [0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71], 'cost': 0}

RAG 구축하기 (Build RAG)

쿼리용 데이터 검색하기 (Retrieve data for a query)

Milvus에 대한 자주 나오는 질문을 하나 정해 볼게요.

question = "How is data stored in milvus?"

컬렉션에서 질문을 검색하고 의미적으로 가장 가까운 상위 3개 결과를 가져와요.

search_res = milvus_client.search(
    collection_name=collection_name,
    data=embedding_model.encode_queries(
        [question]
    ),  # Convert the question to an embedding vector
    limit=3,  # Return top 3 results
    search_params={"metric_type": "IP", "params": {}},  # Inner product distance
    output_fields=["text"],  # Return the text field
)

쿼리의 검색 결과를 한번 살펴볼게요.

import json

retrieved_lines_with_distances = [
    (res["entity"]["text"], res["distance"]) for res in search_res[0]
]
print(json.dumps(retrieved_lines_with_distances, indent=4))
[
    [
        " Where does Milvus store data?\n\nMilvus deals with two types of data, inserted data and metadata. \n\nInserted data, including vector data, scalar data, and collection-specific schema, are stored in persistent storage as incremental log. Milvus supports multiple object storage backends, including [MinIO](https://min.io/), [AWS S3](https://aws.amazon.com/s3/?nc1=h_ls), [Google Cloud Storage](https://cloud.google.com/storage?hl=en#object-storage-for-companies-of-all-sizes) (GCS), [Azure Blob Storage](https://azure.microsoft.com/en-us/products/storage/blobs), [Alibaba Cloud OSS](https://www.alibabacloud.com/product/object-storage-service), and [Tencent Cloud Object Storage](https://www.tencentcloud.com/products/cos) (COS).\n\nMetadata are generated within Milvus. Each Milvus module has its own metadata that are stored in etcd.\n\n###",
        0.6572665572166443
    ],
    [
        "How does Milvus flush data?\n\nMilvus returns success when inserted data are loaded to the message queue. However, the data are not yet flushed to the disk. Then Milvus' data node writes the data in the message queue to persistent storage as incremental logs. If `flush()` is called, the data node is forced to write all data in the message queue to persistent storage immediately.\n\n###",
        0.6312146186828613
    ],
    [
        "How does Milvus handle vector data types and precision?\n\nMilvus supports Binary, Float32, Float16, and BFloat16 vector types.\n\n- Binary vectors: Store binary data as sequences of 0s and 1s, used in image processing and information retrieval.\n- Float32 vectors: Default storage with a precision of about 7 decimal digits. Even Float64 values are stored with Float32 precision, leading to potential precision loss upon retrieval.\n- Float16 and BFloat16 vectors: Offer reduced precision and memory usage. Float16 is suitable for applications with limited bandwidth and storage, while BFloat16 balances range and efficiency, commonly used in deep learning to reduce computational requirements without significantly impacting accuracy.\n\n###",
        0.6115777492523193
    ]
]

LLM으로 RAG 응답 얻기 (Use LLM to get a RAG response)

검색한 문서를 문자열 형식으로 변환해요.

context = "\n".join(
    [line_with_distance[0] for line_with_distance in retrieved_lines_with_distances]
)

언어 모델용 시스템·사용자 프롬프트를 정의해요. 이 프롬프트는 Milvus에서 검색한 문서로 조합돼요.

SYSTEM_PROMPT = """
Human: You are an AI assistant. You are able to find answers to the questions from the contextual passage snippets provided.
"""
USER_PROMPT = f"""
Use the following pieces of information enclosed in <context> tags to provide an answer to the question enclosed in <question> tags.
<context>
{context}
</context>
<question>
{question}
</question>
"""

DeepSeek가 제공하는 deepseek-chat 모델을 사용해 프롬프트 기반 응답을 생성해요.

response = deepseek_client.chat.completions.create(
    model="deepseek-chat",
    messages=[
        {"role": "system", "content": SYSTEM_PROMPT},
        {"role": "user", "content": USER_PROMPT},
    ],
)
print(response.choices[0].message.content)
In Milvus, data is stored in two main categories: inserted data and metadata.

1. **Inserted Data**: This includes vector data, scalar data, and collection-specific schema. The inserted data is stored in persistent storage as incremental logs. Milvus supports various object storage backends for this purpose, such as MinIO, AWS S3, Google Cloud Storage (GCS), Azure Blob Storage, Alibaba Cloud OSS, and Tencent Cloud Object Storage (COS).

2. **Metadata**: Metadata is generated within Milvus and is specific to each Milvus module. This metadata is stored in etcd, a distributed key-value store.

Additionally, when data is inserted, it is first loaded into a message queue, and Milvus returns success at this stage. The data is then written to persistent storage as incremental logs by the data node. If the `flush()` function is called, the data node is forced to write all data in the message queue to persistent storage immediately.

좋아요! Milvus와 DeepSeek로 RAG 파이프라인을 성공적으로 구축했어요.

더 알아보기 (Learn more)