본문 바로가기
WIKI 기술 지식 베이스

Milvus와 Cognee로 RAG 구축하기 (Build RAG with Milvus and Cognee)

원문 보기 위키 갱신

Cognee는 확장 가능하고 모듈식인 ECL(Extract, Cognify, Load) 파이프라인으로 AI 애플리케이션 개발을 간소화하는 개발자 우선 플랫폼이에요. Milvus와 원활하게 통합함으로써 Cognee는 대화, 문서, 전사본의 효율적인 연결과 검색을 가능하게 하고, 할루시네이션을 줄이며 운영 비용을 최적화해요.

Milvus 같은 벡터 저장소, 그래프 데이터베이스, LLM에 대한 강력한 지원으로 Cognee는 검색 증강 생성(RAG) 시스템을 구축하기 위한 유연하고 사용자 지정 가능한 프레임워크를 제공해요. 프로덕션 준비 아키텍처는 AI 기반 애플리케이션의 정확성과 효율성을 향상시켜요.

이 튜토리얼에서는 Milvus와 Cognee로 RAG(Retrieval-Augmented Generation) 파이프라인을 구축하는 방법을 보여 줘요.

출처: Milvus 문서

본문

$ pip install pymilvus git+https://github.com/topoteretes/cognee.git

Google Colab을 사용한다면 방금 설치한 의존성을 활성화하기 위해 런타임을 재시작해야 할 수 있어요(화면 상단의 "Runtime" 메뉴를 클릭하고 드롭다운에서 "Restart session"을 선택).

기본적으로 이 예시에서는 LLM으로 OpenAI를 사용해요. api key를 준비하고 set_llm_api_key() 구성 함수에 설정해야 해요.

Milvus를 벡터 데이터베이스로 구성하려면 VECTOR_DB_PROVIDER를 milvus로 설정하고 VECTOR_DB_URL과 VECTOR_DB_KEY를 지정해요. 이 데모에서는 Milvus Lite를 사용해 데이터를 저장하므로 VECTOR_DB_URL만 제공하면 돼요.

import os

import cognee

cognee.config.set_llm_api_key("YOUR_OPENAI_API_KEY")

os.environ["VECTOR_DB_PROVIDER"] = "milvus"
os.environ["VECTOR_DB_URL"] = "./milvus.db"

환경 변수 VECTOR_DB_URL과 VECTOR_DB_KEY와 관련해서:

  • VECTOR_DB_URL을 로컬 파일(예: ./milvus.db)로 설정하는 것이 가장 편리한 방법이에요. Milvus Lite를 자동으로 사용해 모든 데이터를 이 파일에 저장하기 때문이에요.
  • 대규모 데이터가 있다면 docker 또는 kubernetes에 더 성능이 좋은 Milvus 서버를 설정할 수 있어요. 이 경우 서버 uri(예: http://localhost:19530)를 VECTOR_DB_URL로 사용하세요.
  • Milvus의 완전 관리형 클라우드 서비스인 Zilliz Cloud를 사용하려면 Zilliz Cloud의 Public Endpoint와 Api key에 해당하는 VECTOR_DB_URL과 VECTOR_DB_KEY를 조정하세요.

데이터 준비 (Prepare the data)

Milvus Documentation 2.4.x의 FAQ 페이지를 RAG의 사내 지식으로 사용해요. 이는 간단한 RAG 파이프라인의 좋은 데이터 소스예요.

zip 파일을 다운로드하고 문서를 milvus_docs 폴더에 추출해요.

$ wget https://github.com/milvus-io/milvus-docs/releases/download/v2.4.6-preview/milvus_docs_2.4.x_en.zip
$ unzip -q milvus_docs_2.4.x_en.zip -d milvus_docs

milvus_docs/en/faq 폴더에서 모든 마크다운 파일을 로드해요. 각 문서에 대해 단순히 # 로 파일의 내용을 구분하며, 이는 마크다운 파일의 각 주요 부분 내용을 대략 분리할 수 있어요.

from glob import glob

text_lines = []

for file_path in glob("milvus_docs/en/faq/*.md", recursive=True):
    with open(file_path, "r") as file:
        file_text = file.read()

    text_lines += file_text.split("# ")

RAG 구축 (Build RAG)

Cognee 데이터 리셋 (Resetting Cognee Data)

await cognee.prune.prune_data()
await cognee.prune.prune_system(metadata=True)

깨끗한 상태가 준비되면 이제 데이터셋을 추가하고 지식 그래프로 처리할 수 있어요.

데이터 추가 및 Cognify (Adding Data and Cognifying)

await cognee.add(data=text_lines, dataset_name="milvus_faq")
await cognee.cognify()

# [DocumentChunk(id=UUID('6889e7ef-3670-555c-bb16-3eb50d1d30b0'), updated_at=datetime.datetime(2024, 12, 4, 6, 29, 46, 472907, tzinfo=datetime.timezone.utc), text='Does the query perform in memory? What are incremental data and historical data?\n\nYes. When ...
# ...

add 메서드는 데이터셋(Milvus FAQ)을 Cognee에 로드하고, cognify 메서드는 엔티티, 관계, 요약을 추출하고 지식 그래프를 구축하도록 데이터를 처리해요.

요약 쿼리 (Querying for Summaries)

데이터가 처리되었으니 지식 그래프를 쿼리해 보아요.

from cognee.api.v1.search import SearchType

query_text = "How is data stored in milvus?"
search_results = await cognee.search(SearchType.SUMMARIES, query_text=query_text)

print(search_results[0])
{'id': 'de5c6713-e079-5d0b-b11d-e9bacd1e0d73', 'text': 'Milvus stores two data types: inserted data and metadata.'}

이 쿼리는 쿼리 텍스트와 관련된 요약을 지식 그래프에서 검색하고, 가장 관련성 높은 후보를 출력해요.

청크 쿼리 (Querying for Chunks)

요약은 하이레벨 통찰을 제공하지만, 더 세부적인 정보를 위해 처리된 데이터셋에서 특정 데이터 청크를 직접 쿼리할 수 있어요. 이 청크들은 지식 그래프 생성 중 추가되고 분석된 원본 데이터에서 파생된 것이에요.

from cognee.api.v1.search import SearchType

query_text = "How is data stored in milvus?"
search_results = await cognee.search(SearchType.CHUNKS, query_text=query_text)

더 읽기 좋게 포맷하고 표시해 보아요!

def format_and_print(data):
    print("ID:", data["id"])
    print("\nText:\n")
    paragraphs = data["text"].split("\n\n")
    for paragraph in paragraphs:
        print(paragraph.strip())
        print()

format_and_print(search_results[0])
ID: 4be01c4b-9ee5-541c-9b85-297883934ab3

Text:

Where does Milvus store data?

Milvus deals with two types of data, inserted data and metadata.

Inserted data, including vector data, scalar data, and collection-specific schema, are stored in persistent storage as incremental log. Milvus supports multiple object storage backends, including [MinIO](https://min.io/), [AWS S3](https://aws.amazon.com/s3/?nc1=h_ls), [Google Cloud Storage](https://cloud.google.com/storage?hl=en#object-storage-for-companies-of-all-sizes) (GCS), [Azure Blob Storage](https://azure.microsoft.com/en-us/products/storage/blobs), [Alibaba Cloud OSS](https://www.alibabacloud.com/product/object-storage-service), and [Tencent Cloud Object Storage](https://www.tencentcloud.com/products/cos) (COS).

Metadata are generated within Milvus. Each Milvus module has its own metadata that are stored in etcd.

###

이전 단계에서 우리는 Milvus FAQ 데이터셋의 요약과 특정 데이터 청크를 모두 쿼리했어요. 이는 상세한 통찰과 세부 정보를 제공했지만, 데이터셋이 커서 지식 그래프 내의 의존성을 명확히 시각화하기 어려웠어요.

이를 해결하기 위해 Cognee 환경을 리셋하고 더 작고 집중된 데이터셋으로 작업할 거예요. 이렇게 하면 cognify 과정에서 추출된 관계와 의존성을 더 잘 시연할 수 있어요. 데이터를 단순화함으로써 Cognee가 지식 그래프에서 정보를 어떻게 구성하고 구조화하는지 명확히 볼 수 있어요.

Cognee 리셋 (Reset Cognee)

await cognee.prune.prune_data()
await cognee.prune.prune_system(metadata=True)

집중 데이터셋 추가 (Adding the Focused Dataset)

여기서는 집중적이고 쉽게 해석 가능한 지식 그래프를 보장하기 위해 텍스트 한 줄만 있는 더 작은 데이터셋을 추가하고 처리해요.

# We only use one line of text as the dataset, which simplifies the output later
text = """
    Natural language processing (NLP) is an interdisciplinary
    subfield of computer science and information retrieval.
    """

await cognee.add(text)
await cognee.cognify()

통찰 쿼리 (Querying for Insights)

이 더 작은 데이터셋에 집중함으로써 이제 지식 그래프 내의 관계와 구조를 명확히 분석할 수 있어요.

query_text = "Tell me about NLP"
search_results = await cognee.search(SearchType.INSIGHTS, query_text=query_text)

for result_text in search_results:
    print(result_text)

# Example output:
# ({'id': UUID('bc338a39-64d6-549a-acec-da60846dd90d'), 'updated_at': datetime.datetime(2024, 11, 21, 12, 23, 1, 211808, tzinfo=datetime.timezone.utc), 'name': 'natural language processing', 'description': 'An interdisciplinary subfield of computer science and information retrieval.'}, {'relationship_name': 'is_a_subfield_of', 'source_node_id': UUID('bc338a39-64d6-549a-acec-da60846dd90d'), 'target_node_id': UUID('6218dbab-eb6a-5759-a864-b3419755ffe0'), 'updated_at': datetime.datetime(2024, 11, 21, 12, 23, 15, 473137, tzinfo=datetime.timezone.utc)}, {'id': UUID('6218dbab-eb6a-5759-a864-b3419755ffe0'), 'updated_at': datetime.datetime(2024, 11, 21, 12, 23, 1, 211808, tzinfo=datetime.timezone.utc), 'name': 'computer science', 'description': 'The study of computation and information processing.'})
# (...)
#
# It represents nodes and relationships in the knowledge graph:
# - The first element is the source node (e.g., 'natural language processing').
# - The second element is the relationship between nodes (e.g., 'is_a_subfield_of').
# - The third element is the target node (e.g., 'computer science').

이 출력은 지식 그래프 쿼리 결과를 나타내며, 처리된 데이터셋에서 추출된 엔티티(노드)와 그 관계(엣지)를 보여 줘요. 각 튜플에는 소스 엔티티, 관계 유형, 대상 엔티티와 함께 고유 ID, 설명, 타임스탬프 같은 메타데이터가 포함돼요. 그래프는 핵심 개념과 그 의미적 연결을 강조해 데이터셋에 대한 구조화된 이해를 제공해요.

축하해요! 이제 Milvus와 함께 Cognee의 기본 사용법을 배웠어요. Cognee의 더 고급 사용법을 알고 싶다면 공식 page을 참고하세요.

더 알아보기 (Learn more)