본문 바로가기
WIKI 기술 지식 베이스

Exa와 Milvus로 이중 소스 RAG 에이전트 구축하기 (Building a Dual-Source RAG Agent with Exa and Milvus)

원문 보기 위키 갱신

이 튜토리얼은 공개 웹(via Exa)과 사내 지식 베이스(via Milvus)를 모두 검색한 뒤 통합된 답변을 만들어내는 에이전트를 구축하는 방법을 보여 줘요. 이 에이전트는 OpenAI의 함수 호출(function calling)을 사용해 사용자 질문에 따라 어느 소스를 쿼리할지 자동으로 결정해요.

Exa는 AI 애플리케이션을 위해 설계된 검색 API로, Zilliz Cloud(완전 관리형 Milvus)가 지원해요. 전통적인 키워드 기반 검색 엔진과 달리 Exa는 시맨틱(신경망) 검색을 지원해요. 원하는 것을 자연어로 설명하면 의도를 이해해주거든요. 또한 콘텐츠 추출, 하이라이트, 카테고리 기반 필터링도 제공해요. Milvus는 확장 가능한 유사도 검색을 위해 설계된 오픈소스 벡터 데이터베이스예요. 이 둘을 LLM 에이전트와 결합하면 내부 독점 데이터와 최신 웹 정보를 하나의 워크플로에서 모두 검색하는 시스템을 구축할 수 있어요.

출처: Milvus 문서

본문

사전 요구 사항 (Prerequisites)

이 노트북을 실행하기 전에 다음 의존성이 설치되어 있는지 확인하세요.

$ pip install exa_py pymilvus openai

Google Colab을 사용한다면 방금 설치한 의존성을 활성화하기 위해 런타임을 재시작해야 할 수 있어요(화면 상단의 "Runtime" 메뉴를 클릭하고 드롭다운에서 "Restart session"을 선택).

Exa와 OpenAI의 API 키가 필요해요. 이를 환경 변수로 설정하세요.

import os

os.environ["EXA_API_KEY"] = "***********"
os.environ["OPENAI_API_KEY"] = "sk-***********"

클라이언트 초기화 (Initialize Clients)

Exa, OpenAI, Milvus 클라이언트를 설정해요. OpenAI의 text-embedding-3-small 모델로 벡터 임베딩을 생성하고, 인프라 설정이 전혀 필요 없는 로컬 벡터 저장에는 Milvus Lite를 사용해요.

import json
from openai import OpenAI
from pymilvus import MilvusClient, DataType
from exa_py import Exa

llm = OpenAI()
exa = Exa(api_key=os.environ["EXA_API_KEY"])
milvus = MilvusClient(uri="./milvus_exa_demo.db")

EMBED_MODEL = "text-embedding-3-small"
EMBED_DIM = 1536
COLLECTION = "private_kb"

MilvusVectorAdapter와 MilvusClient의 인자 관련해서:

  • uri를 ./milvus.db 같은 로컬 파일로 설정하는 것이 가장 편리한 방법이에요. Milvus Lite를 자동으로 사용해 모든 데이터를 이 파일에 저장하기 때문이에요.

  • 백만 개 이상의 벡터처럼 대규모 데이터가 있다면 Docker 또는 Kubernetes에 더 성능이 좋은 Milvus 서버를 설정할 수 있어요. 이 경우 서버 주소와 포트를 uri로 사용하세요. 예: http://localhost:19530. Milvus에서 인증 기능을 활성화했다면 :를 token으로 사용하고, 그렇지 않으면 token을 설정하지 마세요.

  • Zilliz Cloud, 즉 Milvus의 완전 관리형 클라우드 서비스를 사용하려면 Zilliz Cloud의 Public Endpoint와 Api key에 해당하는 uri와 token을 조정하세요.

임베딩 생성을 위한 헬퍼 함수를 정의해요. 이 함수는 인덱싱과 쿼리 모두에서 노트북 전반에 걸쳐 재사용할 거예요.

def embed_text(text: str | list[str]) -> list:
    """Generate embedding vector(s) using OpenAI."""
    resp = llm.embeddings.create(
        input=text if isinstance(text, list) else [text],
        model=EMBED_MODEL,
    )
    if isinstance(text, list):
        return [item.embedding for item in resp.data]
    return resp.data[0].embedding

사내 지식 베이스 구축 (Build the Private Knowledge Base, Milvus)

공개 웹에는 나타나지 않을 일련의 내부 회사 문서(제품 사양, 정책, 실적 보고서, API 문서)를 시뮬레이션해요. 실제 시나리오에서는 내부 위키, 데이터베이스, 문서 관리 시스템에서 가져올 수 있어요.

private_docs = [
    {
        "id": 1,
        "text": (
            "Acme Widget Pro supports up to 10,000 concurrent connections. "
            "It uses a proprietary compression algorithm (AcmeZip v3) that "
            "reduces payload size by 72% compared to gzip."
        ),
        "source": "product-spec.pdf",
    },
    {
        "id": 2,
        "text": (
            "Our return policy allows customers to return any product within "
            "30 days of purchase for a full refund. After 30 days, only store "
            "credit is offered. Damaged items must be reported within 48 hours."
        ),
        "source": "return-policy.md",
    },
    {
        "id": 3,
        "text": (
            "Q3 2025 revenue was $4.2M, up 18% from Q2. The growth was "
            "primarily driven by enterprise customers adopting Widget Pro. "
            "Churn rate dropped to 3.1%."
        ),
        "source": "q3-earnings.pdf",
    },
    {
        "id": 4,
        "text": (
            "Internal API rate limits: free tier 100 req/min, pro tier "
            "5,000 req/min, enterprise tier 50,000 req/min. Rate limit "
            "headers are X-RateLimit-Remaining and X-RateLimit-Reset."
        ),
        "source": "api-docs.md",
    },
    {
        "id": 5,
        "text": (
            "Employee onboarding checklist: 1) Sign NDA, 2) Set up VPN access, "
            "3) Enroll in mandatory security training, 4) Request Jira and "
            "Confluence access from IT, 5) Schedule 1:1 with manager."
        ),
        "source": "onboarding-guide.md",
    },
]

명시적 스키마로 Milvus 컬렉션을 만들고, 문서를 임베딩한 뒤 삽입해요.

if milvus.has_collection(COLLECTION):
    milvus.drop_collection(COLLECTION)

schema = milvus.create_schema(auto_id=False, enable_dynamic_field=True)
schema.add_field(field_name="id", datatype=DataType.INT64, is_primary=True)
schema.add_field(field_name="vector", datatype=DataType.FLOAT_VECTOR, dim=EMBED_DIM)
schema.add_field(field_name="text", datatype=DataType.VARCHAR, max_length=65535)
schema.add_field(field_name="source", datatype=DataType.VARCHAR, max_length=512)

index_params = milvus.prepare_index_params()
index_params.add_index(
    field_name="vector", index_type="AUTOINDEX", metric_type="COSINE"
)

milvus.create_collection(
    collection_name=COLLECTION,
    schema=schema,
    index_params=index_params,
    # consistency_level="Strong",
)

# Embed all documents in one batch call
embeddings = embed_text([doc["text"] for doc in private_docs])

milvus.insert(
    collection_name=COLLECTION,
    data=[
        {
            "id": doc["id"],
            "vector": emb,
            "text": doc["text"],
            "source": doc["source"],
        }
        for doc, emb in zip(private_docs, embeddings)
    ],
)

print(f"Inserted {len(private_docs)} documents into Milvus.")
Inserted 5 documents into Milvus.

빠른 테스트 쿼리로 검색이 동작하는지 확인해 보아요.

query = "What is the return policy?"
results = milvus.search(
    collection_name=COLLECTION,
    data=[embed_text(query)],
    limit=2,
    output_fields=["text", "source"],
)

for hit in results[0]:
    print(f"[score={hit['distance']:.3f}] ({hit['entity']['source']})")
    print(f"  {hit['entity']['text'][:120]}...")
    print()
[score=0.665] (return-policy.md)
  Our return policy allows customers to return any product within 30 days of purchase for a full refund. After 30 days, on...

[score=0.119] (q3-earnings.pdf)
  Q3 2025 revenue was $4.2M, up 18% from Q2. The growth was primarily driven by enterprise customers adopting Widget Pro. ...

Exa 검색 기능 탐색 (Explore Exa Search Capabilities)

에이전트를 구축하기 전에 Exa의 검색 기능을 살펴보아요. Exa는 서로 다른 시나리오에 유용한 여러 검색 모드를 지원해요.

콘텐츠 추출이 포함된 시맨틱 검색 — Exa는 링크뿐만 아니라 기사 텍스트, 핵심 하이라이트, AI 생성 요약을 단일 요청으로 반환할 수 있어요.

web_results = exa.search_and_contents(
    query="latest trends in AI agents 2026",
    type="auto",
    num_results=3,
    text={"max_characters": 3000},
    highlights={"num_sentences": 3},
)

for r in web_results.results:
    print(f"[{r.title}]")
    print(f"  URL: {r.url}")
    if r.highlights:
        print(f"  Highlight: {r.highlights[0][:150]}...")
    print()

카테고리 기반 필터링 — 결과를 "research paper", "news", "company", "tweet" 같은 특정 콘텐츠 유형으로 제한할 수 있어요. 고품질 소스를 원하고 노이즈를 피하고 싶을 때 유용해요.

filtered_results = exa.search_and_contents(
    query="retrieval augmented generation real world applications",
    category="research paper",
    num_results=3,
    highlights={"num_sentences": 2},
)

for r in filtered_results.results:
    print(f"- {r.title}")
    print(f"  {r.url}\n")

유사 기사 찾기 — URL이 주어지면 Exa는 비슷한 콘텐츠를 가진 다른 기사를 찾을 수 있어요. 좋은 출발점에서 연구를 확장하는 데 도움이 돼요.

if web_results.results:
    source_url = web_results.results[0].url
    similar = exa.find_similar_and_contents(
        url=source_url,
        num_results=3,
        highlights={"num_sentences": 2},
    )
    print(f"Articles similar to: {source_url}\n")
    for r in similar.results:
        print(f"- {r.title}")
        print(f"  {r.url}\n")

에이전트 도구 정의 (Define the Agent Tools)

이제 에이전트가 사용할 두 가지 도구 함수를 정의해요. 사내 KB 도구는 벡터 유사도를 사용해 Milvus를 검색하고, 웹 도구는 Exa를 통해 공개 인터넷을 검색해요.

def search_private_kb(query: str) -> str:
    """Search the internal knowledge base using Milvus vector search."""
    results = milvus.search(
        collection_name=COLLECTION,
        data=[embed_text(query)],
        limit=3,
        output_fields=["text", "source"],
    )
    chunks = []
    for hit in results[0]:
        chunks.append(f"[{hit['entity']['source']}] {hit['entity']['text']}")
    return "\n\n".join(chunks) if chunks else "No relevant internal documents found."

def search_web(query: str) -> str:
    """Search the public web using Exa for up-to-date information."""
    results = exa.search_and_contents(
        query=query,
        type="auto",
        num_results=3,
        highlights={"num_sentences": 3},
    )
    items = []
    for r in results.results:
        highlight = r.highlights[0] if r.highlights else "No snippet available."
        items.append(f"[{r.title}]({r.url})\n{highlight}")
    return "\n\n".join(items) if items else "No web results found."

TOOL_FNS = {
    "search_private_kb": search_private_kb,
    "search_web": search_web,
}

에이전트 구축 (Build the Agent)

에이전트는 OpenAI의 function calling을 사용해 어떤 도구를 호출할지 결정해요. 간단한 루프를 따르는데, LLM이 사용자 쿼리를 받고, 호출할 도구(있다면)를 결정하고, 실행한 뒤 검색된 컨텍스트에서 최종 답변을 종합해요.

TOOLS = [
    {
        "type": "function",
        "function": {
            "name": "search_private_kb",
            "description": (
                "Search the company's internal knowledge base (product docs, "
                "policies, earnings, API docs, HR guides). Use this for any "
                "question about internal/proprietary information."
            ),
            "parameters": {
                "type": "object",
                "properties": {
                    "query": {"type": "string", "description": "The search query"}
                },
                "required": ["query"],
            },
        },
    },
    {
        "type": "function",
        "function": {
            "name": "search_web",
            "description": (
                "Search the public web for up-to-date external information - "
                "news, trends, competitor analysis, open-source projects, etc. "
                "Use this when the question is about the outside world."
            ),
            "parameters": {
                "type": "object",
                "properties": {
                    "query": {"type": "string", "description": "The search query"}
                },
                "required": ["query"],
            },
        },
    },
]

SYSTEM_PROMPT = """You are a helpful assistant with access to two search tools:

1. **search_private_kb** - searches the company's internal knowledge base.
2. **search_web** - searches the public internet via Exa.

Routing rules:
- Questions about internal products, policies, metrics, or processes: use search_private_kb.
- Questions about external trends, news, competitors, or general knowledge: use search_web.
- Questions that need both internal and external context: call BOTH tools, then synthesize.

Always cite your sources. For internal docs, mention the filename. For web results, include the URL."""

def run_agent(user_query: str) -> str:
    """Run the agent loop: LLM -> tool calls -> LLM -> final answer."""
    messages = [
        {"role": "system", "content": SYSTEM_PROMPT},
        {"role": "user", "content": user_query},
    ]

    print(f"User: {user_query}\n")

    # First LLM call - may request tool calls
    response = llm.chat.completions.create(
        model="gpt-4o",
        messages=messages,
        tools=TOOLS,
    )
    msg = response.choices[0].message
    messages.append(msg)

    # If no tool calls, return directly
    if not msg.tool_calls:
        print(f"Agent (no tools used): {msg.content}")
        return msg.content

    # Execute each tool call
    for tc in msg.tool_calls:
        fn_name = tc.function.name
        fn_args = json.loads(tc.function.arguments)
        print(f"  -> Calling {fn_name}(query={fn_args['query']!r})")

        result = TOOL_FNS[fn_name](**fn_args)
        messages.append(
            {
                "role": "tool",
                "tool_call_id": tc.id,
                "content": result,
            }
        )

    # Second LLM call - synthesize final answer
    response = llm.chat.completions.create(
        model="gpt-4o",
        messages=messages,
        tools=TOOLS,
    )
    answer = response.choices[0].message.content
    print(f"\nAgent:\n{answer}")
    return answer

데모 (Demo)

이제 서로 다른 라우팅 동작을 보여 주는 세 가지 시나리오로 에이전트를 테스트해 보아요.

시나리오 A: 내부 질문 (routes to Milvus)

내부 정책에 대해 질문해요. 에이전트는 search_private_kb를 호출하고 사내 문서에서 답변을 가져와야 해요.

run_agent("What is the return policy for Acme products?")

시나리오 B: 외부 질문 (routes to Exa)

외부 트렌드에 대해 질문해요. 에이전트는 search_web을 호출해 공개 인터넷에서 최신 정보를 가져와야 해요.

run_agent("What are the latest AI agent frameworks trending in 2026?")

시나리오 C: 하이브리드 질문 (routes to both)

내부 사양과 외부 벤치마크를 모두 필요로 하는 질문을 해요. 에이전트는 두 도구를 모두 호출해 비교를 종합해야 해요.

run_agent(
    "How does our Widget Pro's throughput compare to "
    "open-source alternatives on the market?"
)

정리 (Cleanup)

작업이 끝나면 리소스를 해제하기 위해 컬렉션을 삭제해요.

milvus.drop_collection(COLLECTION)

결론 (Conclusion)

이 튜토리얼에서는 사내 지식 검색에 Milvus를, 공개 웹 검색에 Exa를 결합한 이중 소스 RAG 에이전트를 구축했어요. 핵심 구성 요소는 다음과 같아요.

  • Milvus는 벡터 유사도 검색으로 내부 문서를 저장하고 검색해, 독점 데이터가 사적이고 검색 가능하도록 보장해요.
  • Exa는 카테고리 필터링, 콘텐츠 추출, 유사 기사 발견 같은 기능을 갖춘 시맨틱 웹 검색을 제공해요.
  • OpenAI 함수 호출은 LLM이 질문의 의도에 따라 쿼리를 올바른 소스(또는 둘 다)로 자동 라우팅하게 해줘요.

이 패턴은 AI 어시스턴트가 기밀 내부 문서와 실시간 외부 정보에 모두 접근해야 하는 엔터프라이즈 사용 사례에 적용할 수 있어요.

더 알아보기 (Learn more)