Mistral AI와 MongoDB로 RAG 구축하기

Mistral AI와 MongoDB로 RAG 구축하기

Mistral AI와 기업급 벡터 스토어 MongoDB를 결합해 LLM GenAI 애플리케이션을 만드는 단계별 가이드예요. PDF 문서를 임베딩해 MongoDB에 저장하고, 벡터 검색으로 유사 문서를 찾아 질의에 답합니다.

출처: 문서

본문

먼저 MISTRAL_API_KEY를 설정하고 가입을 활성화하며, MongoDB Atlas 클러스터에서 MONGO_URI를 얻어야 해요.

export MONGO_URI="Your_cluster_connection_string"
export MISTRAL_API_KEY="Your_MISTRAL_API_KEY"

필요한 라이브러리 버전은 다음과 같습니다 (이 가이드를 작성한 시점 기준).

mistralai                                         0.0.8
pymongo                                           4.3.3
gradio                                            4.10.0
gradio_client                                     0.7.3
langchain                                         0.0.348
langchain-core                                    0.0.12
pandas                                            2.0.3
# Install necessary packages
!pip install mistralai==0.0.8
!pip install pymongo==4.3.3
!pip install gradio==4.10.0
!pip install gradio_client==0.7.3
!pip install langchain==0.0.348
!pip install langchain-core==0.0.12
!pip install pandas==2.0.3

이 라이브러리들은 데이터 처리, 웹 스크래핑, AI 모델, 데이터베이스 상호작용에 사용됩니다.

import gradio as gr
import os
import pymongo
import pandas as pd
from mistralai.client import MistralClient
from mistralai.models.chat_completion import ChatMessage
from langchain.document_loaders import PyPDFLoader
from langchain.text_splitter import RecursiveCharacterTextSplitter

셸 명령으로 내보낸 API 키를 사용할 수 있어요.

# Check API keys
import os
mistral_api_key = os.environ["MISTRAL_API_KEY"]
mongo_url = os.environ["MONGO_URI"]

데이터 준비 (Data preparation)

data_prep() 함수는 PDF·문서·URL에서 데이터를 로드합니다. 웹페이지/문서에서 텍스트를 추출하고 불필요한 요소를 제거한 뒤, 관리하기 좋은 청크로 나눕니다. 청크가 준비되면 Mistral AI 임베딩 엔드포인트로 각 청크의 임베딩을 계산해 문서에 저장하고, 각 문서를 MongoDB 컬렉션에 추가합니다.

def data_prep(file):
    # Set up Mistral client
    api_key = os.environ["MISTRAL_API_KEY"]
    client = MistralClient(api_key=api_key)

    # Process the uploaded file
    loader = PyPDFLoader(file.name)
    pages = loader.load_and_split()

    # Split data
    text_splitter = RecursiveCharacterTextSplitter(
        chunk_size=100,
        chunk_overlap=20,
        separators=["\n\n", "\n", "(?<=\. )", " ", ""],
        length_function=len,
    )
    docs = text_splitter.split_documents(pages)

    # Calculate embeddings and store into MongoDB
    text_chunks = [text.page_content for text in docs]
    df = pd.DataFrame({'text_chunks': text_chunks})
    df['embedding'] = df.text_chunks.apply(lambda x: get_embedding(x, client))

    collection = connect_mongodb()
    df_dict = df.to_dict(orient='records')
    collection.insert_many(df_dict)

    return "PDF processed and data stored in MongoDB."

MongoDB 서버 연결하기

connect_mongodb() 함수는 MongoDB 서버에 연결해 데이터베이스와 상호작용할 컬렉션 객체를 반환합니다. MongoDB 연결 문자열은 Atlas 콘솔에서 클러스터의 "Connect" 버튼을 누르고 Python 드라이버를 선택하면 얻을 수 있어요.

def connect_mongodb(mongo_url):
    # Your MongoDB connection string
    client = pymongo.MongoClient(mongo_url)
    db = client["mistralpdf"]
    collection = db["pdfRAG"]
    return collection

임베딩 얻기

get_embedding() 함수는 주어진 텍스트의 임베딩을 생성합니다. 새 줄 문자를 공백으로 바꾼 뒤 mistral-embed 엔드포인트로 임베딩을 얻어요. 데이터 준비와 질의응답 과정 모두에서 호출됩니다.

def get_embedding(text, client):
    text = text.replace("\n", " ")
    embeddings_batch_response = client.embeddings(
        model="mistral-embed",
        input=text,
    )
    return embeddings_batch_response.data[0].embedding

MongoDB 벡터 검색 인덱스 구성

벡터 검색 쿼리를 실행하려면 MongoDB Atlas에 다음과 같이 벡터 검색 인덱스를 만들어야 합니다. 벡터 검색 인덱스 생성에 대한 더 자세한 내용은 MongoDB 문서에서 확인할 수 있어요.

{
  "type": "vectorSearch",
  "fields": [
    {
      "numDimensions": 1536,
      "path": "'embedding'",
      "similarity": "cosine",
      "type": "vector"
    }
  ]
}

유사 문서 찾기

find_similar_documents() 함수는 MongoDB 컬렉션에서 벡터 검색 쿼리를 실행합니다. 사용자가 질문할 때 호출돼, 질의응답 과정에서 질문과 유사한 문서를 찾는 데 사용됩니다.

def find_similar_documents(embedding):
    collection = connect_mongodb()
    documents = list(
        collection.aggregate([
            {
                "$vectorSearch": {
                    "index": "vector_index",
                    "path": "embedding",
                    "queryVector": embedding,
                    "numCandidates": 20,
                    "limit": 10
                }
            },
            {"$project": {"_id": 0, "text_chunks": 1}}
        ])
    )
    return documents

질의응답 함수 (Question and answer)

이 함수가 프로그램의 핵심입니다. 사용자 질문을 처리하고 Mistral AI가 공급하는 문맥을 사용해 응답을 만듭니다. 과정은 다음과 같아요.

  1. Mistral AI 임베딩 엔드포인트로 사용자 질문의 수치 표현(임베딩)을 생성합니다.
  2. MongoDB 컬렉션에서 사용자 질문과 유사한 문서를 벡터 검색으로 찾습니다.
  3. 이 유사 문서들의 텍스트 청크를 결합해 문맥적 배경을 구성합니다. 이 정보를 모두 묶어 어시스턴트 지시(instruction)를 준비합니다.
  4. 사용자 질문과 어시스턴트 지시를 Mistral AI 모델용 프롬프트로 조합합니다.
  5. 마지막으로 Mistral AI가 검색 증강 생성(RAG) 과정을 통해 사용자에게 응답을 생성합니다.
def qna(users_question):
    # Set up Mistral client
    api_key = os.environ["MISTRAL_API_KEY"]
    client = MistralClient(api_key=api_key)

    question_embedding = get_embedding(users_question, client)

    print("-----Here is user question------")
    print(users_question)

    documents = find_similar_documents(question_embedding)
    print("-----Retrieved documents------")
    print(documents)

    for doc in documents:
        doc['text_chunks'] = doc['text_chunks'].replace('\n', ' ')

    for document in documents:
        print(str(document) + "\n")

    context = " ".join([doc["text_chunks"] for doc in documents])

    template = f"""
    You are an expert who loves to help people! Given the following context sections, answer the
    question using only the given context. If you are unsure and the answer is not
    explicitly written in the documentation, say "Sorry, I don't know how to help with that."
    Context sections:
    {context}
    Question:
    {users_question}
    Answer:
    """

    messages = [ChatMessage(role="user", content=template)]

    chat_response = client.chat(
        model="mistral-large-latest",
        messages=messages,
    )

    formatted_documents = '\n'.join([doc['text_chunks'] for doc in documents])
    return chat_response.choices[0].message, formatted_documents

더 알아보기 (Learn more)

  • MongoDB Atlas Vector Search 문서 — 벡터 검색 인덱스 생성 가이드
  • mistral-embed — Mistral 임베딩 모델 (이 예제에서는 1536차원 지정)
  • mistral-large-latest — 질의응답에 사용한 LLM 모델
  • pymongo $vectorSearch 집계 파이프라인 — MongoDB 벡터 검색 연산자