Mistral OCR로 PDF에서 텍스트와 이미지를 추출하는 Gradio 앱 구축하기
Mistral OCR로 PDF에서 텍스트와 이미지를 추출하는 Gradio 앱 구축하기 (Building Gradio app with Mistral OCR)
Mistral OCR을 Gradio와 함께 설정하고 사용하는 단계별 가이드예요. PDF와 이미지에서 텍스트와 이미지를 추출하는 애플리케이션을 만듭니다.
출처: 문서
본문
이 쿡북은 Mistral OCR과 Gradio를 설정하고 사용하는 단계별 안내를 제공해요. 이 애플리케이션을 통해 Mistral의 OCR 기능으로 PDF와 이미지에서 텍스트와 이미지를 추출할 수 있어요.
사전 요구사항 (Prerequisites)
시작 전에 다음이 준비되어 있는지 확인하세요:
- 시스템에 Python 설치
- Mistral AI의 API 키
- 필요한 Python 패키지 설치
Step 1: 필요한 패키지 설치
먼저 pip로 필요한 Python 패키지를 설치해요.
pip install gradio requests mistralai
Step 2: 환경 변수 설정
Mistral API 키를 환경 변수로 설정해야 해요. 터미널에서 하거나 스크립트에 추가할 수 있어요. API 키는 [Platforme]에서 만들 수 있어요.
import os
os.environ["MISTRAL_API_KEY"] = "your_mistral_api_key_here"
Step 3: 라이브러리 임포트
Python 스크립트에서 필요한 라이브러리를 임포트해요.
import gradio as gr
import os
import base64
import requests
from mistralai.client import Mistral
Step 4: Mistral 클라이언트 초기화
API 키로 Mistral 클라이언트를 초기화해요.
api_key = os.environ["MISTRAL_API_KEY"]
client = Mistral(api_key=api_key)
Step 5: 헬퍼 함수 정의
이미지를 base64로 인코딩
이 함수는 이미지를 base64 문자열로 인코딩해요. 로컬 이미지를 서비스에 제공하려면 필요해요.
def encode_image(image_path):
"""Encode the image to base64."""
try:
with open(image_path, "rb") as image_file:
return base64.b64encode(image_file.read()).decode('utf-8')
except FileNotFoundError:
return "Error: The file was not found."
except Exception as e:
return f"Error: {e}"
마크다운에서 이미지 교체
이 함수는 마크다운의 이미지 플레이스홀더를 base64 인코딩된 이미지로 교체해요. Mistral OCR은 텍스트와 이미지를 interleave하여 출력할 수 있는데, 이 함수가 플레이스홀더를 교체해 Gradio로 렌더링할 수 있게 해줘요.
def replace_images_in_markdown(markdown_str: str, images_dict: dict) -> str:
for img_name, base64_str in images_dict.items():
markdown_str = markdown_str.replace(f"", f"")
return markdown_str
결합된 마크다운 얻기
이 함수는 OCR 응답의 모든 페이지에서 마크다운을 결합해요. 렌더링 가능한 버전과 이미지가 없는 raw 마크다운 출력을 반환합니다.
def get_combined_markdown(ocr_response) -> tuple:
markdowns = []
raw_markdowns = []
for page in ocr_response.pages:
image_data = {}
for img in page.images:
image_data[img.id] = img.image_base64
markdowns.append(replace_images_in_markdown(page.markdown, image_data))
raw_markdowns.append(page.markdown)
return "\n\n".join(markdowns), "\n\n".join(raw_markdowns)
콘텐츠 타입 가져오기
이 함수는 URL의 콘텐츠 타입을 가져와요. PDF 파일인지 이미지 파일인지 감지하는 것이 목적이에요.
def get_content_type(url):
"""Fetch the content type of the URL."""
try:
response = requests.head(url)
return response.headers.get('Content-Type')
except Exception as e:
return f"Error fetching content type: {e}"
Step 6: OCR 함수 정의
파일에 대해 OCR 수행
로컬 PDF 파일에 OCR을 수행하려면 먼저 Platforme에 업로드하고 OCR 작업에 사용할 서명 URL(signed URL)을 얻어야 해요.
def perform_ocr_file(file, ocr_method="Mistral OCR"):
if ocr_method == "Mistral OCR":
if file.name.endswith('.pdf'):
uploaded_pdf = client.files.upload(
file={
"file_name": file.name,
"content": open(file.name, "rb"),
},
purpose="ocr"
)
signed_url = client.files.get_signed_url(file_id=uploaded_pdf.id)
ocr_response = client.ocr.process(
model="mistral-ocr-latest",
document={
"type": "document_url",
"document_url": signed_url.url,
},
include_image_base64=True
)
client.files.delete(file_id=uploaded_pdf.id)
elif file.name.endswith(('.png', '.jpg', '.jpeg')):
base64_image = encode_image(file.name)
ocr_response = client.ocr.process(
model="mistral-ocr-latest",
document={
"type": "image_url",
"image_url": f"data:image/jpeg;base64,{base64_image}"
},
include_image_base64=True
)
combined_markdown, raw_markdown = get_combined_markdown(ocr_response)
return combined_markdown, raw_markdown
return "## Method not supported.", "Method not supported."
URL에 대해 OCR 수행
다음으로 URL에 대해 OCR을 수행하는 함수를 정의할 수 있어요. 이미지와 PDF 문서에 대해 다른 함수가 필요해요.
def perform_ocr_url(url, ocr_method="Mistral OCR"):
if ocr_method == "Mistral OCR":
content_type = get_content_type(url)
if 'application/pdf' in content_type:
ocr_response = client.ocr.process(
model="mistral-ocr-latest",
document={
"type": "document_url",
"document_url": url,
},
include_image_base64=True
)
elif any(image_type in content_type for image_type in ['image/png', 'image/jpeg', 'image/jpg']):
ocr_response = client.ocr.process(
model="mistral-ocr-latest",
document={
"type": "image_url",
"image_url": url,
},
include_image_base64=True
)
else:
return "Unsupported file type. Please provide a URL to a PDF or an image.", ""
combined_markdown, raw_markdown = get_combined_markdown(ocr_response)
return combined_markdown, raw_markdown
return "## Method not supported.", "Method not supported."
Step 7: Gradio 인터페이스 생성
마지막으로 OCR 함수와 상호작용하는 Gradio 인터페이스를 만들 수 있어요! "Upload File"과 "Enter URL" 두 탭을 제공합니다.
with gr.Blocks() as demo:
gr.Markdown("# Mistral OCR")
gr.Markdown("Upload a PDF or an image, or provide a URL to extract text and images using Mistral OCR capabilities.\n\nLearn more in the blog post [here](https://mistral.ai/news/mistral-ocr).")
with gr.Tab("Upload File"):
file_input = gr.File(label="Upload a PDF or Image")
ocr_method_file = gr.Dropdown(choices=["Mistral OCR"], label="Select OCR Method", value="Mistral OCR")
file_output = gr.Markdown(label="Rendered Markdown")
file_raw_output = gr.Textbox(label="Raw Markdown")
file_button = gr.Button("Process")
example_files = gr.Examples(
examples=[
"pixtral-12b.pdf",
"receipt.png"
],
inputs=[file_input]
)
file_button.click(
fn=perform_ocr_file,
inputs=[file_input, ocr_method_file],
outputs=[file_output, file_raw_output]
)
with gr.Tab("Enter URL"):
url_input = gr.Textbox(label="Enter a URL to a PDF or Image")
ocr_method_url = gr.Dropdown(choices=["Mistral OCR"], label="Select OCR Method", value="Mistral OCR")
url_output = gr.Markdown(label="Rendered Markdown")
url_raw_output = gr.Textbox(label="Raw Markdown")
url_button = gr.Button("Process")
example_urls = gr.Examples(
examples=[
"https://arxiv.org/pdf/2410.07073",
"https://raw.githubusercontent.com/mistralai/cookbook/refs/heads/main/mistral/ocr/receipt.png"
],
inputs=[url_input]
)
url_button.click(
fn=perform_ocr_url,
inputs=[url_input, ocr_method_url],
outputs=[url_output, url_raw_output]
)
demo.launch(max_threads=1)
Step 8: 애플리케이션 실행
스크립트를 실행해 Gradio 인터페이스를 띄워요. 인터페이스와 상호작용해 파일이나 URL에 대해 OCR을 수행할 수 있습니다.
python your_script_name.py
라이브 데모는 [여기]에서 볼 수 있고, 더 많은 정보는 [블로그 포스트]에서 찾을 수 있어요.