Mistral OCR로 PDF에서 텍스트와 이미지를 추출하는 Gradio 앱 구축하기

Mistral OCR로 PDF에서 텍스트와 이미지를 추출하는 Gradio 앱 구축하기 (Building Gradio app with Mistral OCR)

Mistral OCR을 Gradio와 함께 설정하고 사용하는 단계별 가이드예요. PDF와 이미지에서 텍스트와 이미지를 추출하는 애플리케이션을 만듭니다.

출처: 문서

본문

이 쿡북은 Mistral OCR과 Gradio를 설정하고 사용하는 단계별 안내를 제공해요. 이 애플리케이션을 통해 Mistral의 OCR 기능으로 PDF와 이미지에서 텍스트와 이미지를 추출할 수 있어요.

사전 요구사항 (Prerequisites)

시작 전에 다음이 준비되어 있는지 확인하세요:

  • 시스템에 Python 설치
  • Mistral AI의 API 키
  • 필요한 Python 패키지 설치

Step 1: 필요한 패키지 설치

먼저 pip로 필요한 Python 패키지를 설치해요.

pip install gradio requests mistralai

Step 2: 환경 변수 설정

Mistral API 키를 환경 변수로 설정해야 해요. 터미널에서 하거나 스크립트에 추가할 수 있어요. API 키는 [Platforme]에서 만들 수 있어요.

import os

os.environ["MISTRAL_API_KEY"] = "your_mistral_api_key_here"

Step 3: 라이브러리 임포트

Python 스크립트에서 필요한 라이브러리를 임포트해요.

import gradio as gr
import os
import base64
import requests
from mistralai.client import Mistral

Step 4: Mistral 클라이언트 초기화

API 키로 Mistral 클라이언트를 초기화해요.

api_key = os.environ["MISTRAL_API_KEY"]
client = Mistral(api_key=api_key)

Step 5: 헬퍼 함수 정의

이미지를 base64로 인코딩

이 함수는 이미지를 base64 문자열로 인코딩해요. 로컬 이미지를 서비스에 제공하려면 필요해요.

def encode_image(image_path):
    """Encode the image to base64."""
    try:
        with open(image_path, "rb") as image_file:
            return base64.b64encode(image_file.read()).decode('utf-8')
    except FileNotFoundError:
        return "Error: The file was not found."
    except Exception as e:
        return f"Error: {e}"

마크다운에서 이미지 교체

이 함수는 마크다운의 이미지 플레이스홀더를 base64 인코딩된 이미지로 교체해요. Mistral OCR은 텍스트와 이미지를 interleave하여 출력할 수 있는데, 이 함수가 플레이스홀더를 교체해 Gradio로 렌더링할 수 있게 해줘요.

def replace_images_in_markdown(markdown_str: str, images_dict: dict) -> str:
    for img_name, base64_str in images_dict.items():
        markdown_str = markdown_str.replace(f"![{img_name}]({img_name})", f"![{img_name}]({base64_str})")
    return markdown_str

결합된 마크다운 얻기

이 함수는 OCR 응답의 모든 페이지에서 마크다운을 결합해요. 렌더링 가능한 버전과 이미지가 없는 raw 마크다운 출력을 반환합니다.

def get_combined_markdown(ocr_response) -> tuple:
    markdowns = []
    raw_markdowns = []
    for page in ocr_response.pages:
        image_data = {}
        for img in page.images:
            image_data[img.id] = img.image_base64
        markdowns.append(replace_images_in_markdown(page.markdown, image_data))
        raw_markdowns.append(page.markdown)
    return "\n\n".join(markdowns), "\n\n".join(raw_markdowns)

콘텐츠 타입 가져오기

이 함수는 URL의 콘텐츠 타입을 가져와요. PDF 파일인지 이미지 파일인지 감지하는 것이 목적이에요.

def get_content_type(url):
    """Fetch the content type of the URL."""
    try:
        response = requests.head(url)
        return response.headers.get('Content-Type')
    except Exception as e:
        return f"Error fetching content type: {e}"

Step 6: OCR 함수 정의

파일에 대해 OCR 수행

로컬 PDF 파일에 OCR을 수행하려면 먼저 Platforme에 업로드하고 OCR 작업에 사용할 서명 URL(signed URL)을 얻어야 해요.

def perform_ocr_file(file, ocr_method="Mistral OCR"):
    if ocr_method == "Mistral OCR":
        if file.name.endswith('.pdf'):
            uploaded_pdf = client.files.upload(
                file={
                    "file_name": file.name,
                    "content": open(file.name, "rb"),
                },
                purpose="ocr"
            )
            signed_url = client.files.get_signed_url(file_id=uploaded_pdf.id)
            ocr_response = client.ocr.process(
                model="mistral-ocr-latest",
                document={
                    "type": "document_url",
                    "document_url": signed_url.url,
                },
                include_image_base64=True
            )
            client.files.delete(file_id=uploaded_pdf.id)

        elif file.name.endswith(('.png', '.jpg', '.jpeg')):
            base64_image = encode_image(file.name)
            ocr_response = client.ocr.process(
                model="mistral-ocr-latest",
                document={
                    "type": "image_url",
                    "image_url": f"data:image/jpeg;base64,{base64_image}"
                },
                include_image_base64=True
            )

        combined_markdown, raw_markdown = get_combined_markdown(ocr_response)
        return combined_markdown, raw_markdown

    return "## Method not supported.", "Method not supported."

URL에 대해 OCR 수행

다음으로 URL에 대해 OCR을 수행하는 함수를 정의할 수 있어요. 이미지와 PDF 문서에 대해 다른 함수가 필요해요.

def perform_ocr_url(url, ocr_method="Mistral OCR"):
    if ocr_method == "Mistral OCR":
        content_type = get_content_type(url)
        if 'application/pdf' in content_type:
            ocr_response = client.ocr.process(
                model="mistral-ocr-latest",
                document={
                    "type": "document_url",
                    "document_url": url,
                },
                include_image_base64=True
            )

        elif any(image_type in content_type for image_type in ['image/png', 'image/jpeg', 'image/jpg']):
            ocr_response = client.ocr.process(
                model="mistral-ocr-latest",
                document={
                    "type": "image_url",
                    "image_url": url,
                },
                include_image_base64=True
            )
        else:
            return "Unsupported file type. Please provide a URL to a PDF or an image.", ""

        combined_markdown, raw_markdown = get_combined_markdown(ocr_response)
        return combined_markdown, raw_markdown

    return "## Method not supported.", "Method not supported."

Step 7: Gradio 인터페이스 생성

마지막으로 OCR 함수와 상호작용하는 Gradio 인터페이스를 만들 수 있어요! "Upload File"과 "Enter URL" 두 탭을 제공합니다.

with gr.Blocks() as demo:
    gr.Markdown("# Mistral OCR")
    gr.Markdown("Upload a PDF or an image, or provide a URL to extract text and images using Mistral OCR capabilities.\n\nLearn more in the blog post [here](https://mistral.ai/news/mistral-ocr).")

    with gr.Tab("Upload File"):
        file_input = gr.File(label="Upload a PDF or Image")
        ocr_method_file = gr.Dropdown(choices=["Mistral OCR"], label="Select OCR Method", value="Mistral OCR")
        file_output = gr.Markdown(label="Rendered Markdown")
        file_raw_output = gr.Textbox(label="Raw Markdown")
        file_button = gr.Button("Process")

        example_files = gr.Examples(
            examples=[
                "pixtral-12b.pdf",
                "receipt.png"
            ],
            inputs=[file_input]
        )

        file_button.click(
            fn=perform_ocr_file,
            inputs=[file_input, ocr_method_file],
            outputs=[file_output, file_raw_output]
        )

    with gr.Tab("Enter URL"):
        url_input = gr.Textbox(label="Enter a URL to a PDF or Image")
        ocr_method_url = gr.Dropdown(choices=["Mistral OCR"], label="Select OCR Method", value="Mistral OCR")
        url_output = gr.Markdown(label="Rendered Markdown")
        url_raw_output = gr.Textbox(label="Raw Markdown")
        url_button = gr.Button("Process")

        example_urls = gr.Examples(
            examples=[
                "https://arxiv.org/pdf/2410.07073",
                "https://raw.githubusercontent.com/mistralai/cookbook/refs/heads/main/mistral/ocr/receipt.png"
            ],
            inputs=[url_input]
        )

        url_button.click(
            fn=perform_ocr_url,
            inputs=[url_input, ocr_method_url],
            outputs=[url_output, url_raw_output]
        )

demo.launch(max_threads=1)

Step 8: 애플리케이션 실행

스크립트를 실행해 Gradio 인터페이스를 띄워요. 인터페이스와 상호작용해 파일이나 URL에 대해 OCR을 수행할 수 있습니다.

python your_script_name.py

라이브 데모는 [여기]에서 볼 수 있고, 더 많은 정보는 [블로그 포스트]에서 찾을 수 있어요.

더 알아보기 (Learn more)