Chuyển đến nội dung chính

レッスン 10: マルチモーダル RAG — ドキュメント内の画像、表、グラフ

画像、表、グラフを含むドキュメントの場合は RAG。 PDF スキャン、OCR、表抽出から情報を抽出します。マルチモーダルの Vision LLM + ベクトル検索。

🧠 AI と ML — レッスン 9 レッスン 10: マルチモーダル RAG — 画像、表、チャート ドキュメント内のマップ

リアルバトルRAG:基礎から上級まで

パート 4: 高度な RAG パターン

xdev.asia

はじめに

実際のドキュメントはテキストだけではなく、画像、表、チャート、図もあります。従来の RAG はすべてを省略します。マルチモーダル RAG はこの問題を解決します。

例: 50 ページの財務報告書: テキストが 40%、データ表が 30%、グラフが 20%、画像が 10%。 RAG テキストのみでは情報の 60% が欠落しています。

この記事の内容は次のとおりです。

  1. テーブル抽出 — テーブルの抽出とインデックス付け
  2. 画像の理解 — Vision LLM を使用して画像/グラフを記述する
  3. マルチモーダル埋め込み — テキストと画像の両方を同じベクトル空間に埋め込みます

1. 問題: マルチモーダルなドキュメント

1.1 ドキュメント内のコンテンツの種類

┌─────────────────────────────────────────┐
│  Typical Business Document              │
│                                          │
│  [Text paragraph]                        │  ← RAG text OK
│  [Text paragraph]                        │  ← RAG text OK
│                                          │
│  ┌────────────────────────────┐          │
│  │  Revenue  │ Q1  │ Q2  │ Q3│          │  ← RAG text BỎ QUA!
│  │  Product A│ 100 │ 120 │ 95│          │
│  │  Product B│ 200 │ 180 │ 220│         │
│  └────────────────────────────┘          │
│                                          │
│  [Bar chart: Revenue trends]  📊        │  ← RAG text BỎ QUA!
│                                          │
│  [Architecture diagram]       🖼️        │  ← RAG text BỎ QUA!
│                                          │
└─────────────────────────────────────────┘

1.2 処理戦略

コンテンツタイプ戦略ツール
テキストチャンクスライブLangChain スプリッター
表抽出 → テキスト/マークダウンに変換非構造化、キャメロット
チャート/図Vision LLM → テキスト説明GPT-4o、クロード
スキャンされた PDFOCR → テキストTesseract、Azure OCR

2. テーブルの抽出

2.1 非構造化の使用

"""Extract tables từ PDF bằng Unstructured"""
from unstructured.partition.pdf import partition_pdf

elements = partition_pdf(
    filename="financial-report.pdf",
    strategy="hi_res",           # Dùng model detection
    infer_table_structure=True,  # Detect và extract tables
    extract_images_in_pdf=True,  # Extract images
)

# Phân loại elements
tables = []
texts = []
images = []

for el in elements:
    if el.category == "Table":
        tables.append(el)
        print(f"Table found: {el.metadata.text_as_html[:200]}...")
    elif el.category == "Image":
        images.append(el)
    else:
        texts.append(el)

print(f"Found: {len(texts)} texts, {len(tables)} tables, {len(images)} images")

2.2 表→テキストの要約

"""Dùng LLM summarize bảng thành text cho RAG"""
from langchain_openai import ChatOpenAI

llm = ChatOpenAI(model="gpt-4o-mini", temperature=0)

def summarize_table(table_html: str) -> str:
    prompt = f"""Đây là bảng dữ liệu (HTML):
{table_html}

Tóm tắt nội dung bảng thành đoạn văn (50-100 từ).
Bao gồm: tên bảng, các cột, xu hướng nổi bật, giá trị đặc biệt."""
    
    return llm.invoke(prompt).content

# Tạo document cho mỗi table
from langchain.schema import Document

table_docs = []
for table in tables:
    summary = summarize_table(table.metadata.text_as_html)
    table_docs.append(Document(
        page_content=summary,
        metadata={
            "source": "financial-report.pdf",
            "type": "table",
            "original_html": table.metadata.text_as_html,
            "page": table.metadata.page_number,
        }
    ))

2.3 マルチベクトル: サマリー データと生データの両方を保存する

"""Multi-vector store: search bằng summary, trả về raw table"""
from langchain.storage import InMemoryByteStore
from langchain.retrievers.multi_vector import MultiVectorRetriever
from langchain_community.vectorstores import Chroma
from langchain_openai import OpenAIEmbeddings
import uuid

# Vector store: chứa summaries (để search)
vectorstore = Chroma(
    collection_name="multimodal",
    embedding_function=OpenAIEmbeddings(),
)

# Doc store: chứa raw data (để trả về cho LLM)
docstore = InMemoryByteStore()

retriever = MultiVectorRetriever(
    vectorstore=vectorstore,
    byte_store=docstore,
    id_key="doc_id",
)

# Index: summary → vector store, raw → doc store
for table in tables:
    doc_id = str(uuid.uuid4())
    summary = summarize_table(table.metadata.text_as_html)
    
    # Summary vào vector store (search)
    retriever.vectorstore.add_documents([
        Document(page_content=summary, metadata={"doc_id": doc_id, "type": "table"})
    ])
    
    # Raw table vào doc store (return)
    retriever.docstore.mset([(doc_id, table.metadata.text_as_html)])

💡 演習 1: 少なくとも 3 つの表を含む PDF から表を抽出します。マルチベクトル ストアの作成: サマリーを使用して検索し、生のテーブルを返します。テーブル データに関連する 5 つの質問をテストします。


3. 画像の理解

3.1 Vision LLM イメージの説明

"""Dùng GPT-4o mô tả ảnh/biểu đồ trong tài liệu"""
import base64
from langchain_openai import ChatOpenAI

llm = ChatOpenAI(model="gpt-4o", temperature=0)

def describe_image(image_path: str) -> str:
    with open(image_path, "rb") as f:
        image_data = base64.b64encode(f.read()).decode()
    
    response = llm.invoke([
        {"role": "system", "content": "Mô tả chi tiết nội dung ảnh/biểu đồ. "
         "Nếu là biểu đồ: liệt kê data points, xu hướng, kết luận."},
        {"role": "user", "content": [
            {"type": "text", "text": "Mô tả ảnh này:"},
            {"type": "image_url", "image_url": {
                "url": f"data:image/png;base64,{image_data}"
            }},
        ]},
    ])
    return response.content

# Mô tả biểu đồ revenue
desc = describe_image("charts/revenue-q3.png")
# "Biểu đồ cột so sánh doanh thu Q1-Q3 2024.
#  Product A: giảm 20% từ Q1 (100M) xuống Q3 (80M).
#  Product B: tăng 10% ổn định, đạt 220M Q3.
#  Tổng doanh thu Q3: 300M, tăng 5% so với Q2..."

3.2 パイプライン: PDF → 画像の抽出 → 説明 → インデックス

"""Full pipeline cho multimodal PDF"""
import os

def process_multimodal_pdf(pdf_path: str, output_dir: str):
    # 1. Extract elements
    elements = partition_pdf(
        filename=pdf_path,
        strategy="hi_res",
        extract_images_in_pdf=True,
        image_output_dir_path=output_dir,
    )
    
    all_docs = []
    
    for el in elements:
        if el.category == "Table":
            # Summarize table
            summary = summarize_table(el.metadata.text_as_html)
            all_docs.append(Document(
                page_content=summary,
                metadata={"type": "table", "page": el.metadata.page_number}
            ))
        elif el.category == "Image":
            # Describe image
            img_path = os.path.join(output_dir, el.metadata.image_path)
            description = describe_image(img_path)
            all_docs.append(Document(
                page_content=description,
                metadata={"type": "image", "page": el.metadata.page_number}
            ))
        else:
            all_docs.append(Document(
                page_content=str(el),
                metadata={"type": "text", "page": el.metadata.page_number}
            ))
    
    return all_docs

docs = process_multimodal_pdf("report.pdf", "./extracted_images")
# Index all_docs vào vector store → search bình thường!

4. マルチモーダルな埋め込み

4.1 CLIP ベース: ベクトル空間を使用したテキスト + 画像

"""Embed text và ảnh vào cùng vector space"""
from langchain_experimental.open_clip import OpenCLIPEmbeddings

# CLIP embeddings: text và image → cùng 1 vector space
clip_embeddings = OpenCLIPEmbeddings(
    model_name="ViT-B-32",
    checkpoint="openai",
)

# Embed text
text_emb = clip_embeddings.embed_documents(["biểu đồ doanh thu tăng"])

# Embed image
img_emb = clip_embeddings.embed_image(["charts/revenue.png"])

# Cả 2 vectors có thể so sánh cosine similarity!
# → Search bằng text, tìm được ảnh liên quan

4.2 いつどのアプローチを使用するか?

アプローチ利点デメリット使用例
ビジョン LLM → テキスト柔軟、詳細高価な API、遅いチャート、ダイアグラム
OCR → テキスト早い、安い画像内のテキストを読み取り専用スキャンされたドキュメント
CLIP 埋め込み直接検索詳細は少し画像検索
マルチベクトル両方の長所複雑なセットアップ制作

💡 演習 2: テキスト + 表 + 画像を含む PDF レポート用のマルチモーダル RAG を作成します。テストは、(a) テキストについて、(b) 表データについて、(c) グラフの内容についての質問に答えます。


5. スキャンした PDF の処理 (OCR)

5.1 OCR パイプライン

"""OCR cho PDF scan — không có text layer"""
from unstructured.partition.pdf import partition_pdf

# strategy="ocr_only" cho scanned PDFs
elements = partition_pdf(
    filename="scanned-contract.pdf",
    strategy="ocr_only",
    languages=["vie", "eng"],   # Hỗ trợ tiếng Việt
    ocr_languages="vie+eng",
)

# Elements đã được OCR → có text content
for el in elements:
    print(el.text[:100])

5.2 OCR 品質の向上

Kết quả OCR thô: "Điều 5. Quvền vá nghia vụ cùa người lao dộng"
                                 ↑ sai     ↑ sai         ↑ sai

Post-processing bằng LLM:
"Điều 5. Quyền và nghĩa vụ của người lao động"
→ LLM fix lỗi OCR dựa trên context!
"""LLM post-process OCR text"""
def fix_ocr_text(raw_text: str) -> str:
    prompt = f"""Text sau được OCR từ tài liệu tiếng Việt, có thể có lỗi.
Sửa lỗi chính tả, giữ nguyên nội dung:

{raw_text}

Text đã sửa:"""
    return llm.invoke(prompt).content

概要

コンセプト覚えておいてください
マルチモーダル RAGテキスト + 表 + 画像 + グラフの RAG
テーブル抽出非構造化高解像度 → HTML → LLM の概要
画像の説明Vision LLM (GPT-4o) は画像をテキストに記述します
マルチベクトル概要を使用して検索し、生データを返す
クリップベクトル空間でテキスト + 画像を埋め込む
OCRスキャンされた PDF → テキスト、LLM 修正エラー

一般的な演習

  1. ✅ 2 つの小さな演習 (1、2) を完了します。
  2. 完全なマルチモーダル パイプライン: 1 つの複雑な PDF (年次報告書) を処理: テキスト + 表 + グラフを抽出 → すべてにインデックスを作成 → あらゆる種類の質問に答える Q&A チャットボットを構築。
  3. マルチベクター ストア: Chroma + InMemoryByteStore を使用して実装します。サマリーを使用して検索し、生の HTML テーブルを返します。回答の品質をテキストのみの RAG と比較します。
  4. OCR パイプライン: プロセス 5 でスキャンされたベトナム語 PDF → OCR → LLM 修正 → インデックス。 10 問の正解率を測定します。

次の記事: Agentic RAG — エージェント + RAG がパワーを結合します — RAG パイプラインが自身で決定する必要がある場合: どこを検索するか、どのような追加情報が必要か、いつ停止するか。