Chuyển đến nội dung chính

レッスン 12: マルチモーダル AI — GPT-4o ビジョン、ジェミニ ビジョン

視覚言語モデル: GPT-4o、Gemini 1.5、Claude Vision。画像理解、チャート分析、ドキュメント QA のための API 統合。モデル間の精度を比較します。コストの最適化。

🧠 AI と ML — レッスン 11 レッスン 12: マルチモーダル AI — GPT-4o ビジョン、 ジェミニビジョン

深層学習によるコンピューター ビジョン: CNN から Vision Transformer まで

パート 4: 実際のアプリケーションと展開

xdev.asia

はじめに

マルチモーダル AI = モデルは テキストと画像/ビデオの両方を同時に理解します。 GPT-4o、Gemini 2.0、Claude 3.5 はすべて、自然言語で画像を見て質問に答える機能を備えています。トレーニングも複雑なコードも必要なく、API 呼び出しだけで済みます。

🎯 マルチモーダル AI は、多くの従来の CV パイプラインを置き換えています。 1 つの API 呼び出しが OCR + 分類 + 分析に置き換わります。


1. 視覚言語モデル (VLM)

1.1 比較表

モデルプロバイダー強み価格
GPT-4oオープンAIオールラウンドで強力な推論$2.50/1M 入力トークン
GPT-4o-miniオープンAI安くて早くて十分$0.15/1M 入力トークン
ジェミニ 2.0 フラッシュグーグル迅速、安価、ビデオサポート$0.075/1M 入力トークン
ジェミニ 1.5 プログーグル長いコンテキスト (2M)、ビデオ$1.25/1M 入力トークン
クロード 3.5 ソネット人類優れたチャート/図分析$3/1M 入力トークン
Qwen2-VLアリババオープンソース、良い無料 (自己ホスト型)

1.2 従来の CV ではなく VLM を使用する理由は何ですか?

Pipeline truyền thống:
  YOLO detect → OCR extract text → NLP classify → Custom logic → Output
  = Nhiều models, phức tạp, maintenance cao

VLM:
  Image + Question → GPT-4o → Answer
  = 1 API call, đơn giản, chính xác!

2. ハンズオン: OpenAI GPT-4o ビジョン

2.1 画像の基本的な理解

"""GPT-4o Vision — hiểu ảnh bằng 1 API call"""
import base64
from openai import OpenAI

client = OpenAI()

def encode_image(image_path):
    with open(image_path, "rb") as f:
        return base64.b64encode(f.read()).decode("utf-8")

# Gửi ảnh + câu hỏi
image_b64 = encode_image("street_vietnam.jpg")

response = client.chat.completions.create(
    model="gpt-4o",
    messages=[
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "Mô tả chi tiết bức ảnh này. Có những gì?"},
                {
                    "type": "image_url",
                    "image_url": {
                        "url": f"data:image/jpeg;base64,{image_b64}",
                        "detail": "high",  # "low" = rẻ hơn, "high" = chi tiết hơn
                    }
                }
            ],
        }
    ],
    max_tokens=1000,
)

print(response.choices[0].message.content)

2.2 構造化出力 — 情報の抽出

"""Extract thông tin có cấu trúc từ ảnh"""
from pydantic import BaseModel
from openai import OpenAI

client = OpenAI()

class InvoiceInfo(BaseModel):
    invoice_number: str
    date: str
    vendor_name: str
    items: list[dict]
    total_amount: float
    currency: str

image_b64 = encode_image("invoice.jpg")

response = client.beta.chat.completions.parse(
    model="gpt-4o",
    messages=[
        {
            "role": "system",
            "content": "Extract invoice information from the image. Return structured JSON."
        },
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "Extract all information from this invoice:"},
                {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{image_b64}"}}
            ],
        }
    ],
    response_format=InvoiceInfo,
)

invoice = response.choices[0].message.parsed
print(f"Invoice #: {invoice.invoice_number}")
print(f"Date: {invoice.date}")
print(f"Total: {invoice.total_amount} {invoice.currency}")
for item in invoice.items:
    print(f"  - {item}")

2.3 複数の画像 — 写真を比較する

"""So sánh nhiều ảnh"""
image1_b64 = encode_image("product_v1.jpg")
image2_b64 = encode_image("product_v2.jpg")

response = client.chat.completions.create(
    model="gpt-4o",
    messages=[
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "So sánh 2 sản phẩm này. Khác biệt gì?"},
                {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{image1_b64}"}},
                {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{image2_b64}"}},
            ],
        }
    ],
)

print(response.choices[0].message.content)

3. Google Gemini ビジョン

"""Gemini Vision — hỗ trợ video, long context"""
import google.generativeai as genai
from PIL import Image

genai.configure(api_key="YOUR_API_KEY")
model = genai.GenerativeModel("gemini-2.0-flash")

# Image understanding
image = Image.open("chart.png")
response = model.generate_content([
    "Phân tích biểu đồ này chi tiết. Xu hướng chính là gì?",
    image,
])
print(response.text)

# Video understanding
video_file = genai.upload_file("presentation.mp4")
response = model.generate_content([
    "Tóm tắt nội dung video này. Các slide chính nói về gì?",
    video_file,
])
print(response.text)

4. 実用化

4.1 チャート/グラフ分析

"""Phân tích biểu đồ — VLM rất mạnh!"""
def analyze_chart(image_path, question=None):
    image_b64 = encode_image(image_path)
    prompt = question or "Phân tích biểu đồ này. Xu hướng chính? Insights quan trọng?"

    response = client.chat.completions.create(
        model="gpt-4o",
        messages=[{
            "role": "user",
            "content": [
                {"type": "text", "text": prompt},
                {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{image_b64}"}},
            ],
        }],
    )
    return response.choices[0].message.content

# Test
print(analyze_chart("revenue_chart.png", "Revenue Q4 tăng hay giảm?"))

4.2 製品の品質検査

"""Kiểm tra chất lượng sản phẩm bằng VLM"""
def inspect_product(image_path):
    image_b64 = encode_image(image_path)

    response = client.chat.completions.create(
        model="gpt-4o",
        messages=[
            {
                "role": "system",
                "content": """Bạn là chuyên gia QC. Kiểm tra sản phẩm trong ảnh.
                Output JSON: {"status": "pass/fail", "defects": [...], "confidence": "high/medium/low"}"""
            },
            {
                "role": "user",
                "content": [
                    {"type": "text", "text": "Kiểm tra sản phẩm này có lỗi không?"},
                    {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{image_b64}"}},
                ],
            }
        ],
    )
    return response.choices[0].message.content

result = inspect_product("product_sample.jpg")
print(result)

4.3 医用画像の予備分析

"""Phân tích ảnh y tế sơ bộ (KHÔNG thay thế bác sĩ!)"""
# ⚠️ DISCLAIMER: Chỉ dùng để hỗ trợ, KHÔNG thay thế chẩn đoán y khoa

response = client.chat.completions.create(
    model="gpt-4o",
    messages=[
        {
            "role": "system",
            "content": "You are a medical imaging assistant. Provide preliminary observations only. Always recommend professional medical review."
        },
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "Describe what you observe in this X-ray image."},
                {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{image_b64}"}},
            ],
        }
    ],
)

5. コストの最適化

5.1 コストを比較する

"""Tính cost cho image processing"""

# GPT-4o: 1 ảnh 768×768 (high detail) ≈ 765 tokens ≈ $0.0019
# GPT-4o-mini: cùng ảnh ≈ $0.0001
# Gemini Flash: cùng ảnh ≈ $0.00006

# 1000 ảnh/ngày:
# GPT-4o:      $1.90/day  = $57/month
# GPT-4o-mini: $0.10/day  = $3/month
# Gemini Flash: $0.06/day = $1.80/month

strategies = """
Cost Optimization:
1. Dùng "low" detail cho ảnh đơn giản (giảm 50% tokens)
2. Resize ảnh trước khi gửi (max 1024px)
3. Dùng GPT-4o-mini cho tasks đơn giản
4. Batch processing: gửi nhiều ảnh 1 lúc
5. Cache results: không gửi lại ảnh đã xử lý
6. Hybrid: VLM cho complex tasks, YOLO/OCR cho simple tasks
"""

5.2 VLM と従来の CV はどちらを使用するべきですか?

Dùng VLM (GPT-4o, Gemini):
✅ Complex understanding (hiểu context, suy luận)
✅ Unstructured documents
✅ Prototype nhanh
✅ Tasks thay đổi thường xuyên

Dùng Traditional CV (YOLO, OCR):
✅ High volume (>10K ảnh/giờ)
✅ Real-time (< 50ms)
✅ Offline processing
✅ Cost-sensitive
✅ Privacy (không gửi data lên cloud)

概要

コンセプト覚えておいてください
VLMモデルはテキストと画像の両方を同時に理解します。
GPT-4oオールラウンドで構造化された出力、推論
ジェミニ フラッシュ最安、迅速、ビデオサポート
クロード ソネット優れたチャートと図
コストGPT-4o-mini または Gemini フラッシュ (実稼働用)
ハイブリッド複雑な場合は VLM + 単純な場合は YOLO/OCR

一般的な演習

  1. 画像 Q&A: 5 つの異なる画像 (街路、食べ物、チャート、文書、製品) を GPT-4o に送信します。具体的な質問をしてください。
  2. 請求書の抽出: GPT-4o を使用した請求書抽出と PaddleOCR + 正規表現を使用した請求書抽出を比較します。どちらがより正確ですか?
  3. コスト計算ツール: 各プロバイダー (GPT-4o、mini、Gemini) の 1 日あたり 1000 枚のイメージのコストを計算します。
  4. モデルの比較: 同じ写真 + 質問 → GPT-4o、ジェミニ、クロードを送信します。答えを比較してください。

次の記事: エッジ デプロイメント — TensorRT、ONNX を使用して CV モデルをエッジ デバイスに導入します。