Chuyển đến nội dung chính

レッスン 21: マルチモーダル AI — 画像理解、文書 OCR、視覚的な質問応答

ビジョン モデルによる画像理解、ドキュメント OCR パイプライン、チャート/グラフ分析、視覚的な質問応答、マルチモーダル RAG、ナレッジ ベースの画像からテキストへの変換。

🏗️ アーキテクチャ — レッスン 21 レッスン 21: マルチモーダル AI — イメージ 理解、ドキュメント OCR、ビジュアル 質問への回答

エンタープライズ AI チャットボット プラットフォームのアーキテクチャ — プロトタイプから本番まで

パート 6: 高度な AI 機能

xdev.asia

1. エンタープライズチャットボットにおけるマルチモーダル AI

企業ユーザーはテキストだけでなくテキストも理解できるチャットボットを必要としています 画像、ドキュメント、チャート、スクリーンショット。マルチモーダル AI は、チャットボットを「テキストのみのアシスタント」から「視覚認識型インテリジェント エージェント」に変えます。


┌─────────── MULTIMODAL PIPELINE ──────────────────────┐
│                                                       │
│  Input Types:                                         │
│  ┌────────┐ ┌────────┐ ┌────────┐ ┌────────────┐     │
│  │ Photo  │ │  PDF   │ │ Chart  │ │ Screenshot │     │
│  │        │ │  Scan  │ │ Graph  │ │            │     │
│  └───┬────┘ └───┬────┘ └───┬────┘ └─────┬──────┘     │
│      │          │          │            │             │
│  ┌───▼──────────▼──────────▼────────────▼──────┐      │
│  │         MULTIMODAL ROUTER                   │      │
│  │  (Detect content type → route to pipeline)  │      │
│  └───┬──────────┬──────────┬────────────┬──────┘      │
│      │          │          │            │             │
│      ▼          ▼          ▼            ▼             │
│  ┌──────┐  ┌──────┐  ┌──────┐    ┌──────────┐        │
│  │Vision│  │ OCR  │  │Chart │    │  Screen  │        │
│  │Model │  │Engine│  │Parser│    │ Understanding│     │
│  └──┬───┘  └──┬───┘  └──┬───┘    └─────┬────┘        │
│     └─────────┴─────────┴──────────────┘             │
│                     │                                 │
│              ┌──────▼──────┐                          │
│              │ Unified     │                          │
│              │ Context     │──▶ LLM                   │
│              └─────────────┘                          │
└───────────────────────────────────────────────────────┘

2. ビジョンモデルの統合


class VisionProcessor {
  async processImage(
    image: ImageInput,
    query: string,
    context: VisionContext,
  ): Promise<VisionResult> {
    // 1. Optimize image for vision model
    const optimized = await this.optimizeForVision(image);

    // 2. Route to appropriate vision pipeline
    const contentType = await this.classifyImageContent(optimized);

    switch (contentType) {
      case 'document':
        return this.processDocument(optimized, query);
      case 'chart':
        return this.processChart(optimized, query);
      case 'product_photo':
        return this.processProductPhoto(optimized, query);
      case 'screenshot':
        return this.processScreenshot(optimized, query);
      default:
        return this.processGeneral(optimized, query);
    }
  }

  private async classifyImageContent(image: OptimizedImage): Promise<string> {
    const response = await this.llm.chat({
      messages: [{
        role: 'user',
        content: [
          { type: 'text', text: 'Classify this image into one of: document, chart, product_photo, screenshot, general. Output only the category.' },
          {
            type: 'image_url',
            image_url: {
              url: `data:${image.mimeType};base64,${image.base64}`,
              detail: 'low', // Low detail for classification (cheaper)
            },
          },
        ],
      }],
      model: 'gpt-4o-mini',
      maxTokens: 20,
    });

    return response.content.trim().toLowerCase();
  }

  private async processGeneral(
    image: OptimizedImage,
    query: string,
  ): Promise<VisionResult> {
    const response = await this.llm.chat({
      messages: [{
        role: 'user',
        content: [
          { type: 'text', text: query || 'Describe this image in detail.' },
          {
            type: 'image_url',
            image_url: {
              url: `data:${image.mimeType};base64,${image.base64}`,
              detail: 'high',
            },
          },
        ],
      }],
      model: 'gpt-4o',
    });

    return {
      type: 'general',
      description: response.content,
      extractedText: null,
      structuredData: null,
    };
  }
}

3. ドキュメント OCR パイプライン


class DocumentOCRPipeline {
  async process(document: DocumentInput): Promise<OCRResult> {
    const pages: PageResult[] = [];

    // 1. Convert document to images (if PDF)
    const images = document.type === 'pdf'
      ? await this.pdfToImages(document.data)
      : [{ data: document.data, mimeType: document.mimeType }];

    // 2. OCR each page
    for (let i = 0; i < images.length; i++) {
      // Strategy A: Vision model OCR (higher accuracy, slower)
      const visionOCR = await this.visionModelOCR(images[i]);

      // Strategy B: Tesseract OCR (faster, good for clear text)
      const tesseractOCR = await this.tesseractOCR(images[i]);

      // 3. Merge results (use vision for complex layouts, Tesseract for simple)
      const mergedText = this.mergeOCRResults(visionOCR, tesseractOCR);

      // 4. Structure extraction (tables, forms, lists)
      const structured = await this.extractStructure(images[i], mergedText);

      pages.push({
        pageNumber: i + 1,
        text: mergedText,
        tables: structured.tables,
        formFields: structured.formFields,
        confidence: structured.confidence,
      });
    }

    return {
      totalPages: pages.length,
      fullText: pages.map(p => p.text).join('\n\n---\n\n'),
      pages,
      language: await this.detectLanguage(pages[0].text),
    };
  }

  private async visionModelOCR(image: ImageData): Promise<string> {
    const response = await this.llm.chat({
      messages: [{
        role: 'user',
        content: [
          {
            type: 'text',
            text: `Extract ALL text from this document image. 
Preserve the original formatting and structure.
For tables, use markdown table format.
For forms, extract field labels and values as "Label: Value".
Output the raw extracted text only.`,
          },
          {
            type: 'image_url',
            image_url: {
              url: `data:${image.mimeType};base64,${Buffer.from(image.data).toString('base64')}`,
              detail: 'high',
            },
          },
        ],
      }],
      model: 'gpt-4o',
      maxTokens: 4096,
    });

    return response.content;
  }

  private async extractStructure(
    image: ImageData,
    text: string,
  ): Promise<StructuredContent> {
    const response = await this.llm.chat({
      messages: [{
        role: 'user',
        content: [
          {
            type: 'text',
            text: `Analyze this document and extract structured data.
Return JSON with: tables (as arrays), formFields (key-value pairs), lists.`,
          },
          {
            type: 'image_url',
            image_url: {
              url: `data:${image.mimeType};base64,${Buffer.from(image.data).toString('base64')}`,
              detail: 'high',
            },
          },
        ],
      }],
      response_format: { type: 'json_object' },
      model: 'gpt-4o',
    });

    return JSON.parse(response.content);
  }
}

4. チャートとグラフの分析


class ChartAnalyzer {
  async analyzeChart(
    chartImage: ImageInput,
    question: string,
  ): Promise<ChartAnalysis> {
    const response = await this.llm.chat({
      messages: [{
        role: 'user',
        content: [
          {
            type: 'text',
            text: `Analyze this chart/graph and answer the question.

1. Identify chart type (bar, line, pie, scatter, etc.)
2. Extract data points and labels
3. Identify trends and insights
4. Answer the specific question

Question: ${question || 'What are the key insights from this chart?'}

Output JSON:
{
  "chartType": "...",
  "title": "...",
  "dataPoints": [{"label": "...", "value": N}],
  "trends": ["..."],
  "insights": ["..."],
  "answer": "..."
}`,
          },
          {
            type: 'image_url',
            image_url: {
              url: `data:${chartImage.mimeType};base64,${chartImage.base64}`,
              detail: 'high',
            },
          },
        ],
      }],
      response_format: { type: 'json_object' },
      model: 'gpt-4o',
    });

    return JSON.parse(response.content);
  }
}

5. マルチモーダル RAG — ナレッジベース内の画像のインデックス付け


class MultimodalRAG {
  // Index images alongside text in knowledge base
  async indexDocumentWithImages(
    document: DocumentWithImages,
    tenantId: string,
  ): Promise<void> {
    // 1. Index text chunks as usual
    for (const chunk of document.textChunks) {
      const embedding = await this.textEmbedder.embed(chunk.content);
      await this.vectorStore.upsert({
        id: `${document.id}:text:${chunk.index}`,
        vector: embedding,
        metadata: {
          tenantId,
          documentId: document.id,
          type: 'text',
          content: chunk.content,
        },
      });
    }

    // 2. Generate descriptions for images → index as text
    for (const image of document.images) {
      const description = await this.describeImage(image);

      const embedding = await this.textEmbedder.embed(description);
      await this.vectorStore.upsert({
        id: `${document.id}:image:${image.index}`,
        vector: embedding,
        metadata: {
          tenantId,
          documentId: document.id,
          type: 'image',
          content: description,
          imageUrl: image.url,
          pageNumber: image.pageNumber,
        },
      });
    }
  }

  // Retrieve with image context
  async retrieve(
    query: string,
    tenantId: string,
  ): Promise<MultimodalSearchResult[]> {
    const results = await this.vectorStore.search({
      vector: await this.textEmbedder.embed(query),
      filter: { tenantId },
      topK: 10,
    });

    return results.map(r => ({
      content: r.metadata.content as string,
      type: r.metadata.type as 'text' | 'image',
      imageUrl: r.metadata.type === 'image' ? r.metadata.imageUrl as string : undefined,
      score: r.score,
      documentId: r.metadata.documentId as string,
    }));
  }
}

レッスン 21 のまとめ

  • ビジョンルーター: 画像タイプ (ドキュメント、チャート、写真、スクリーンショット) を自動分類 → パイプラインにルーティング
  • ドキュメントOCR: ビジョンモデル + Tesseract フォールバック、テーブル/フォーム抽出、マルチページサポート
  • チャート分析: グラフの種類の識別、データ ポイントの抽出、傾向の検出、質問への回答
  • マルチモーダル RAG: 画像をテキスト説明としてインデックス化 → テキストドキュメントと一緒に検索可能
  • コストの最適化: 使用する 詳細: 「低い」 分類のために、 詳細: 「高い」 抽出用

次の記事: ワークフローの自動化 — チャットボットによってトリガーされるワークフロー、承認フロー、n8n/Temporal との統合、イベント駆動型の自動化。