Chuyển đến nội dung chính

レッスン 9: トランスフォーマー — 「必要なのは注意だけです」

詳細な Transformer アーキテクチャ: エンコーダ-デコーダ、位置エンコーディング、レイヤー正規化、フィードフォワード ネットワーク。 Transformer が RNN に勝る理由: 並列化、長距離依存性。 PyTorch を使用してゼロからコード Transformer を作成します。注釈付きのトランスフォーマーのウォークスルー。

🧠 AI と ML — レッスン 8 レッスン 9: トランスフォーマー — 「注意力だけです」 必要です」

NLP の基礎から上級まで: 自然言語処理をマスターする

パート 3: NLP のための深層学習 — RNN、LSTM、Transformer へ

xdev.asia

はじめに

論文「Attention Is All You Need」(Vaswani et al.、2017)は、NLP の歴史の中で 最大の転換点 です。 Transformer は RNN/LSTM を完全に排除し、注意のみを使用し、すべて の最新の LLM (GPT-4、Gemini、Claude、LLaMA) の基礎となります。


1. なぜ変圧器が必要なのでしょうか?

RNN/LSTMの問題点変圧器の解決
逐次処理(遅い)並列処理 (高速)
消失勾配 (長い→忘れる)セルフアテンション(あらゆる場所への直接接続)
コンテキストウィンドウを修正全体 シーケンスに注意
GPU クラスターに合わせて拡張するのが難しい簡単な並列化

2. トランスのアーキテクチャ

┌──────────────────────────────────────────────┐
│                TRANSFORMER                    │
│                                              │
│  ┌─────────────────┐  ┌─────────────────┐   │
│  │    ENCODER       │  │    DECODER       │   │
│  │  (stack of N)    │  │  (stack of N)    │   │
│  │                  │  │                  │   │
│  │ ┌──────────────┐│  │ ┌──────────────┐ │   │
│  │ │Multi-Head    ││  │ │Masked MH     │ │   │
│  │ │Self-Attention││  │ │Self-Attention │ │   │
│  │ └──────┬───────┘│  │ └──────┬───────┘ │   │
│  │ ┌──────┴───────┐│  │ ┌──────┴───────┐ │   │
│  │ │Add & Norm    ││  │ │Add & Norm    │ │   │
│  │ └──────┬───────┘│  │ └──────┬───────┘ │   │
│  │ ┌──────┴───────┐│  │ ┌──────┴───────┐ │   │
│  │ │Feed-Forward  ││  │ │Cross-Attention│ │   │
│  │ └──────┬───────┘│  │ └──────┬───────┘ │   │
│  │ ┌──────┴───────┐│  │ ┌──────┴───────┐ │   │
│  │ │Add & Norm    ││  │ │Feed-Forward  │ │   │
│  │ └──────────────┘│  │ └──────────────┘ │   │
│  └────────┬────────┘  └────────┬────────┘   │
│           │                    │              │
│  ┌────────┴────────┐  ┌───────┴────────┐    │
│  │Positional       │  │Positional      │    │
│  │Encoding + Embed │  │Encoding + Embed│    │
│  └─────────────────┘  └────────────────┘    │
└──────────────────────────────────────────────┘

3. 位置エンコーディング

Transformer には「順序」の概念がありません。埋め込みに 位置を追加する必要があります。

$$PE_{(pos, 2i)} = \sin\left(\frac{pos}{10000^{2i/d_{モデル}}}\right)$$ $$PE_{(pos, 2i+1)} = \cos\left(\frac{pos}{10000^{2i/d_{model}}}\right)$$

import torch
import math

class PositionalEncoding(nn.Module):
    def __init__(self, d_model, max_len=5000):
        super().__init__()
        pe = torch.zeros(max_len, d_model)
        position = torch.arange(0, max_len, dtype=torch.float).unsqueeze(1)
        div_term = torch.exp(
            torch.arange(0, d_model, 2).float() * (-math.log(10000.0) / d_model)
        )
        pe[:, 0::2] = torch.sin(position * div_term)
        pe[:, 1::2] = torch.cos(position * div_term)
        pe = pe.unsqueeze(0)  # (1, max_len, d_model)
        self.register_buffer('pe', pe)

    def forward(self, x):
        return x + self.pe[:, :x.size(1)]

4. エンコーダーブロック

class TransformerEncoderBlock(nn.Module):
    def __init__(self, d_model, num_heads, d_ff, dropout=0.1):
        super().__init__()
        self.self_attn = MultiHeadAttention(d_model, num_heads)
        self.ffn = nn.Sequential(
            nn.Linear(d_model, d_ff),
            nn.ReLU(),
            nn.Linear(d_ff, d_model),
        )
        self.norm1 = nn.LayerNorm(d_model)
        self.norm2 = nn.LayerNorm(d_model)
        self.dropout = nn.Dropout(dropout)

    def forward(self, x, mask=None):
        # Self-Attention + Residual + LayerNorm
        attn_out = self.self_attn(x, x, x, mask)
        x = self.norm1(x + self.dropout(attn_out))

        # Feed-Forward + Residual + LayerNorm
        ffn_out = self.ffn(x)
        x = self.norm2(x + self.dropout(ffn_out))
        return x

5. さまざまな問題に対応する変圧器

建築使用セクションモデル数学の問題
エンコーダのみエンコーダロベルタ・バート分類、NER、QA
デコーダのみデコーダGPT、LLaMAテキスト生成
エンコーダ-デコーダ両方T5、BART、mBART翻訳、要約
BERT (Encoder-only):
  Input: "The [MASK] sat on the mat"
  Output: "The cat sat on the mat"

GPT (Decoder-only):
  Input: "Once upon a time"
  Output: "Once upon a time, there was a..."

T5 (Encoder-Decoder):
  Input: "translate English to Vietnamese: Hello"
  Output: "Xin chào"

6. トランスと RNN の比較

特長RNN/LSTM変圧器
並列処理シーケンシャル完全並列
長距離消えるグラデーション直接の注意
スピード(トレーニング)遅いはるかに高速
メモリステップごとに O(1)O(n²) アテンション行列
ポジション暗黙的 (シーケンス)明示的 (位置エンコーディング)

概要

成分役割
自己注意すべてのトークンをすべてのトークンに接続する
マルチヘッドさまざまな視点
位置エンコーディング位置情報の追加
追加と標準化残留接続 + 層正規化
フィードフォワード非線形変換

次の記事

レッスン 10: BERT — NLP を完全に変える最初の事前トレーニング済み言語モデル: 一度トレーニングすれば、問題ごとに微調整できます。