はじめに
論文「Attention Is All You Need」(Vaswani et al.、2017)は、NLP の歴史の中で 最大の転換点 です。 Transformer は RNN/LSTM を完全に排除し、注意のみを使用し、すべて の最新の LLM (GPT-4、Gemini、Claude、LLaMA) の基礎となります。
1. なぜ変圧器が必要なのでしょうか?
| RNN/LSTMの問題点 | 変圧器の解決 |
|---|---|
| 逐次処理(遅い) | 並列処理 (高速) |
| 消失勾配 (長い→忘れる) | セルフアテンション(あらゆる場所への直接接続) |
| コンテキストウィンドウを修正 | 全体 シーケンスに注意 |
| GPU クラスターに合わせて拡張するのが難しい | 簡単な並列化 |
2. トランスのアーキテクチャ
┌──────────────────────────────────────────────┐
│ TRANSFORMER │
│ │
│ ┌─────────────────┐ ┌─────────────────┐ │
│ │ ENCODER │ │ DECODER │ │
│ │ (stack of N) │ │ (stack of N) │ │
│ │ │ │ │ │
│ │ ┌──────────────┐│ │ ┌──────────────┐ │ │
│ │ │Multi-Head ││ │ │Masked MH │ │ │
│ │ │Self-Attention││ │ │Self-Attention │ │ │
│ │ └──────┬───────┘│ │ └──────┬───────┘ │ │
│ │ ┌──────┴───────┐│ │ ┌──────┴───────┐ │ │
│ │ │Add & Norm ││ │ │Add & Norm │ │ │
│ │ └──────┬───────┘│ │ └──────┬───────┘ │ │
│ │ ┌──────┴───────┐│ │ ┌──────┴───────┐ │ │
│ │ │Feed-Forward ││ │ │Cross-Attention│ │ │
│ │ └──────┬───────┘│ │ └──────┬───────┘ │ │
│ │ ┌──────┴───────┐│ │ ┌──────┴───────┐ │ │
│ │ │Add & Norm ││ │ │Feed-Forward │ │ │
│ │ └──────────────┘│ │ └──────────────┘ │ │
│ └────────┬────────┘ └────────┬────────┘ │
│ │ │ │
│ ┌────────┴────────┐ ┌───────┴────────┐ │
│ │Positional │ │Positional │ │
│ │Encoding + Embed │ │Encoding + Embed│ │
│ └─────────────────┘ └────────────────┘ │
└──────────────────────────────────────────────┘
3. 位置エンコーディング
Transformer には「順序」の概念がありません。埋め込みに 位置を追加する必要があります。
$$PE_{(pos, 2i)} = \sin\left(\frac{pos}{10000^{2i/d_{モデル}}}\right)$$ $$PE_{(pos, 2i+1)} = \cos\left(\frac{pos}{10000^{2i/d_{model}}}\right)$$
import torch
import math
class PositionalEncoding(nn.Module):
def __init__(self, d_model, max_len=5000):
super().__init__()
pe = torch.zeros(max_len, d_model)
position = torch.arange(0, max_len, dtype=torch.float).unsqueeze(1)
div_term = torch.exp(
torch.arange(0, d_model, 2).float() * (-math.log(10000.0) / d_model)
)
pe[:, 0::2] = torch.sin(position * div_term)
pe[:, 1::2] = torch.cos(position * div_term)
pe = pe.unsqueeze(0) # (1, max_len, d_model)
self.register_buffer('pe', pe)
def forward(self, x):
return x + self.pe[:, :x.size(1)]
4. エンコーダーブロック
class TransformerEncoderBlock(nn.Module):
def __init__(self, d_model, num_heads, d_ff, dropout=0.1):
super().__init__()
self.self_attn = MultiHeadAttention(d_model, num_heads)
self.ffn = nn.Sequential(
nn.Linear(d_model, d_ff),
nn.ReLU(),
nn.Linear(d_ff, d_model),
)
self.norm1 = nn.LayerNorm(d_model)
self.norm2 = nn.LayerNorm(d_model)
self.dropout = nn.Dropout(dropout)
def forward(self, x, mask=None):
# Self-Attention + Residual + LayerNorm
attn_out = self.self_attn(x, x, x, mask)
x = self.norm1(x + self.dropout(attn_out))
# Feed-Forward + Residual + LayerNorm
ffn_out = self.ffn(x)
x = self.norm2(x + self.dropout(ffn_out))
return x
5. さまざまな問題に対応する変圧器
| 建築 | 使用セクション | モデル | 数学の問題 |
|---|---|---|---|
| エンコーダのみ | エンコーダ | ロベルタ・バート | 分類、NER、QA |
| デコーダのみ | デコーダ | GPT、LLaMA | テキスト生成 |
| エンコーダ-デコーダ | 両方 | T5、BART、mBART | 翻訳、要約 |
BERT (Encoder-only):
Input: "The [MASK] sat on the mat"
Output: "The cat sat on the mat"
GPT (Decoder-only):
Input: "Once upon a time"
Output: "Once upon a time, there was a..."
T5 (Encoder-Decoder):
Input: "translate English to Vietnamese: Hello"
Output: "Xin chào"
6. トランスと RNN の比較
| 特長 | RNN/LSTM | 変圧器 |
|---|---|---|
| 並列処理 | シーケンシャル | 完全並列 |
| 長距離 | 消えるグラデーション | 直接の注意 |
| スピード(トレーニング) | 遅い | はるかに高速 |
| メモリ | ステップごとに O(1) | O(n²) アテンション行列 |
| ポジション | 暗黙的 (シーケンス) | 明示的 (位置エンコーディング) |
概要
| 成分 | 役割 |
|---|---|
| 自己注意 | すべてのトークンをすべてのトークンに接続する |
| マルチヘッド | さまざまな視点 |
| 位置エンコーディング | 位置情報の追加 |
| 追加と標準化 | 残留接続 + 層正規化 |
| フィードフォワード | 非線形変換 |
次の記事
レッスン 10: BERT — NLP を完全に変える最初の事前トレーニング済み言語モデル: 一度トレーニングすれば、問題ごとに微調整できます。