簡介
論文《Attention Is All You Need》(Vaswani 等人,2017 年)是 NLP 史上最大的轉捩點。 Transformer 完全消除了 RNN/LSTM,僅使用注意力,並成為每個現代 LLM 的基礎:GPT-4、Gemini、Claude、LLaMA。
1. 為什麼我們需要 Transformer?
| RNN/LSTM 的問題 | 變壓器解決 |
|---|---|
| 順序處理(慢) | 並行處理(快速) |
| 消失梯度(長→忘記) | 自註意力(直接連接到每個位置) |
| 固定上下文視窗 | 注意整個序列 |
| 難以擴展到GPU叢集 | 輕鬆並行化 |
2.Transformer架構
┌──────────────────────────────────────────────┐
│ TRANSFORMER │
│ │
│ ┌─────────────────┐ ┌─────────────────┐ │
│ │ ENCODER │ │ DECODER │ │
│ │ (stack of N) │ │ (stack of N) │ │
│ │ │ │ │ │
│ │ ┌──────────────┐│ │ ┌──────────────┐ │ │
│ │ │Multi-Head ││ │ │Masked MH │ │ │
│ │ │Self-Attention││ │ │Self-Attention │ │ │
│ │ └──────┬───────┘│ │ └──────┬───────┘ │ │
│ │ ┌──────┴───────┐│ │ ┌──────┴───────┐ │ │
│ │ │Add & Norm ││ │ │Add & Norm │ │ │
│ │ └──────┬───────┘│ │ └──────┬───────┘ │ │
│ │ ┌──────┴───────┐│ │ ┌──────┴───────┐ │ │
│ │ │Feed-Forward ││ │ │Cross-Attention│ │ │
│ │ └──────┬───────┘│ │ └──────┬───────┘ │ │
│ │ ┌──────┴───────┐│ │ ┌──────┴───────┐ │ │
│ │ │Add & Norm ││ │ │Feed-Forward │ │ │
│ │ └──────────────┘│ │ └──────────────┘ │ │
│ └────────┬────────┘ └────────┬────────┘ │
│ │ │ │
│ ┌────────┴────────┐ ┌───────┴────────┐ │
│ │Positional │ │Positional │ │
│ │Encoding + Embed │ │Encoding + Embed│ │
│ └─────────────────┘ └────────────────┘ │
└──────────────────────────────────────────────┘
3. 位置編碼
Transformer 沒有「順序」的概念—需要新增位置到嵌入:
$$PE_{(pos, 2i)} = \sin\left(\frac{pos}{10000^{2i/d_{model}}}\right)$$ $$PE_{(pos, 2i+1)} = \cos\left(\frac{pos}{10000^{2i/d_{模型}}}\right)$$
import torch
import math
class PositionalEncoding(nn.Module):
def __init__(self, d_model, max_len=5000):
super().__init__()
pe = torch.zeros(max_len, d_model)
position = torch.arange(0, max_len, dtype=torch.float).unsqueeze(1)
div_term = torch.exp(
torch.arange(0, d_model, 2).float() * (-math.log(10000.0) / d_model)
)
pe[:, 0::2] = torch.sin(position * div_term)
pe[:, 1::2] = torch.cos(position * div_term)
pe = pe.unsqueeze(0) # (1, max_len, d_model)
self.register_buffer('pe', pe)
def forward(self, x):
return x + self.pe[:, :x.size(1)]
4. 編碼器區塊
class TransformerEncoderBlock(nn.Module):
def __init__(self, d_model, num_heads, d_ff, dropout=0.1):
super().__init__()
self.self_attn = MultiHeadAttention(d_model, num_heads)
self.ffn = nn.Sequential(
nn.Linear(d_model, d_ff),
nn.ReLU(),
nn.Linear(d_ff, d_model),
)
self.norm1 = nn.LayerNorm(d_model)
self.norm2 = nn.LayerNorm(d_model)
self.dropout = nn.Dropout(dropout)
def forward(self, x, mask=None):
# Self-Attention + Residual + LayerNorm
attn_out = self.self_attn(x, x, x, mask)
x = self.norm1(x + self.dropout(attn_out))
# Feed-Forward + Residual + LayerNorm
ffn_out = self.ffn(x)
x = self.norm2(x + self.dropout(ffn_out))
return x
5. 不同問題的 Transformer
| 建築 | 使用部分 | 型號 | 數學問題 |
|---|---|---|---|
| 僅編碼器 | 編碼器 | 伯特,羅伯塔 | 分類、NER、QA |
| 僅解碼器 | 解碼器 | GPT、駱駝 | 文字產生 |
| 編碼器-解碼器 | 兩者 | T5、BART、mBART | 翻譯、摘要 |
BERT (Encoder-only):
Input: "The [MASK] sat on the mat"
Output: "The cat sat on the mat"
GPT (Decoder-only):
Input: "Once upon a time"
Output: "Once upon a time, there was a..."
T5 (Encoder-Decoder):
Input: "translate English to Vietnamese: Hello"
Output: "Xin chào"
6. 比較 Transformer 與 RNN
| 特點 | RNN/LSTM | 變壓器 |
|---|---|---|
| 並行度 | 順序 | 完全並行 |
| 遠端 | 梯度消失 | 直接關注 |
| 速度(訓練) | 慢 | 更快 |
| 記憶體 | 每步 O(1) | O(n²) 注意力矩陣 |
| 職位 | 隱式(序列) | 顯式(位置編碼) |
總結
| 成分 | 角色 |
|---|---|
| 自我關注 | 將每個令牌連接到每個令牌 |
| 多頭 | 許多不同的觀點 |
| 位置編碼 | 新增位置資訊 |
| 新增與規格 | 剩餘連接+層歸一化 |
| 前饋 | 非線性變換 |
下一篇文章
第 10 課:BERT — 第一個徹底改變 NLP 的預訓練語言模型:訓練一次,針對每個問題進行微調。