Chuyển đến nội dung chính

第2課:TransformerアーキテクチャとAttentionメカニズム

Self-attention、Multi-Head Attention、Positional Encoding。 Encoder-Decoderアーキテクチャ。BERT、GPT、T5モデルファミリー。 トークン化:BPE、WordPiece、SentencePiece。 NLPタスク:分類、NER、QA、要約。

1. はじめに

Transformerは、GPT、BERT、Stable DiffusionからLLaMAまで、すべての現代のGenerative AIモデルの基盤となるアーキテクチャです。NVIDIA DLIアセスメントでは、Attentionメカニズムの仕組みを十分に理解し、PyTorchで実装できる必要があります。

この課では、Scaled Dot-Product Attentionから完全なTransformerアーキテクチャまでを網羅し、具体的なモデルファミリーやNLPタスクとの対応を解説します。

試験のヒント: NVIDIA DLIアセスメントでは、Attentionメカニズムのコードを完成させたり、Transformerの次元不一致エラーをデバッグする問題がよく出題されます。Attentionの各ステップでのテンソルの形状をマスターすることが、アセスメント合格の鍵です。

Transformerアーキテクチャ — Encoder-Decoder、Self-Attention、Cross-Attention
Transformerアーキテクチャ — Encoder-Decoder、Self-Attention、Cross-Attention

2. Attentionメカニズム

2.1 直感的理解 — 「最も重要なものは何か?」

Attentionは次の質問に答えます:「現在のトークンを処理するとき、入力シーケンスのどのトークンが最も重要か?」RNNのようにシーケンス全体を固定サイズのベクトルに圧縮するのではなく、Attentionはモデルが入力のすべての位置を直接「見る」ことを可能にします。

このメカニズムは3つのコンポーネントで動作します:

  • Query (Q) — 「何を探しているのか?」— 現在のトークンが質問をします
  • Key (K) — 「どんな情報を持っているか?」— 各トークンが自身の内容を提示します
  • Value (V) — 「これが実際の情報です」— マッチした際に返される内容です
例: "The cat sat on the mat because it was tired"
                                              ↑
                                        トークン "it" (Query)
                                              │
                ┌─────────────────────────────┤
                │     Attentionスコア:        │
                │   "cat"  = 0.72  ← 高い!   │
                │   "mat"  = 0.11            │
                │   "sat"  = 0.08            │
                │   "The"  = 0.03            │
                │   ...                       │
                └─────────────────────────────┘
                → "it" は主に "cat" に注目する

2.2 Scaled Dot-Product Attention

Attentionの核となる公式:

Attention(Q, K, V) = softmax(Q · K^T / √d_k) · V

各要素の説明:
  Q: Queryマトリクス  — 形状 (seq_len, d_k)
  K: Keyマトリクス    — 形状 (seq_len, d_k)
  V: Valueマトリクス  — 形状 (seq_len, d_v)
  d_k: Keyベクトルの次元数
  √d_k: softmaxでの勾配消失を防ぐためのスケーリングファクター

なぜ√d_kによるスケーリングが必要なのでしょうか?d_kが大きいと、内積Q·K^Tが非常に大きくなり、softmaxが飽和して勾配が0に近づきます。√d_kで割ることで分散を安定させます。

Scaled Dot-Product Attentionの流れ:

  Q ──┐
      │──→ MatMul ──→ Scale (÷√d_k) ──→ Mask (任意) ──→ Softmax ──→ MatMul ──→ 出力
  K ──┘                                                                ↑
                                                                       │
  V ────────────────────────────────────────────────────────────────────┘

形状 (batch_size=B, seq_len=S, d_k=D):
  Q:        (B, S, D)
  K^T:      (B, D, S)
  Q·K^T:    (B, S, S)   ← Attentionスコアマトリクス
  softmax:  (B, S, S)   ← Attention重み(各行の合計が1)
  × V:      (B, S, D)   ← 重み付き出力

2.3 コード:Attentionをゼロから実装

import torch
import torch.nn.functional as F
import math

def scaled_dot_product_attention(Q, K, V, mask=None):
    """
    Q: (batch, seq_len, d_k)
    K: (batch, seq_len, d_k)
    V: (batch, seq_len, d_v)
    mask: (batch, 1, seq_len) or (batch, seq_len, seq_len)
    """
    d_k = Q.size(-1)

    # ステップ1: Attentionスコアを計算
    scores = torch.matmul(Q, K.transpose(-2, -1)) / math.sqrt(d_k)
    # scoresの形状: (batch, seq_len, seq_len)

    # ステップ2: マスクを適用(デコーダのcausal attentionに使用)
    if mask is not None:
        scores = scores.masked_fill(mask == 0, float('-inf'))

    # ステップ3: softmaxでAttention重みを取得
    attn_weights = F.softmax(scores, dim=-1)
    # attn_weightsの形状: (batch, seq_len, seq_len)

    # ステップ4: Valueの重み付き和
    output = torch.matmul(attn_weights, V)
    # outputの形状: (batch, seq_len, d_v)

    return output, attn_weights

# デモ
batch_size, seq_len, d_k = 2, 5, 64
Q = torch.randn(batch_size, seq_len, d_k)
K = torch.randn(batch_size, seq_len, d_k)
V = torch.randn(batch_size, seq_len, d_k)

output, weights = scaled_dot_product_attention(Q, K, V)
print(f"Output shape: {output.shape}")    # (2, 5, 64)
print(f"Weights shape: {weights.shape}")  # (2, 5, 5)
print(f"Weights sum per row: {weights.sum(dim=-1)}")  # 各行 = 1.0

試験のヒント: Attention実装で最も多いエラー:K.transpose(-2, -1)を忘れる、または間違った次元で割り算をする。各ステップ後に形状を必ず確認してください — scoresは(batch, seq_len, seq_len)の形状でなければなりません。

3. Multi-Head Attention

3.1 なぜ複数のヘッドが必要なのか?

単一のヘッドでは1種類の関係しか学習できません。Multi-Head Attentionにより、モデルは異なる表現サブスペースに同時に注目できます:

  • ヘッド1:構文関係を学習(主語-動詞)
  • ヘッド2:共参照を学習(代名詞→名詞)
  • ヘッド3:位置的近接性を学習
  • ヘッド4:意味的類似性を学習
Multi-Head Attentionの流れ:

入力 (batch, seq_len, d_model)
    │
    ├── Linear → Q ──┐
    ├── Linear → K ──┼── hヘッドに分割
    └── Linear → V ──┘
                      │
        ┌─────────────┼─────────────┐
        ▼             ▼             ▼
    ヘッド1       ヘッド2    ...  ヘッドh
 Attention()   Attention()    Attention()
        │             │             │
        └─────────────┼─────────────┘
                      │
                   結合
                      │
                   Linear
                      │
                   出力 (batch, seq_len, d_model)

形状 (d_model=512, h=8, d_k = d_model/h = 64):
  入力:             (B, S, 512)
  各ヘッドのQ/K/V:  (B, h, S, 64)   ← Linear後にreshape
  各ヘッドの出力:    (B, h, S, 64)
  結合:             (B, S, 512)      ← h × d_k = d_model
  最終出力:         (B, S, 512)

3.2 コード:MultiHeadAttentionモジュール

import torch
import torch.nn as nn
import math

class MultiHeadAttention(nn.Module):
    def __init__(self, d_model, num_heads):
        super().__init__()
        assert d_model % num_heads == 0, "d_model must be divisible by num_heads"

        self.d_model = d_model
        self.num_heads = num_heads
        self.d_k = d_model // num_heads

        # Q, K, V および出力のLinear射影
        self.W_q = nn.Linear(d_model, d_model)
        self.W_k = nn.Linear(d_model, d_model)
        self.W_v = nn.Linear(d_model, d_model)
        self.W_o = nn.Linear(d_model, d_model)

    def forward(self, query, key, value, mask=None):
        batch_size = query.size(0)

        # ステップ1: Linear射影
        Q = self.W_q(query)  # (B, S, d_model)
        K = self.W_k(key)
        V = self.W_v(value)

        # ステップ2: マルチヘッド形式にreshape
        # (B, S, d_model) → (B, S, h, d_k) → (B, h, S, d_k)
        Q = Q.view(batch_size, -1, self.num_heads, self.d_k).transpose(1, 2)
        K = K.view(batch_size, -1, self.num_heads, self.d_k).transpose(1, 2)
        V = V.view(batch_size, -1, self.num_heads, self.d_k).transpose(1, 2)

        # ステップ3: 各ヘッドでScaled Dot-Product Attention
        scores = torch.matmul(Q, K.transpose(-2, -1)) / math.sqrt(self.d_k)

        if mask is not None:
            scores = scores.masked_fill(mask == 0, float('-inf'))

        attn_weights = torch.softmax(scores, dim=-1)
        context = torch.matmul(attn_weights, V)
        # context: (B, h, S, d_k)

        # ステップ4: ヘッドを結合
        # (B, h, S, d_k) → (B, S, h, d_k) → (B, S, d_model)
        context = context.transpose(1, 2).contiguous().view(
            batch_size, -1, self.d_model
        )

        # ステップ5: 最終Linear射影
        output = self.W_o(context)
        return output

# デモ
d_model, num_heads = 512, 8
mha = MultiHeadAttention(d_model, num_heads)

x = torch.randn(2, 10, d_model)  # batch=2, seq_len=10
output = mha(x, x, x)  # self-attention: Q=K=V=x
print(f"Output shape: {output.shape}")  # (2, 10, 512)
パラメータ一般的な値備考
d_model512, 768, 1024埋め込みサイズ
num_heads8, 12, 16d_modelはnum_headsで割り切れる必要がある
d_k = d_model / h64, 64, 64各ヘッドは通常d_k = 64
総パラメータ数(MHA)4 × d_model²4つのLinear層:W_q, W_k, W_v, W_o

試験のヒント: DLIアセスメントでAttentionのRuntimeError: shape mismatchに遭遇した場合、以下を確認してください:(1) d_model % num_heads == 0、(2) view()とtranspose()が正しい順序になっている、(3) transpose後に.view()を呼ぶ前に.contiguous()を呼んでいる。

4. Transformerアーキテクチャ

4.1 Encoderブロック

Transformer Encoderの各レイヤーは、残差接続とレイヤー正規化を持つ2つのサブレイヤーで構成されます:

Transformer Encoderブロック:

  入力
    │
    ▼
┌─────────────────────────────┐
│  Multi-Head Self-Attention  │
└──────────────┬──────────────┘
               │
    ┌──────────┴──────────┐
    │     Add & Norm      │ ← x + Sublayer(x)、次にLayerNorm
    └──────────┬──────────┘
               │
    ┌──────────┴──────────┐
    │  Feed-Forward Net   │ ← 2つのLinear層 + ReLU/GELU
    │  FFN(x) = W₂·σ(W₁x + b₁) + b₂
    └──────────┬──────────┘
               │
    ┌──────────┴──────────┐
    │     Add & Norm      │
    └──────────┬──────────┘
               │
             出力

4.2 Decoderブロック

Decoderには、Encoderの出力に注目するためのCross-Attentionサブレイヤーが追加されています:

Transformer Decoderブロック:

  入力(右シフト)
    │
    ▼
┌────────────────────────────────┐
│  Masked Multi-Head Attention   │ ← causal mask: 以前のトークンのみ参照可能
└──────────────┬─────────────────┘
               │
    ┌──────────┴──────────┐
    │     Add & Norm      │
    └──────────┬──────────┘
               │
    ┌──────────┴──────────────────────────┐
    │  Cross Multi-Head Attention         │ ← QはDecoder、K/VはEncoderから
    │  Q = Decoderの隠れ状態              │
    │  K, V = Encoderの出力               │
    └──────────────────┬──────────────────┘
               │
    ┌──────────┴──────────┐
    │     Add & Norm      │
    └──────────┬──────────┘
               │
    ┌──────────┴──────────┐
    │  Feed-Forward Net   │
    └──────────┬──────────┘
               │
    ┌──────────┴──────────┐
    │     Add & Norm      │
    └──────────┬──────────┘
               │
             出力

4.3 完全なTransformerアーキテクチャ

┌─────────────────────────────────────────────────────────────────┐
│                    TRANSFORMERアーキテクチャ                     │
├───────────────────────┬─────────────────────────────────────────┤
│       ENCODER         │              DECODER                    │
│                       │                                         │
│   Input Embedding     │         Output Embedding                │
│        +              │              +                          │
│   Positional Enc.     │         Positional Enc.                 │
│        │              │              │                          │
│   ┌────┴────┐         │         ┌────┴──────────┐              │
│   │  MH     │         │         │  Masked MH    │              │
│   │  Self-  │         │         │  Self-        │              │
│   │  Attn   │         │         │  Attention    │              │
│   └────┬────┘         │         └────┬──────────┘              │
│   Add & Norm          │         Add & Norm                     │
│        │              │              │                          │
│   ┌────┴────┐         │    ┌────────┴────────────┐             │
│   │  Feed   │         │    │  Cross MH Attention │             │
│   │ Forward │    ─────┼───►│  Q=dec, K/V=enc    │             │
│   │   Net   │         │    └─────────┬───────────┘             │
│   └────┬────┘         │         Add & Norm                     │
│   Add & Norm          │              │                          │
│        │              │         ┌────┴────┐                    │
│      × N層            │         │  Feed   │                    │
│        │              │         │ Forward │                    │
│     Encoder           │         └────┬────┘                    │
│     出力              │         Add & Norm                     │
│                       │              │                          │
│                       │           × N層                        │
│                       │              │                          │
│                       │         Linear + Softmax               │
│                       │              │                          │
│                       │       出力確率分布                       │
└───────────────────────┴─────────────────────────────────────────┘

4.4 Positional Encoding

TransformerにはRNNのような「順序」の概念がありません。Positional Encodingはsin/cos関数を使って、埋め込みに位置情報を追加します:

PE(pos, 2i)   = sin(pos / 10000^(2i/d_model))
PE(pos, 2i+1) = cos(pos / 10000^(2i/d_model))

pos: シーケンス内のトークン位置 (0, 1, 2, ...)
i:   次元のインデックス (0, 1, 2, ..., d_model/2)

これはDiffusion Modelsで再利用されるまったく同じ概念です — タイムステップtをエンコードするための正弦波埋め込みです。ここでPositional Encodingをマスターすれば、Diffusionのセクションで大いに役立ちます。

import torch
import math

class PositionalEncoding(nn.Module):
    def __init__(self, d_model, max_len=5000):
        super().__init__()
        pe = torch.zeros(max_len, d_model)
        position = torch.arange(0, max_len).unsqueeze(1).float()
        div_term = torch.exp(
            torch.arange(0, d_model, 2).float()
            * (-math.log(10000.0) / d_model)
        )
        pe[:, 0::2] = torch.sin(position * div_term)  # 偶数インデックス
        pe[:, 1::2] = torch.cos(position * div_term)  # 奇数インデックス
        pe = pe.unsqueeze(0)  # (1, max_len, d_model)
        self.register_buffer('pe', pe)

    def forward(self, x):
        # x: (batch, seq_len, d_model)
        return x + self.pe[:, :x.size(1), :]

# デモ
pe = PositionalEncoding(d_model=512)
x = torch.randn(2, 100, 512)  # batch=2, seq_len=100
output = pe(x)
print(f"Output shape: {output.shape}")  # (2, 100, 512)

4.5 Layer Normalization vs Batch Normalization

特徴Batch NormalizationLayer Normalization
正規化の対象バッチ次元特徴次元
バッチサイズへの依存あり — 十分なバッチサイズが必要なし — 各サンプルに対して独立に動作
使用される分野CNN(コンピュータビジョン)Transformer、RNN(NLP)
推論時の挙動移動統計量を使用直接計算 — 学習時と同じ
Transformerでの使用いいえはい — すべてのサブレイヤー

試験のヒント: 「TransformerがBatchNormではなくLayerNormを使う理由は?」という問題が出た場合→回答:(1) NLPではシーケンス長が可変であるため、バッチ統計が不安定になる、(2) LayerNormはバッチサイズに依存せず、小さなバッチサイズや推論時にもうまく動作する。

5. モデルファミリー

5.1 Encoderのみ:BERT

BERT(Bidirectional Encoder Representations from Transformers)はEncoderのみを使用します。以下の方法で事前学習されています:

  • MLM(Masked Language Modeling) — トークンの15%をマスクし、マスクされたトークンを予測する
  • NSP(Next Sentence Prediction) — 2つの文が連続しているかどうかを判定する

BERTは両方向を見ることができる(双方向性)ため、理解タスクに優れています:分類、NER、QA。

5.2 Decoderのみ:GPT

GPT(Generative Pre-trained Transformer)はcausal maskingを持つDecoderのみを使用します — 各トークンは以前のトークンにのみ注目できます。次トークン予測で事前学習されています。

Causal Attentionマスク(GPT):

トークン: [The]  [cat]  [sat]  [on]
The       ✓      ✗      ✗      ✗
cat       ✓      ✓      ✗      ✗
sat       ✓      ✓      ✓      ✗
on        ✓      ✓      ✓      ✓

✓ = 注目可能    ✗ = マスク済み(= softmax前に-inf)

→ 各トークンは前のトークンのみ「見る」ことができる
→ 生成に適している:次のトークンを予測

5.3 Encoder-Decoder:T5

T5(Text-to-Text Transfer Transformer)は完全なEncoder-Decoderアーキテクチャを使用します。すべてのタスクが「テキスト入力→テキスト出力」として定式化されます:

  • 翻訳:"translate English to French: The cat sat" → "Le chat s'est assis"
  • 要約:"summarize: {長いテキスト}" → "{要約}"
  • 分類:"classify: {テキスト}" → "positive"

5.4 比較表

特徴BERT(Encoder)GPT(Decoder)T5(Enc-Dec)
アーキテクチャEncoderのみDecoderのみEncoder-Decoder
方向性双方向左から右(causal)双方向enc + causal dec
事前学習MLM + NSP次トークン予測Span corruption
Attentionマスク全体(すべてのトークンを参照)Causal(過去のみ参照)全体enc + causal dec
得意分野理解:NER、QA、分類生成:テキスト、コードSeq2seq:翻訳、要約
出力文脈的埋め込み次トークンの確率ターゲットシーケンス
代表的なモデルBERT, RoBERTa, DeBERTaGPT-2/3/4, LLaMAT5, BART, mBART
決定木 — モデルファミリーの選択:

                ┌─ 長いテキストの生成が必要?
                │   はい → Decoderのみ(GPT, LLaMA)
                │
タスク ─────────┤
                │   ┌─ 入力→出力のシーケンス変換?
                │   │   はい → Encoder-Decoder(T5, BART)
                いいえ┤
                    │   いいえ → Encoderのみ(BERT)
                    │   (分類、NER、埋め込み)
                    └──────────────────────────────────

試験のヒント: よくある問題:「タスクXに対して、どのモデルファミリーを使うべきか?」クイックルール:(1) 理解/分類 → BERT、(2) 生成 → GPT、(3) Seq2seq(翻訳、要約)→ T5。注意:十分に大きなGPTはプロンプティングを通じてあらゆるタスクも処理できます。

6. トークン化

6.1 トークン vs 単語

言語モデルは「単語」ではなくトークンで動作します — トークンは単語以下のサイズの単位です。トークン化は、テキストをどのようにトークンに分割するかを決定します。

"unbelievable"のトークン化の例:

単語レベル:       ["unbelievable"]         → 語彙が大きすぎ、OOVが多い
文字レベル:       ["u","n","b","e",...]    → シーケンスが長すぎ
サブワード (BPE): ["un", "believ", "able"] → 語彙サイズとシーケンス長のバランス

6.2 BPE、WordPiece、SentencePiece

アルゴリズム使用モデル仕組み特徴
BPE(Byte-Pair Encoding)GPT-2/3/4, RoBERTa最も頻度の高いバイトペアを繰り返しマージボトムアップ、貪欲マージ
WordPieceBERT, DistilBERT尤度を最大化するペアをマージサブワードに##プレフィックスを使用
SentencePieceT5, LLaMA, mT5生テキストに対するUnigram LMまたはBPE言語非依存、事前トークン化不要
比較の例:

入力: "I love tokenization"

BPE (GPT-2):       ["I", " love", " token", "ization"]
WordPiece (BERT):   ["I", "love", "token", "##ization"]
SentencePiece (T5): ["▁I", "▁love", "▁token", "ization"]

WordPieceは接続に##を使用
SentencePieceは単語の開始に▁を使用

6.3 コード:HuggingFaceによるトークン化

from transformers import AutoTokenizer

# 各モデルファミリーのトークナイザーを読み込み
bert_tok = AutoTokenizer.from_pretrained("bert-base-uncased")
gpt2_tok = AutoTokenizer.from_pretrained("gpt2")
t5_tok = AutoTokenizer.from_pretrained("t5-small")

text = "Transformers are amazing for NLP tasks!"

# BERT (WordPiece)
bert_tokens = bert_tok.tokenize(text)
print(f"BERT:  {bert_tokens}")
# ['transformers', 'are', 'amazing', 'for', 'nl', '##p', 'tasks', '!']

# GPT-2 (BPE)
gpt2_tokens = gpt2_tok.tokenize(text)
print(f"GPT-2: {gpt2_tokens}")
# ['Trans', 'formers', 'Ġare', 'Ġamazing', 'Ġfor', 'ĠNLP', 'Ġtasks', '!']

# T5 (SentencePiece)
t5_tokens = t5_tok.tokenize(text)
print(f"T5:    {t5_tokens}")
# ['▁Transform', 'ers', '▁are', '▁amazing', '▁for', '▁NLP', '▁tasks', '!']

# エンコード → トークンID
ids = bert_tok.encode(text, return_tensors="pt")
print(f"Token IDs shape: {ids.shape}")

# テキストにデコード
decoded = bert_tok.decode(ids[0])
print(f"Decoded: {decoded}")

語彙サイズは埋め込みサイズとモデルの能力に影響します:

モデルトークナイザー語彙サイズ備考
BERT-baseWordPiece30,522小文字英語
GPT-2BPE50,257大文字小文字区別あり
T5SentencePiece32,100多言語対応
LLaMA-2SentencePiece32,000BPEの変種
GPT-4BPE (cl100k)100,256コード+多言語に最適化

7. NLPタスクのマッピング

7.1 タスク → モデル → 出力

NLPタスクタスク種別最適なモデルファミリー出力
テキスト分類シーケンス分類Encoder(BERT)単一ラベル
感情分析シーケンス分類Encoder(BERT)Positive/Negative
固有表現抽出トークン分類Encoder(BERT)トークンごとのラベル
質問応答抽出型/生成型EncoderまたはEnc-Decスパンまたはテキスト
要約Seq2seq生成Enc-Dec(T5, BART)要約テキスト
翻訳Seq2seq生成Enc-Dec(T5, mBART)翻訳テキスト
テキスト生成自己回帰Decoder(GPT)続きのテキスト
コード生成自己回帰Decoder(CodeGen)コード

7.2 コード:BERTをテキスト分類用にファインチューニング

import torch
import torch.nn as nn
from transformers import BertModel, BertTokenizer

class BertClassifier(nn.Module):
    def __init__(self, num_classes, model_name='bert-base-uncased'):
        super().__init__()
        self.bert = BertModel.from_pretrained(model_name)
        self.dropout = nn.Dropout(0.1)
        self.classifier = nn.Linear(self.bert.config.hidden_size, num_classes)

    def forward(self, input_ids, attention_mask):
        # BERTの出力: last_hidden_state, pooler_output
        outputs = self.bert(input_ids=input_ids,
                           attention_mask=attention_mask)

        # [CLS]トークンの表現を分類に使用
        cls_output = outputs.pooler_output  # (batch, hidden_size)
        cls_output = self.dropout(cls_output)
        logits = self.classifier(cls_output)  # (batch, num_classes)
        return logits

# 使用例
tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')
model = BertClassifier(num_classes=3)

# 入力をトークン化
text = "This movie was absolutely wonderful!"
encoded = tokenizer(text, return_tensors='pt', padding=True,
                    truncation=True, max_length=128)

# フォワードパス
logits = model(encoded['input_ids'], encoded['attention_mask'])
prediction = torch.argmax(logits, dim=-1)
print(f"Predicted class: {prediction.item()}")
# BERT分類器のトレーニングループ
from torch.utils.data import DataLoader

optimizer = torch.optim.AdamW(model.parameters(), lr=2e-5)
loss_fn = nn.CrossEntropyLoss()

model.train()
for epoch in range(3):
    total_loss = 0
    for batch in train_loader:
        optimizer.zero_grad()

        logits = model(batch['input_ids'], batch['attention_mask'])
        loss = loss_fn(logits, batch['labels'])

        loss.backward()
        torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)
        optimizer.step()

        total_loss += loss.item()

    print(f"Epoch {epoch+1}, Loss: {total_loss / len(train_loader):.4f}")

試験のヒント: BERTのファインチューニングで重要な3つのポイント:(1) 小さな学習率(2e-5〜5e-5)— モデルが既に事前学習済みのため、(2) 勾配クリッピング(clip_grad_norm_)で学習を安定化、(3) 分類にはpooler_output(CLSトークン)を使用し、トークンレベルタスク(NER)にはlast_hidden_stateを使用する。

8. チートシート

概念重要な公式/パターン覚えておくこと
Scaled Dot-Productsoftmax(QK^T / √d_k) Vsoftmaxの飽和を防ぐため√d_kで割る
Multi-Head分割 → h × Attention → 結合 → Lineard_k = d_model / num_heads
Positional Encodingsin/cos関数Diffusionのタイムステップ埋め込みでも再利用
Encoder(BERT)双方向、MLM理解タスク:NER、QA、分類
Decoder(GPT)Causalマスク、次トークン予測生成タスク:テキスト、コード
Enc-Dec(T5)Cross-attention、seq2seq翻訳、要約
BPE頻度の高いバイトペアをマージGPTファミリー
WordPiece尤度最大化マージBERTファミリー、##プレフィックスを使用
SentencePiece生テキストに対して言語非依存T5、LLaMA、▁プレフィックスを使用
Causalマスク下三角行列GPT:各トークンは前のトークンのみ参照可能
Layer Norm特徴次元で正規化Transformerで使用(BatchNormではない)
ファインチューニングLR2e-5 → 5e-5事前学習済み重みのため小さなLR

9. 練習問題

NVIDIA DLIに類似したコーディングアセスメント形式の問題:

Q1: scaled_dot_product_attention関数を実装してください。この関数はQ、K、Vテンソルとオプションのマスクを受け取り、出力とAttention重みを返します。

Q1の解答を表示
import torch
import torch.nn.functional as F
import math

def scaled_dot_product_attention(Q, K, V, mask=None):
    """
    Args:
        Q: (batch, seq_len, d_k)
        K: (batch, seq_len, d_k)
        V: (batch, seq_len, d_v)
        mask: optional (batch, 1, seq_len) or (batch, seq_len, seq_len)
    Returns:
        output: (batch, seq_len, d_v)
        attn_weights: (batch, seq_len, seq_len)
    """
    d_k = Q.size(-1)

    # Attentionスコアを計算
    scores = torch.matmul(Q, K.transpose(-2, -1)) / math.sqrt(d_k)

    # マスクがあれば適用
    if mask is not None:
        scores = scores.masked_fill(mask == 0, float('-inf'))

    # 最後の次元(Key次元)でsoftmax
    attn_weights = F.softmax(scores, dim=-1)

    # Valueの重み付き和
    output = torch.matmul(attn_weights, V)

    return output, attn_weights

# 検証:
B, S, D = 2, 4, 64
Q = torch.randn(B, S, D)
K = torch.randn(B, S, D)
V = torch.randn(B, S, D)
out, w = scaled_dot_product_attention(Q, K, V)
assert out.shape == (B, S, D)
assert w.shape == (B, S, S)
assert torch.allclose(w.sum(dim=-1), torch.ones(B, S), atol=1e-6)
print("All assertions passed!")

解説:重要なポイントは (1) 正しい行列乗算のためのK.transpose(-2, -1)、(2) スケーリングのためのmath.sqrt(d_k)による除算、(3) softmax前に-infでmasked_fill、(4) dim=-1でのsoftmax。よくあるエラー:Kの転置を忘れる、またはsoftmaxを間違った次元で適用する。

Q2: TransformerからPositional Encodingを削除するとどうなりますか?コードで証明してください。

Q2の解答を表示
import torch

# Positional EncodingなしのSelf-Attention
# → 出力は順列不変(トークンの順序は関係ない)

def self_attention_no_pos(x):
    """x: (batch, seq_len, d_model)"""
    d_k = x.size(-1)
    scores = torch.matmul(x, x.transpose(-2, -1)) / (d_k ** 0.5)
    weights = torch.softmax(scores, dim=-1)
    return torch.matmul(weights, x)

# 入力を作成
x = torch.randn(1, 4, 8)  # 4トークン、d_model=8

# 元の出力
out1 = self_attention_no_pos(x)

# トークンの順序をシャッフル: [0,1,2,3] → [2,0,3,1]
perm = [2, 0, 3, 1]
x_shuffled = x[:, perm, :]
out2 = self_attention_no_pos(x_shuffled)

# 確認: 出力も同じ順序でシャッフルされている
inv_perm = [1, 3, 0, 2]  # 逆順列
out2_reordered = out2[:, inv_perm, :]

print(f"Difference: {(out1 - out2_reordered).abs().max().item():.10f}")
# → ほぼ0! Attentionは順序を区別しない
# → "The cat sat on mat" = "mat on sat cat The"
# だからPositional Encodingが必須なのです!

解説:Positional Encodingがない場合、Self-Attentionは順列同変です — 「The cat sat」と「sat The cat」を同一に処理します。Positional Encodingはこの対称性を破り、モデルがトークンの順序を区別できるようにします。Diffusion Modelsでも、同じ概念がタイムステップ埋め込みに使われています。

Q3: GPTはcausal maskingを使用します。causal maskを作成するコードを書き、マスク行列の各要素を説明してください。

Q3の解答を表示
import torch

def create_causal_mask(seq_len):
    """
    Decoder用のcausal(先読み防止)マスクを作成。
    mask[i][j] = 1: トークンiがトークンjに注目できる(j <= i)
    mask[i][j] = 0: トークンiがトークンjを見ることができない(j > i)
    """
    mask = torch.tril(torch.ones(seq_len, seq_len))
    return mask

seq_len = 5
mask = create_causal_mask(seq_len)
print("Causal Mask:")
print(mask)
# tensor([[1., 0., 0., 0., 0.],
#         [1., 1., 0., 0., 0.],
#         [1., 1., 1., 0., 0.],
#         [1., 1., 1., 1., 0.],
#         [1., 1., 1., 1., 1.]])

# Attentionに適用
def causal_attention(Q, K, V):
    d_k = Q.size(-1)
    seq_len = Q.size(1)
    scores = torch.matmul(Q, K.transpose(-2, -1)) / (d_k ** 0.5)

    # causal maskを適用
    mask = create_causal_mask(seq_len).unsqueeze(0)  # (1, S, S)
    scores = scores.masked_fill(mask == 0, float('-inf'))
    # 結果: 未来の位置 → -inf → softmax → 0

    weights = torch.softmax(scores, dim=-1)
    print("Attention重み(causal):")
    print(weights[0].detach())
    # 行0: [1.0, 0.0, 0.0, 0.0, 0.0]  ← トークン0は自分自身のみ参照
    # 行1: [0.4, 0.6, 0.0, 0.0, 0.0]  ← トークン1はトークン0, 1を参照
    # 行4: [0.1, 0.2, 0.3, 0.2, 0.2]  ← トークン4はすべてのトークンを参照

    return torch.matmul(weights, V)

B, S, D = 1, 5, 32
Q = torch.randn(B, S, D)
K = torch.randn(B, S, D)
V = torch.randn(B, S, D)
out = causal_attention(Q, K, V)
print(f"Output shape: {out.shape}")  # (1, 5, 32)

解説:causal maskは下三角行列(torch.tril)です。mask[i][j]=0(j > i)の位置はsoftmax前に-infで埋められ、softmax後に0になります。これにより各トークンは前のトークンにのみ注目できます — GPTの自己回帰生成に不可欠です。

Q4: 以下のユースケースに対して、最も適切なモデルファミリーを選び、理由を説明してください:

  • (a) メールのスパム/非スパム分類
  • (b) ベトナム語から英語への翻訳
  • (c) 自由形式のテキスト生成チャットボット
  • (d) テキストからの人名抽出(NER)
Q4の解答を表示
# ユースケース → モデルファミリーのマッピング

tasks = {
    "(a) メールスパム分類": {
        "model_family": "Encoderのみ(BERT)",
        "reason": "分類タスク — メール全体を理解する必要がある"
                  "(双方向)。出力 = 1つのラベル(スパム/非スパム)。"
                  "BERT + Linear分類ヘッド。",
        "code_hint": "BertForSequenceClassification.from_pretrained('bert-base-uncased', num_labels=2)"
    },
    "(b) ベトナム語 → 英語翻訳": {
        "model_family": "Encoder-Decoder(T5, mBART)",
        "reason": "Seq2seqタスク — 入力シーケンス(ベトナム語)→ 出力"
                  "シーケンス(英語)。Encoderが入力を理解し、Decoderが"
                  "出力を生成。多言語対応のT5またはmBART。",
        "code_hint": "T5ForConditionalGeneration.from_pretrained('t5-base')"
    },
    "(c) 自由形式チャットボット": {
        "model_family": "Decoderのみ(GPT, LLaMA)",
        "reason": "自己回帰生成 — トークンごとにテキストを生成し、"
                  "別のEncoderは不要。チャットボット用途には"
                  "指示チューニング済みのGPT/LLaMA。",
        "code_hint": "AutoModelForCausalLM.from_pretrained('meta-llama/Llama-2-7b-chat-hf')"
    },
    "(d) 固有表現抽出": {
        "model_family": "Encoderのみ(BERT)",
        "reason": "トークン分類 — 各トークンにラベルを割り当てる必要がある"
                  "(B-PER, I-PER, O, B-LOC,...)。BERTの双方向Attentionにより"
                  "各トークンが両方向のコンテキストを参照可能。",
        "code_hint": "BertForTokenClassification.from_pretrained('bert-base-uncased', num_labels=9)"
    }
}

for task, info in tasks.items():
    print(f"\n{task}")
    print(f"  → {info['model_family']}")
    print(f"  理由: {info['reason']}")
    print(f"  コード: {info['code_hint']}")

解説:一般的なルール:(1) 出力が入力全体に対する1つのラベル → Encoder(BERT)、(2) 出力が各トークンに対するラベル → Encoder(BERT)+ トークン分類ヘッド、(3) 出力が異なる言語/形式のシーケンス → Encoder-Decoder(T5)、(4) 継続的なテキスト生成が必要 → Decoder(GPT)。ただし実際には、十分に大きなDecoder-only LLM(GPT-4、LLaMA-70B)はプロンプティングを通じてあらゆるタスクをうまく処理できます。