Chuyển đến nội dung chính

Bài 2: Transformer Architecture & Attention Mechanism

Self-attention, multi-head attention, positional encoding. Encoder-decoder architecture. BERT, GPT, T5 model families. Tokenization: BPE, WordPiece, SentencePiece. NLP tasks: classification, NER, QA, summarization.

1. Giới thiệu

Transformer là kiến trúc nền tảng đằng sau mọi mô hình Generative AI hiện đại — từ GPT, BERT, Stable Diffusion đến LLaMA. Trong bài assessment của NVIDIA DLI, bạn phải hiểu rõ cách Attention Mechanism hoạt động và có khả năng implement nó bằng PyTorch.

Bài học này sẽ đi từ Scaled Dot-Product Attention đến toàn bộ kiến trúc Transformer, sau đó mapping sang các model families và NLP tasks cụ thể.

Exam tip: NVIDIA DLI assessment thường yêu cầu bạn hoàn thành code cho attention mechanism hoặc debug lỗi dimension mismatch trong Transformer. Hãy nắm vững tensor shapes qua từng bước của attention — đây là chìa khóa để pass assessment.

Kiến trúc Transformer — Encoder-Decoder, Self-Attention, Cross-Attention
Kiến trúc Transformer — Encoder-Decoder, Self-Attention, Cross-Attention

2. Attention Mechanism

2.1 Trực giác — "Phần nào quan trọng nhất?"

Attention trả lời câu hỏi: "Khi xử lý token hiện tại, những token nào trong input sequence là quan trọng nhất?" Thay vì nén toàn bộ sequence vào một fixed-size vector như RNN, Attention cho phép mô hình "nhìn" trực tiếp vào mọi vị trí của input.

Cơ chế hoạt động qua 3 thành phần:

  • Query (Q) — "Tôi đang tìm gì?" — token hiện tại đặt câu hỏi
  • Key (K) — "Tôi chứa thông tin gì?" — mỗi token quảng cáo nội dung của mình
  • Value (V) — "Đây là thông tin thực sự" — nội dung được trả về khi match
Ví dụ: "The cat sat on the mat because it was tired"
                                          ↑
                                    Token "it" (Query)
                                          │
            ┌─────────────────────────────┤
            │     Attention scores:       │
            │   "cat"  = 0.72  ← cao!    │
            │   "mat"  = 0.11            │
            │   "sat"  = 0.08            │
            │   "The"  = 0.03            │
            │   ...                       │
            └─────────────────────────────┘
            → "it" attend chủ yếu vào "cat"

2.2 Scaled Dot-Product Attention

Công thức chính của attention:

Attention(Q, K, V) = softmax(Q · K^T / √d_k) · V

Trong đó:
  Q: Query matrix  — shape (seq_len, d_k)
  K: Key matrix    — shape (seq_len, d_k)
  V: Value matrix  — shape (seq_len, d_v)
  d_k: dimension của Key vectors
  √d_k: scaling factor để tránh gradient vanishing trong softmax

Tại sao cần scaling bằng √d_k? Khi d_k lớn, dot product Q·K^T có thể rất lớn, khiến softmax bão hòa → gradient gần bằng 0. Chia cho √d_k giữ variance ổn định.

Flow của Scaled Dot-Product Attention:

  Q ──┐
      │──→ MatMul ──→ Scale (÷√d_k) ──→ Mask (opt.) ──→ Softmax ──→ MatMul ──→ Output
  K ──┘                                                                ↑
                                                                       │
  V ────────────────────────────────────────────────────────────────────┘

Shapes (batch_size=B, seq_len=S, d_k=D):
  Q:        (B, S, D)
  K^T:      (B, D, S)
  Q·K^T:    (B, S, S)   ← attention score matrix
  softmax:  (B, S, S)   ← attention weights (mỗi hàng sum = 1)
  × V:      (B, S, D)   ← weighted output

2.3 Code: Implement Attention từ Scratch

import torch
import torch.nn.functional as F
import math

def scaled_dot_product_attention(Q, K, V, mask=None):
    """
    Q: (batch, seq_len, d_k)
    K: (batch, seq_len, d_k)
    V: (batch, seq_len, d_v)
    mask: (batch, 1, seq_len) or (batch, seq_len, seq_len)
    """
    d_k = Q.size(-1)

    # Step 1: Tính attention scores
    scores = torch.matmul(Q, K.transpose(-2, -1)) / math.sqrt(d_k)
    # scores shape: (batch, seq_len, seq_len)

    # Step 2: Apply mask (cho causal attention trong decoder)
    if mask is not None:
        scores = scores.masked_fill(mask == 0, float('-inf'))

    # Step 3: Softmax để có attention weights
    attn_weights = F.softmax(scores, dim=-1)
    # attn_weights shape: (batch, seq_len, seq_len)

    # Step 4: Weighted sum of values
    output = torch.matmul(attn_weights, V)
    # output shape: (batch, seq_len, d_v)

    return output, attn_weights

# Demo
batch_size, seq_len, d_k = 2, 5, 64
Q = torch.randn(batch_size, seq_len, d_k)
K = torch.randn(batch_size, seq_len, d_k)
V = torch.randn(batch_size, seq_len, d_k)

output, weights = scaled_dot_product_attention(Q, K, V)
print(f"Output shape: {output.shape}")    # (2, 5, 64)
print(f"Weights shape: {weights.shape}")  # (2, 5, 5)
print(f"Weights sum per row: {weights.sum(dim=-1)}")  # mỗi hàng = 1.0

Exam tip: Lỗi phổ biến nhất khi implement attention: quên K.transpose(-2, -1) hoặc chia sai dimension. Hãy luôn kiểm tra shape sau mỗi bước — scores phải có shape (batch, seq_len, seq_len).

3. Multi-Head Attention

3.1 Tại sao cần nhiều heads?

Một head duy nhất chỉ học được một loại relationship. Multi-Head Attention cho phép mô hình attend đồng thời vào nhiều representation subspaces khác nhau:

  • Head 1: học syntactic relationships (subject-verb)
  • Head 2: học coreference (pronoun → noun)
  • Head 3: học positional proximity
  • Head 4: học semantic similarity
Multi-Head Attention Flow:

Input (batch, seq_len, d_model)
    │
    ├── Linear → Q ──┐
    ├── Linear → K ──┼── Split thành h heads
    └── Linear → V ──┘
                      │
        ┌─────────────┼─────────────┐
        ▼             ▼             ▼
    Head 1        Head 2    ...  Head h
 Attention()   Attention()    Attention()
        │             │             │
        └─────────────┼─────────────┘
                      │
                  Concatenate
                      │
                   Linear
                      │
                   Output (batch, seq_len, d_model)

Shapes (d_model=512, h=8, d_k = d_model/h = 64):
  Input:          (B, S, 512)
  Per-head Q/K/V: (B, h, S, 64)   ← reshape sau linear
  Per-head out:   (B, h, S, 64)
  Concat:         (B, S, 512)      ← h × d_k = d_model
  Final output:   (B, S, 512)

3.2 Code: MultiHeadAttention Module

import torch
import torch.nn as nn
import math

class MultiHeadAttention(nn.Module):
    def __init__(self, d_model, num_heads):
        super().__init__()
        assert d_model % num_heads == 0, "d_model phải chia hết cho num_heads"

        self.d_model = d_model
        self.num_heads = num_heads
        self.d_k = d_model // num_heads

        # Linear projections cho Q, K, V và output
        self.W_q = nn.Linear(d_model, d_model)
        self.W_k = nn.Linear(d_model, d_model)
        self.W_v = nn.Linear(d_model, d_model)
        self.W_o = nn.Linear(d_model, d_model)

    def forward(self, query, key, value, mask=None):
        batch_size = query.size(0)

        # Step 1: Linear projections
        Q = self.W_q(query)  # (B, S, d_model)
        K = self.W_k(key)
        V = self.W_v(value)

        # Step 2: Reshape thành multi-head format
        # (B, S, d_model) → (B, S, h, d_k) → (B, h, S, d_k)
        Q = Q.view(batch_size, -1, self.num_heads, self.d_k).transpose(1, 2)
        K = K.view(batch_size, -1, self.num_heads, self.d_k).transpose(1, 2)
        V = V.view(batch_size, -1, self.num_heads, self.d_k).transpose(1, 2)

        # Step 3: Scaled dot-product attention cho mỗi head
        scores = torch.matmul(Q, K.transpose(-2, -1)) / math.sqrt(self.d_k)

        if mask is not None:
            scores = scores.masked_fill(mask == 0, float('-inf'))

        attn_weights = torch.softmax(scores, dim=-1)
        context = torch.matmul(attn_weights, V)
        # context: (B, h, S, d_k)

        # Step 4: Concatenate heads
        # (B, h, S, d_k) → (B, S, h, d_k) → (B, S, d_model)
        context = context.transpose(1, 2).contiguous().view(
            batch_size, -1, self.d_model
        )

        # Step 5: Final linear projection
        output = self.W_o(context)
        return output

# Demo
d_model, num_heads = 512, 8
mha = MultiHeadAttention(d_model, num_heads)

x = torch.randn(2, 10, d_model)  # batch=2, seq_len=10
output = mha(x, x, x)  # self-attention: Q=K=V=x
print(f"Output shape: {output.shape}")  # (2, 10, 512)
ParameterGiá trị thường gặpGhi chú
d_model512, 768, 1024Kích thước embedding
num_heads8, 12, 16d_model phải chia hết cho num_heads
d_k = d_model / h64, 64, 64Mỗi head thường có d_k = 64
Total params (MHA)4 × d_model²4 linear layers: W_q, W_k, W_v, W_o

Exam tip: Trong DLI assessment, nếu bạn gặp lỗi RuntimeError: shape mismatch ở attention, hãy kiểm tra: (1) d_model % num_heads == 0, (2) view() và transpose() đúng thứ tự, (3) nhớ gọi .contiguous() trước .view() sau transpose.

4. Transformer Architecture

4.1 Encoder Block

Mỗi layer trong Transformer Encoder gồm 2 sub-layers với residual connections và layer normalization:

Transformer Encoder Block:

  Input
    │
    ▼
┌─────────────────────────────┐
│  Multi-Head Self-Attention  │
└──────────────┬──────────────┘
               │
    ┌──────────┴──────────┐
    │     Add & Norm      │ ← x + Sublayer(x), rồi LayerNorm
    └──────────┬──────────┘
               │
    ┌──────────┴──────────┐
    │  Feed-Forward Net   │ ← 2 linear layers + ReLU/GELU
    │  FFN(x) = W₂·σ(W₁x + b₁) + b₂
    └──────────┬──────────┘
               │
    ┌──────────┴──────────┐
    │     Add & Norm      │
    └──────────┬──────────┘
               │
             Output

4.2 Decoder Block

Decoder có thêm một sub-layer Cross-Attention — attend vào output của encoder:

Transformer Decoder Block:

  Input (shifted right)
    │
    ▼
┌────────────────────────────────┐
│  Masked Multi-Head Attention   │ ← causal mask: chỉ nhìn tokens trước
└──────────────┬─────────────────┘
               │
    ┌──────────┴──────────┐
    │     Add & Norm      │
    └──────────┬──────────┘
               │
    ┌──────────┴──────────────────────────┐
    │  Cross Multi-Head Attention         │ ← Q từ decoder, K/V từ encoder
    │  Q = decoder hidden states          │
    │  K, V = encoder output              │
    └──────────────────┬──────────────────┘
               │
    ┌──────────┴──────────┐
    │     Add & Norm      │
    └──────────┬──────────┘
               │
    ┌──────────┴──────────┐
    │  Feed-Forward Net   │
    └──────────┬──────────┘
               │
    ┌──────────┴──────────┐
    │     Add & Norm      │
    └──────────┬──────────┘
               │
             Output

4.3 Full Transformer Architecture

┌─────────────────────────────────────────────────────────────────┐
│                    TRANSFORMER ARCHITECTURE                     │
├───────────────────────┬─────────────────────────────────────────┤
│       ENCODER         │              DECODER                    │
│                       │                                         │
│   Input Embedding     │         Output Embedding                │
│        +              │              +                          │
│   Positional Enc.     │         Positional Enc.                 │
│        │              │              │                          │
│   ┌────┴────┐         │         ┌────┴──────────┐              │
│   │  MH     │         │         │  Masked MH    │              │
│   │  Self-  │         │         │  Self-        │              │
│   │  Attn   │         │         │  Attention    │              │
│   └────┬────┘         │         └────┬──────────┘              │
│   Add & Norm          │         Add & Norm                     │
│        │              │              │                          │
│   ┌────┴────┐         │    ┌────────┴────────────┐             │
│   │  Feed   │         │    │  Cross MH Attention │             │
│   │ Forward │    ─────┼───►│  Q=dec, K/V=enc    │             │
│   │   Net   │         │    └─────────┬───────────┘             │
│   └────┬────┘         │         Add & Norm                     │
│   Add & Norm          │              │                          │
│        │              │         ┌────┴────┐                    │
│      × N layers       │         │  Feed   │                    │
│        │              │         │ Forward │                    │
│     Encoder           │         └────┬────┘                    │
│     Output            │         Add & Norm                     │
│                       │              │                          │
│                       │           × N layers                   │
│                       │              │                          │
│                       │         Linear + Softmax               │
│                       │              │                          │
│                       │       Output Probabilities              │
└───────────────────────┴─────────────────────────────────────────┘

4.4 Positional Encoding

Transformer không có khái niệm "thứ tự" như RNN. Positional Encoding thêm thông tin vị trí vào embedding bằng hàm sin/cos:

PE(pos, 2i)   = sin(pos / 10000^(2i/d_model))
PE(pos, 2i+1) = cos(pos / 10000^(2i/d_model))

pos: vị trí token trong sequence (0, 1, 2, ...)
i:   chỉ số dimension (0, 1, 2, ..., d_model/2)

Đây chính xác là concept được reuse trong Diffusion Models — sinusoidal embeddings để encode timestep t. Nắm vững positional encoding ở đây sẽ giúp bạn rất nhiều ở phần Diffusion.

import torch
import math

class PositionalEncoding(nn.Module):
    def __init__(self, d_model, max_len=5000):
        super().__init__()
        pe = torch.zeros(max_len, d_model)
        position = torch.arange(0, max_len).unsqueeze(1).float()
        div_term = torch.exp(
            torch.arange(0, d_model, 2).float()
            * (-math.log(10000.0) / d_model)
        )
        pe[:, 0::2] = torch.sin(position * div_term)  # even indices
        pe[:, 1::2] = torch.cos(position * div_term)  # odd indices
        pe = pe.unsqueeze(0)  # (1, max_len, d_model)
        self.register_buffer('pe', pe)

    def forward(self, x):
        # x: (batch, seq_len, d_model)
        return x + self.pe[:, :x.size(1), :]

# Demo
pe = PositionalEncoding(d_model=512)
x = torch.randn(2, 100, 512)  # batch=2, seq_len=100
output = pe(x)
print(f"Output shape: {output.shape}")  # (2, 100, 512)

4.5 Layer Normalization vs Batch Normalization

FeatureBatch NormalizationLayer Normalization
Normalize acrossBatch dimensionFeature dimension
Phụ thuộc batch sizeCó — cần batch đủ lớnKhông — hoạt động trên từng sample
Dùng trongCNN (computer vision)Transformer, RNN (NLP)
Inference behaviorDùng running statisticsTính trực tiếp — giống train
Transformer sử dụngKhôngCó — mọi sub-layer

Exam tip: Nếu câu hỏi hỏi "vì sao Transformer dùng LayerNorm thay vì BatchNorm?" → trả lời: (1) NLP có variable sequence length → batch stats không ổn định, (2) LayerNorm không phụ thuộc batch size, hoạt động tốt với batch size nhỏ và inference.

5. Model Families

5.1 Encoder-only: BERT

BERT (Bidirectional Encoder Representations from Transformers) chỉ sử dụng phần Encoder. Pre-training bằng:

  • MLM (Masked Language Modeling) — che 15% tokens, dự đoán token bị che
  • NSP (Next Sentence Prediction) — hai câu có liên tiếp không?

Vì BERT nhìn được cả hai hướng (bidirectional), nó rất mạnh cho các task understanding: classification, NER, QA.

5.2 Decoder-only: GPT

GPT (Generative Pre-trained Transformer) chỉ sử dụng Decoder với causal masking — mỗi token chỉ attend vào các token trước nó. Pre-training bằng next-token prediction.

Causal Attention Mask (GPT):

Token:   [The]  [cat]  [sat]  [on]
The       ✓      ✗      ✗      ✗
cat       ✓      ✓      ✗      ✗
sat       ✓      ✓      ✓      ✗
on        ✓      ✓      ✓      ✓

✓ = có thể attend    ✗ = bị mask (= -inf trước softmax)

→ Mỗi token chỉ "thấy" tokens trước nó
→ Phù hợp cho generation: sinh token tiếp theo

5.3 Encoder-Decoder: T5

T5 (Text-to-Text Transfer Transformer) sử dụng đầy đủ kiến trúc Encoder-Decoder. Mọi task đều được frame thành "text-in → text-out":

  • Translation: "translate English to French: The cat sat" → "Le chat s'est assis"
  • Summarization: "summarize: {long text}" → "{summary}"
  • Classification: "classify: {text}" → "positive"

5.4 Comparison Table

FeatureBERT (Encoder)GPT (Decoder)T5 (Enc-Dec)
ArchitectureEncoder-onlyDecoder-onlyEncoder-Decoder
DirectionalityBidirectionalLeft-to-right (causal)Bidirectional enc + causal dec
Pre-trainingMLM + NSPNext-token predictionSpan corruption
Attention maskFull (nhìn mọi token)Causal (chỉ nhìn trước)Full enc + causal dec
Mạnh choUnderstanding: NER, QA, classificationGeneration: text, codeSeq2seq: translation, summary
OutputContextual embeddingsNext token probabilityTarget sequence
Ví dụ modelsBERT, RoBERTa, DeBERTaGPT-2/3/4, LLaMAT5, BART, mBART
Decision Tree — Chọn Model Family:

                ┌─ Cần sinh text dài?
                │   YES → Decoder-only (GPT, LLaMA)
                │
Task ───────────┤
                │   ┌─ Input→Output seq?
                │   │   YES → Encoder-Decoder (T5, BART)
                NO ─┤
                    │   NO → Encoder-only (BERT)
                    │   (classification, NER, embedding)
                    └──────────────────────────────────

Exam tip: Câu hỏi thường gặp: "Cho task X, nên dùng model family nào?" Rule nhanh: (1) Understanding / classification → BERT, (2) Generation → GPT, (3) Seq2seq (translation, summarization) → T5. Nhưng lưu ý: GPT đủ lớn cũng làm được mọi task qua prompting.

6. Tokenization

6.1 Token vs Word

Mô hình ngôn ngữ không làm việc với "từ" mà với tokens — đơn vị nhỏ hơn hoặc bằng một từ. Tokenization quyết định cách chia text thành tokens.

Ví dụ tokenization cho "unbelievable":

Word-level:     ["unbelievable"]         → vocab quá lớn, OOV nhiều
Character:      ["u","n","b","e",...]    → sequence quá dài
Subword (BPE):  ["un", "believ", "able"] → cân bằng vocab size và seq length

6.2 BPE, WordPiece, SentencePiece

AlgorithmSử dụng bởiCách hoạt độngĐặc điểm
BPE (Byte-Pair Encoding)GPT-2/3/4, RoBERTaMerge cặp byte phổ biến nhất lặp lạiBottom-up, greedy merging
WordPieceBERT, DistilBERTMerge pair maximize likelihoodDùng ## prefix cho subword
SentencePieceT5, LLaMA, mT5Unigram LM hoặc BPE trên raw textLanguage-agnostic, không cần pre-tokenize
Ví dụ so sánh:

Input: "I love tokenization"

BPE (GPT-2):       ["I", " love", " token", "ization"]
WordPiece (BERT):   ["I", "love", "token", "##ization"]
SentencePiece (T5): ["▁I", "▁love", "▁token", "ization"]

WordPiece dùng ## cho continuation
SentencePiece dùng ▁ cho word start

6.3 Code: Tokenization với HuggingFace

from transformers import AutoTokenizer

# Load tokenizers của mỗi model family
bert_tok = AutoTokenizer.from_pretrained("bert-base-uncased")
gpt2_tok = AutoTokenizer.from_pretrained("gpt2")
t5_tok = AutoTokenizer.from_pretrained("t5-small")

text = "Transformers are amazing for NLP tasks!"

# BERT (WordPiece)
bert_tokens = bert_tok.tokenize(text)
print(f"BERT:  {bert_tokens}")
# ['transformers', 'are', 'amazing', 'for', 'nl', '##p', 'tasks', '!']

# GPT-2 (BPE)
gpt2_tokens = gpt2_tok.tokenize(text)
print(f"GPT-2: {gpt2_tokens}")
# ['Trans', 'formers', 'Ġare', 'Ġamazing', 'Ġfor', 'ĠNLP', 'Ġtasks', '!']

# T5 (SentencePiece)
t5_tokens = t5_tok.tokenize(text)
print(f"T5:    {t5_tokens}")
# ['▁Transform', 'ers', '▁are', '▁amazing', '▁for', '▁NLP', '▁tasks', '!']

# Encode → token IDs
ids = bert_tok.encode(text, return_tensors="pt")
print(f"Token IDs shape: {ids.shape}")

# Decode ngược lại
decoded = bert_tok.decode(ids[0])
print(f"Decoded: {decoded}")

Vocab size ảnh hưởng đến embedding size và model capacity:

ModelTokenizerVocab SizeGhi chú
BERT-baseWordPiece30,522Lowercase English
GPT-2BPE50,257Case-sensitive
T5SentencePiece32,100Multilingual capable
LLaMA-2SentencePiece32,000BPE variant
GPT-4BPE (cl100k)100,256Optimized for code + multilingual

7. NLP Tasks Mapping

7.1 Task → Model → Output

NLP TaskTask TypeBest Model FamilyOutput
Text ClassificationSequence classificationEncoder (BERT)Single label
Sentiment AnalysisSequence classificationEncoder (BERT)Positive/Negative
Named Entity RecognitionToken classificationEncoder (BERT)Label per token
Question AnsweringExtractive / GenerativeEncoder or Enc-DecSpan or text
SummarizationSeq2seq generationEnc-Dec (T5, BART)Summary text
TranslationSeq2seq generationEnc-Dec (T5, mBART)Translated text
Text GenerationAutoregressiveDecoder (GPT)Continuation
Code GenerationAutoregressiveDecoder (CodeGen)Code

7.2 Code: Fine-tune BERT cho Text Classification

import torch
import torch.nn as nn
from transformers import BertModel, BertTokenizer

class BertClassifier(nn.Module):
    def __init__(self, num_classes, model_name='bert-base-uncased'):
        super().__init__()
        self.bert = BertModel.from_pretrained(model_name)
        self.dropout = nn.Dropout(0.1)
        self.classifier = nn.Linear(self.bert.config.hidden_size, num_classes)

    def forward(self, input_ids, attention_mask):
        # BERT output: last_hidden_state, pooler_output
        outputs = self.bert(input_ids=input_ids,
                           attention_mask=attention_mask)

        # Lấy [CLS] token representation cho classification
        cls_output = outputs.pooler_output  # (batch, hidden_size)
        cls_output = self.dropout(cls_output)
        logits = self.classifier(cls_output)  # (batch, num_classes)
        return logits

# Sử dụng
tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')
model = BertClassifier(num_classes=3)

# Tokenize input
text = "This movie was absolutely wonderful!"
encoded = tokenizer(text, return_tensors='pt', padding=True,
                    truncation=True, max_length=128)

# Forward pass
logits = model(encoded['input_ids'], encoded['attention_mask'])
prediction = torch.argmax(logits, dim=-1)
print(f"Predicted class: {prediction.item()}")
# Training loop cho BERT classifier
from torch.utils.data import DataLoader

optimizer = torch.optim.AdamW(model.parameters(), lr=2e-5)
loss_fn = nn.CrossEntropyLoss()

model.train()
for epoch in range(3):
    total_loss = 0
    for batch in train_loader:
        optimizer.zero_grad()

        logits = model(batch['input_ids'], batch['attention_mask'])
        loss = loss_fn(logits, batch['labels'])

        loss.backward()
        torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)
        optimizer.step()

        total_loss += loss.item()

    print(f"Epoch {epoch+1}, Loss: {total_loss / len(train_loader):.4f}")

Exam tip: Khi fine-tune BERT, 3 điểm quan trọng: (1) Learning rate nhỏ (2e-5 đến 5e-5) vì model đã pre-trained, (2) Gradient clipping (clip_grad_norm_) để ổn định training, (3) Dùng pooler_output (CLS token) cho classification, last_hidden_state cho token-level tasks (NER).

8. Cheat Sheet

ConceptKey Formula / PatternGhi nhớ
Scaled Dot-Productsoftmax(QK^T / √d_k) VChia √d_k để tránh softmax bão hòa
Multi-HeadSplit → h × Attention → Concat → Lineard_k = d_model / num_heads
Positional Encodingsin/cos functionsReused trong Diffusion timestep embedding
Encoder (BERT)Bidirectional, MLMUnderstanding tasks: NER, QA, classification
Decoder (GPT)Causal mask, next-tokenGeneration tasks: text, code
Enc-Dec (T5)Cross-attention, seq2seqTranslation, summarization
BPEMerge frequent byte pairsGPT family
WordPieceMaximize likelihood mergeBERT family, dùng ## prefix
SentencePieceLanguage-agnostic on raw textT5, LLaMA, dùng ▁ prefix
Causal MaskLower-triangular matrixGPT: mỗi token chỉ thấy trước nó
Layer NormNormalize across featuresDùng trong Transformer (không phải BatchNorm)
Fine-tune LR2e-5 → 5e-5LR nhỏ vì pre-trained weights

9. Practice Questions

Các câu hỏi dạng coding assessment tương tự NVIDIA DLI:

Q1: Implement scaled_dot_product_attention function. Hàm nhận Q, K, V tensors và optional mask, trả về output và attention weights.

Xem đáp án Q1
import torch
import torch.nn.functional as F
import math

def scaled_dot_product_attention(Q, K, V, mask=None):
    """
    Args:
        Q: (batch, seq_len, d_k)
        K: (batch, seq_len, d_k)
        V: (batch, seq_len, d_v)
        mask: optional (batch, 1, seq_len) or (batch, seq_len, seq_len)
    Returns:
        output: (batch, seq_len, d_v)
        attn_weights: (batch, seq_len, seq_len)
    """
    d_k = Q.size(-1)

    # Compute attention scores
    scores = torch.matmul(Q, K.transpose(-2, -1)) / math.sqrt(d_k)

    # Apply mask if provided
    if mask is not None:
        scores = scores.masked_fill(mask == 0, float('-inf'))

    # Softmax over last dimension (key dimension)
    attn_weights = F.softmax(scores, dim=-1)

    # Weighted sum of values
    output = torch.matmul(attn_weights, V)

    return output, attn_weights

# Kiểm tra:
B, S, D = 2, 4, 64
Q = torch.randn(B, S, D)
K = torch.randn(B, S, D)
V = torch.randn(B, S, D)
out, w = scaled_dot_product_attention(Q, K, V)
assert out.shape == (B, S, D)
assert w.shape == (B, S, S)
assert torch.allclose(w.sum(dim=-1), torch.ones(B, S), atol=1e-6)
print("All assertions passed!")

Giải thích: Điểm mấu chốt là (1) K.transpose(-2, -1) để matrix multiply đúng, (2) chia math.sqrt(d_k) để scale, (3) masked_fill với -inf trước softmax, (4) softmax trên dim=-1. Lỗi phổ biến: quên transpose K hoặc softmax sai dim.

Q2: Điều gì xảy ra nếu loại bỏ Positional Encoding khỏi Transformer? Viết code chứng minh.

Xem đáp án Q2
import torch

# Self-attention KHÔNG có positional encoding
# → output là permutation invariant (thứ tự token không ảnh hưởng)

def self_attention_no_pos(x):
    """x: (batch, seq_len, d_model)"""
    d_k = x.size(-1)
    scores = torch.matmul(x, x.transpose(-2, -1)) / (d_k ** 0.5)
    weights = torch.softmax(scores, dim=-1)
    return torch.matmul(weights, x)

# Tạo input
x = torch.randn(1, 4, 8)  # 4 tokens, d_model=8

# Output gốc
out1 = self_attention_no_pos(x)

# Shuffle thứ tự tokens: [0,1,2,3] → [2,0,3,1]
perm = [2, 0, 3, 1]
x_shuffled = x[:, perm, :]
out2 = self_attention_no_pos(x_shuffled)

# Kiểm tra: output cũng bị shuffle theo cùng thứ tự
inv_perm = [1, 3, 0, 2]  # inverse permutation
out2_reordered = out2[:, inv_perm, :]

print(f"Difference: {(out1 - out2_reordered).abs().max().item():.10f}")
# → gần 0! Attention không phân biệt thứ tự
# → "The cat sat on mat" = "mat on sat cat The"
# Đó là lý do PHẢI có Positional Encoding!

Giải thích: Không có Positional Encoding, self-attention là permutation equivariant — nó xử lý "The cat sat" giống hệt "sat The cat". Positional Encoding phá vỡ tính đối xứng này, cho phép model phân biệt thứ tự tokens. Trong Diffusion Models, cùng concept này được dùng cho timestep embedding.

Q3: GPT sử dụng causal masking. Hãy viết code tạo causal mask và giải thích mỗi phần tử trong ma trận mask.

Xem đáp án Q3
import torch

def create_causal_mask(seq_len):
    """
    Tạo causal (look-ahead) mask cho decoder.
    mask[i][j] = 1 nếu token i được attend vào token j (j <= i)
    mask[i][j] = 0 nếu token i KHÔNG được nhìn token j (j > i)
    """
    mask = torch.tril(torch.ones(seq_len, seq_len))
    return mask

seq_len = 5
mask = create_causal_mask(seq_len)
print("Causal Mask:")
print(mask)
# tensor([[1., 0., 0., 0., 0.],
#         [1., 1., 0., 0., 0.],
#         [1., 1., 1., 0., 0.],
#         [1., 1., 1., 1., 0.],
#         [1., 1., 1., 1., 1.]])

# Áp dụng vào attention
def causal_attention(Q, K, V):
    d_k = Q.size(-1)
    seq_len = Q.size(1)
    scores = torch.matmul(Q, K.transpose(-2, -1)) / (d_k ** 0.5)

    # Apply causal mask
    mask = create_causal_mask(seq_len).unsqueeze(0)  # (1, S, S)
    scores = scores.masked_fill(mask == 0, float('-inf'))
    # Kết quả: vị trí tương lai → -inf → softmax → 0

    weights = torch.softmax(scores, dim=-1)
    print("Attention weights (causal):")
    print(weights[0].detach())
    # Row 0: [1.0, 0.0, 0.0, 0.0, 0.0]  ← token 0 chỉ thấy chính nó
    # Row 1: [0.4, 0.6, 0.0, 0.0, 0.0]  ← token 1 thấy token 0,1
    # Row 4: [0.1, 0.2, 0.3, 0.2, 0.2]  ← token 4 thấy tất cả

    return torch.matmul(weights, V)

B, S, D = 1, 5, 32
Q = torch.randn(B, S, D)
K = torch.randn(B, S, D)
V = torch.randn(B, S, D)
out = causal_attention(Q, K, V)
print(f"Output shape: {out.shape}")  # (1, 5, 32)

Giải thích: Causal mask là ma trận tam giác dưới (torch.tril). Vị trí mask[i][j]=0 (j > i) được fill bằng -inf trước softmax, biến thành 0 sau softmax. Điều này đảm bảo mỗi token chỉ attend vào các token trước nó — essential cho autoregressive generation trong GPT.

Q4: Cho các use cases sau, hãy chọn model family phù hợp nhất và giải thích lý do:

  • (a) Phân loại email spam/not-spam
  • (b) Dịch tiếng Việt sang tiếng Anh
  • (c) Chatbot tạo text tự do
  • (d) Trích xuất tên người từ văn bản (NER)
Xem đáp án Q4
# Mapping use cases → model families

tasks = {
    "(a) Email spam classification": {
        "model_family": "Encoder-only (BERT)",
        "reason": "Classification task — cần hiểu toàn bộ email "
                  "(bidirectional). Output = 1 label (spam/not-spam). "
                  "BERT + Linear classifier head.",
        "code_hint": "BertForSequenceClassification.from_pretrained('bert-base-uncased', num_labels=2)"
    },
    "(b) Vietnamese → English translation": {
        "model_family": "Encoder-Decoder (T5, mBART)",
        "reason": "Seq2seq task — input sequence (Vietnamese) → output "
                  "sequence (English). Encoder hiểu input, decoder sinh "
                  "output. T5 hoặc mBART cho multilingual.",
        "code_hint": "T5ForConditionalGeneration.from_pretrained('t5-base')"
    },
    "(c) Free-form chatbot": {
        "model_family": "Decoder-only (GPT, LLaMA)",
        "reason": "Autoregressive generation — sinh text token-by-token, "
                  "không cần encoder riêng. GPT/LLaMA với instruction "
                  "tuning cho chatbot use case.",
        "code_hint": "AutoModelForCausalLM.from_pretrained('meta-llama/Llama-2-7b-chat-hf')"
    },
    "(d) Named Entity Recognition": {
        "model_family": "Encoder-only (BERT)",
        "reason": "Token classification — cần gán label cho TỪNG token "
                  "(B-PER, I-PER, O, B-LOC,...). BERT bidirectional "
                  "giúp mỗi token nhìn context cả hai phía.",
        "code_hint": "BertForTokenClassification.from_pretrained('bert-base-uncased', num_labels=9)"
    }
}

for task, info in tasks.items():
    print(f"\n{task}")
    print(f"  → {info['model_family']}")
    print(f"  Lý do: {info['reason']}")
    print(f"  Code: {info['code_hint']}")

Giải thích: Rule tổng quát: (1) Nếu output là 1 label cho toàn bộ input → Encoder (BERT), (2) Nếu output là label cho mỗi token → Encoder (BERT) + token classification head, (3) Nếu output là sequence khác input language/format → Encoder-Decoder (T5), (4) Nếu cần sinh text liên tục → Decoder (GPT). Tuy nhiên, trong thực tế, LLM decoder-only đủ lớn (GPT-4, LLaMA-70B) có thể làm tốt mọi task qua prompting.

Q5: Debug lỗi sau trong Transformer. Code có bug ở dimension — tìm và sửa:

# BUG CODE — Tìm và sửa lỗi
class BrokenMultiHeadAttention(nn.Module):
    def __init__(self, d_model=512, num_heads=8):
        super().__init__()
        self.d_model = d_model
        self.num_heads = num_heads
        self.d_k = d_model // num_heads  # 64

        self.W_q = nn.Linear(d_model, d_model)
        self.W_k = nn.Linear(d_model, d_model)
        self.W_v = nn.Linear(d_model, d_model)
        self.W_o = nn.Linear(d_model, d_model)

    def forward(self, x, mask=None):
        B = x.size(0)
        Q = self.W_q(x)
        K = self.W_k(x)
        V = self.W_v(x)

        # BUG: reshape sai thứ tự dimensions
        Q = Q.view(B, self.num_heads, -1, self.d_k)  # ← Sai!
        K = K.view(B, self.num_heads, -1, self.d_k)
        V = V.view(B, self.num_heads, -1, self.d_k)

        scores = torch.matmul(Q, K.transpose(-2, -1)) / (self.d_k ** 0.5)
        weights = torch.softmax(scores, dim=-1)
        context = torch.matmul(weights, V)

        # BUG: quên contiguous() trước view
        context = context.transpose(1, 2).view(B, -1, self.d_model)  # ← Sai!
        return self.W_o(context)
Xem đáp án Q5
import torch
import torch.nn as nn

class FixedMultiHeadAttention(nn.Module):
    def __init__(self, d_model=512, num_heads=8):
        super().__init__()
        self.d_model = d_model
        self.num_heads = num_heads
        self.d_k = d_model // num_heads

        self.W_q = nn.Linear(d_model, d_model)
        self.W_k = nn.Linear(d_model, d_model)
        self.W_v = nn.Linear(d_model, d_model)
        self.W_o = nn.Linear(d_model, d_model)

    def forward(self, x, mask=None):
        B = x.size(0)
        Q = self.W_q(x)
        K = self.W_k(x)
        V = self.W_v(x)

        # FIX 1: Đúng thứ tự: (B, S, d_model) → (B, S, h, d_k) → (B, h, S, d_k)
        # Phải view thành (B, S, h, d_k) TRƯỚC, rồi transpose(1,2)
        Q = Q.view(B, -1, self.num_heads, self.d_k).transpose(1, 2)
        K = K.view(B, -1, self.num_heads, self.d_k).transpose(1, 2)
        V = V.view(B, -1, self.num_heads, self.d_k).transpose(1, 2)

        scores = torch.matmul(Q, K.transpose(-2, -1)) / (self.d_k ** 0.5)
        weights = torch.softmax(scores, dim=-1)
        context = torch.matmul(weights, V)  # (B, h, S, d_k)

        # FIX 2: Thêm .contiguous() sau transpose trước .view()
        context = context.transpose(1, 2).contiguous().view(B, -1, self.d_model)
        return self.W_o(context)

# Verify
model = FixedMultiHeadAttention(d_model=512, num_heads=8)
x = torch.randn(2, 10, 512)
out = model(x)
print(f"Output shape: {out.shape}")  # (2, 10, 512) ✓
assert out.shape == (2, 10, 512)
print("Fixed! All correct.")

Giải thích: Có 2 lỗi: Bug 1: view(B, num_heads, -1, d_k) sai vì tensor layout trong memory là (B, S, d_model). Phải view thành (B, S, num_heads, d_k) trước rồi transpose(1, 2) để có (B, num_heads, S, d_k). View trực tiếp thành (B, num_heads, S, d_k) sẽ trộn lẫn data giữa các heads. Bug 2: Sau transpose(1, 2), tensor không còn contiguous trong memory. Gọi .view() trên non-contiguous tensor gây RuntimeError. Phải thêm .contiguous() trước .view().