Chuyển đến nội dung chính

Lesson 2: Transformer Architecture & Attention Mechanism

Self-attention, multi-head attention, positional encoding. Encoder-decoder architecture. BERT, GPT, T5 model families. Tokenization: BPE, WordPiece, SentencePiece. NLP tasks: classification, NER, QA, summarization.

1. Introduction

Transformer is the foundational architecture behind every modern Generative AI model — from GPT, BERT, Stable Diffusion to LLaMA. In the NVIDIA DLI assessment, you must thoroughly understand how the Attention Mechanism works and be able to implement it in PyTorch.

This lesson covers everything from Scaled Dot-Product Attention to the full Transformer architecture, then maps to specific model families and NLP tasks.

Exam tip: The NVIDIA DLI assessment often requires you to complete attention mechanism code or debug dimension mismatch errors in Transformers. Master the tensor shapes at each step of attention — this is the key to passing the assessment.

Transformer Architecture — Encoder-Decoder, Self-Attention, Cross-Attention
Transformer Architecture — Encoder-Decoder, Self-Attention, Cross-Attention

2. Attention Mechanism

2.1 Intuition — "What's most important?"

Attention answers the question: "When processing the current token, which tokens in the input sequence are most important?" Instead of compressing the entire sequence into a fixed-size vector like RNN, Attention allows the model to "look" directly at every position in the input.

The mechanism works through 3 components:

  • Query (Q) — "What am I looking for?" — the current token asks a question
  • Key (K) — "What information do I contain?" — each token advertises its content
  • Value (V) — "Here is the actual information" — the content returned when matched
Example: "The cat sat on the mat because it was tired"
                                              ↑
                                        Token "it" (Query)
                                              │
                ┌─────────────────────────────┤
                │     Attention scores:       │
                │   "cat"  = 0.72  ← high!   │
                │   "mat"  = 0.11            │
                │   "sat"  = 0.08            │
                │   "The"  = 0.03            │
                │   ...                       │
                └─────────────────────────────┘
                → "it" attends mainly to "cat"

2.2 Scaled Dot-Product Attention

The core attention formula:

Attention(Q, K, V) = softmax(Q · K^T / √d_k) · V

Where:
  Q: Query matrix  — shape (seq_len, d_k)
  K: Key matrix    — shape (seq_len, d_k)
  V: Value matrix  — shape (seq_len, d_v)
  d_k: dimension of Key vectors
  √d_k: scaling factor to prevent gradient vanishing in softmax

Why do we need scaling by √d_k? When d_k is large, the dot product Q·K^T can be very large, causing softmax to saturate → gradient approaches 0. Dividing by √d_k keeps the variance stable.

Scaled Dot-Product Attention Flow:

  Q ──┐
      │──→ MatMul ──→ Scale (÷√d_k) ──→ Mask (opt.) ──→ Softmax ──→ MatMul ──→ Output
  K ──┘                                                                ↑
                                                                       │
  V ────────────────────────────────────────────────────────────────────┘

Shapes (batch_size=B, seq_len=S, d_k=D):
  Q:        (B, S, D)
  K^T:      (B, D, S)
  Q·K^T:    (B, S, S)   ← attention score matrix
  softmax:  (B, S, S)   ← attention weights (each row sums to 1)
  × V:      (B, S, D)   ← weighted output

2.3 Code: Implement Attention from Scratch

import torch
import torch.nn.functional as F
import math

def scaled_dot_product_attention(Q, K, V, mask=None):
    """
    Q: (batch, seq_len, d_k)
    K: (batch, seq_len, d_k)
    V: (batch, seq_len, d_v)
    mask: (batch, 1, seq_len) or (batch, seq_len, seq_len)
    """
    d_k = Q.size(-1)

    # Step 1: Compute attention scores
    scores = torch.matmul(Q, K.transpose(-2, -1)) / math.sqrt(d_k)
    # scores shape: (batch, seq_len, seq_len)

    # Step 2: Apply mask (for causal attention in decoder)
    if mask is not None:
        scores = scores.masked_fill(mask == 0, float('-inf'))

    # Step 3: Softmax to get attention weights
    attn_weights = F.softmax(scores, dim=-1)
    # attn_weights shape: (batch, seq_len, seq_len)

    # Step 4: Weighted sum of values
    output = torch.matmul(attn_weights, V)
    # output shape: (batch, seq_len, d_v)

    return output, attn_weights

# Demo
batch_size, seq_len, d_k = 2, 5, 64
Q = torch.randn(batch_size, seq_len, d_k)
K = torch.randn(batch_size, seq_len, d_k)
V = torch.randn(batch_size, seq_len, d_k)

output, weights = scaled_dot_product_attention(Q, K, V)
print(f"Output shape: {output.shape}")    # (2, 5, 64)
print(f"Weights shape: {weights.shape}")  # (2, 5, 5)
print(f"Weights sum per row: {weights.sum(dim=-1)}")  # each row = 1.0

Exam tip: The most common error when implementing attention: forgetting K.transpose(-2, -1) or dividing by the wrong dimension. Always check the shape after each step — scores must have shape (batch, seq_len, seq_len).

3. Multi-Head Attention

3.1 Why multiple heads?

A single head can only learn one type of relationship. Multi-Head Attention allows the model to attend simultaneously to different representation subspaces:

  • Head 1: learns syntactic relationships (subject-verb)
  • Head 2: learns coreference (pronoun → noun)
  • Head 3: learns positional proximity
  • Head 4: learns semantic similarity
Multi-Head Attention Flow:

Input (batch, seq_len, d_model)
    │
    ├── Linear → Q ──┐
    ├── Linear → K ──┼── Split into h heads
    └── Linear → V ──┘
                      │
        ┌─────────────┼─────────────┐
        ▼             ▼             ▼
    Head 1        Head 2    ...  Head h
 Attention()   Attention()    Attention()
        │             │             │
        └─────────────┼─────────────┘
                      │
                  Concatenate
                      │
                   Linear
                      │
                   Output (batch, seq_len, d_model)

Shapes (d_model=512, h=8, d_k = d_model/h = 64):
  Input:          (B, S, 512)
  Per-head Q/K/V: (B, h, S, 64)   ← reshape after linear
  Per-head out:   (B, h, S, 64)
  Concat:         (B, S, 512)      ← h × d_k = d_model
  Final output:   (B, S, 512)

3.2 Code: MultiHeadAttention Module

import torch
import torch.nn as nn
import math

class MultiHeadAttention(nn.Module):
    def __init__(self, d_model, num_heads):
        super().__init__()
        assert d_model % num_heads == 0, "d_model must be divisible by num_heads"

        self.d_model = d_model
        self.num_heads = num_heads
        self.d_k = d_model // num_heads

        # Linear projections for Q, K, V and output
        self.W_q = nn.Linear(d_model, d_model)
        self.W_k = nn.Linear(d_model, d_model)
        self.W_v = nn.Linear(d_model, d_model)
        self.W_o = nn.Linear(d_model, d_model)

    def forward(self, query, key, value, mask=None):
        batch_size = query.size(0)

        # Step 1: Linear projections
        Q = self.W_q(query)  # (B, S, d_model)
        K = self.W_k(key)
        V = self.W_v(value)

        # Step 2: Reshape to multi-head format
        # (B, S, d_model) → (B, S, h, d_k) → (B, h, S, d_k)
        Q = Q.view(batch_size, -1, self.num_heads, self.d_k).transpose(1, 2)
        K = K.view(batch_size, -1, self.num_heads, self.d_k).transpose(1, 2)
        V = V.view(batch_size, -1, self.num_heads, self.d_k).transpose(1, 2)

        # Step 3: Scaled dot-product attention for each head
        scores = torch.matmul(Q, K.transpose(-2, -1)) / math.sqrt(self.d_k)

        if mask is not None:
            scores = scores.masked_fill(mask == 0, float('-inf'))

        attn_weights = torch.softmax(scores, dim=-1)
        context = torch.matmul(attn_weights, V)
        # context: (B, h, S, d_k)

        # Step 4: Concatenate heads
        # (B, h, S, d_k) → (B, S, h, d_k) → (B, S, d_model)
        context = context.transpose(1, 2).contiguous().view(
            batch_size, -1, self.d_model
        )

        # Step 5: Final linear projection
        output = self.W_o(context)
        return output

# Demo
d_model, num_heads = 512, 8
mha = MultiHeadAttention(d_model, num_heads)

x = torch.randn(2, 10, d_model)  # batch=2, seq_len=10
output = mha(x, x, x)  # self-attention: Q=K=V=x
print(f"Output shape: {output.shape}")  # (2, 10, 512)
ParameterCommon ValuesNotes
d_model512, 768, 1024Embedding size
num_heads8, 12, 16d_model must be divisible by num_heads
d_k = d_model / h64, 64, 64Each head typically has d_k = 64
Total params (MHA)4 × d_model²4 linear layers: W_q, W_k, W_v, W_o

Exam tip: In the DLI assessment, if you encounter RuntimeError: shape mismatch in attention, check: (1) d_model % num_heads == 0, (2) view() and transpose() are in the correct order, (3) remember to call .contiguous() before .view() after transpose.

4. Transformer Architecture

4.1 Encoder Block

Each layer in the Transformer Encoder consists of 2 sub-layers with residual connections and layer normalization:

Transformer Encoder Block:

  Input
    │
    ▼
┌─────────────────────────────┐
│  Multi-Head Self-Attention  │
└──────────────┬──────────────┘
               │
    ┌──────────┴──────────┐
    │     Add & Norm      │ ← x + Sublayer(x), then LayerNorm
    └──────────┬──────────┘
               │
    ┌──────────┴──────────┐
    │  Feed-Forward Net   │ ← 2 linear layers + ReLU/GELU
    │  FFN(x) = W₂·σ(W₁x + b₁) + b₂
    └──────────┬──────────┘
               │
    ┌──────────┴──────────┐
    │     Add & Norm      │
    └──────────┬──────────┘
               │
             Output

4.2 Decoder Block

The Decoder has an additional sub-layer for Cross-Attention — attending to the encoder output:

Transformer Decoder Block:

  Input (shifted right)
    │
    ▼
┌────────────────────────────────┐
│  Masked Multi-Head Attention   │ ← causal mask: only sees previous tokens
└──────────────┬─────────────────┘
               │
    ┌──────────┴──────────┐
    │     Add & Norm      │
    └──────────┬──────────┘
               │
    ┌──────────┴──────────────────────────┐
    │  Cross Multi-Head Attention         │ ← Q from decoder, K/V from encoder
    │  Q = decoder hidden states          │
    │  K, V = encoder output              │
    └──────────────────┬──────────────────┘
               │
    ┌──────────┴──────────┐
    │     Add & Norm      │
    └──────────┬──────────┘
               │
    ┌──────────┴──────────┐
    │  Feed-Forward Net   │
    └──────────┬──────────┘
               │
    ┌──────────┴──────────┐
    │     Add & Norm      │
    └──────────┬──────────┘
               │
             Output

4.3 Full Transformer Architecture

┌─────────────────────────────────────────────────────────────────┐
│                    TRANSFORMER ARCHITECTURE                     │
├───────────────────────┬─────────────────────────────────────────┤
│       ENCODER         │              DECODER                    │
│                       │                                         │
│   Input Embedding     │         Output Embedding                │
│        +              │              +                          │
│   Positional Enc.     │         Positional Enc.                 │
│        │              │              │                          │
│   ┌────┴────┐         │         ┌────┴──────────┐              │
│   │  MH     │         │         │  Masked MH    │              │
│   │  Self-  │         │         │  Self-        │              │
│   │  Attn   │         │         │  Attention    │              │
│   └────┬────┘         │         └────┬──────────┘              │
│   Add & Norm          │         Add & Norm                     │
│        │              │              │                          │
│   ┌────┴────┐         │    ┌────────┴────────────┐             │
│   │  Feed   │         │    │  Cross MH Attention │             │
│   │ Forward │    ─────┼───►│  Q=dec, K/V=enc    │             │
│   │   Net   │         │    └─────────┬───────────┘             │
│   └────┬────┘         │         Add & Norm                     │
│   Add & Norm          │              │                          │
│        │              │         ┌────┴────┐                    │
│      × N layers       │         │  Feed   │                    │
│        │              │         │ Forward │                    │
│     Encoder           │         └────┬────┘                    │
│     Output            │         Add & Norm                     │
│                       │              │                          │
│                       │           × N layers                   │
│                       │              │                          │
│                       │         Linear + Softmax               │
│                       │              │                          │
│                       │       Output Probabilities              │
└───────────────────────┴─────────────────────────────────────────┘

4.4 Positional Encoding

Transformers have no concept of "order" like RNNs. Positional Encoding adds position information to embeddings using sin/cos functions:

PE(pos, 2i)   = sin(pos / 10000^(2i/d_model))
PE(pos, 2i+1) = cos(pos / 10000^(2i/d_model))

pos: token position in sequence (0, 1, 2, ...)
i:   dimension index (0, 1, 2, ..., d_model/2)

This is the exact same concept reused in Diffusion Models — sinusoidal embeddings to encode timestep t. Mastering positional encoding here will help you greatly in the Diffusion section.

import torch
import math

class PositionalEncoding(nn.Module):
    def __init__(self, d_model, max_len=5000):
        super().__init__()
        pe = torch.zeros(max_len, d_model)
        position = torch.arange(0, max_len).unsqueeze(1).float()
        div_term = torch.exp(
            torch.arange(0, d_model, 2).float()
            * (-math.log(10000.0) / d_model)
        )
        pe[:, 0::2] = torch.sin(position * div_term)  # even indices
        pe[:, 1::2] = torch.cos(position * div_term)  # odd indices
        pe = pe.unsqueeze(0)  # (1, max_len, d_model)
        self.register_buffer('pe', pe)

    def forward(self, x):
        # x: (batch, seq_len, d_model)
        return x + self.pe[:, :x.size(1), :]

# Demo
pe = PositionalEncoding(d_model=512)
x = torch.randn(2, 100, 512)  # batch=2, seq_len=100
output = pe(x)
print(f"Output shape: {output.shape}")  # (2, 100, 512)

4.5 Layer Normalization vs Batch Normalization

FeatureBatch NormalizationLayer Normalization
Normalizes acrossBatch dimensionFeature dimension
Depends on batch sizeYes — requires large enough batchNo — operates on each sample individually
Used inCNN (computer vision)Transformer, RNN (NLP)
Inference behaviorUses running statisticsComputed directly — same as training
Used in TransformerNoYes — every sub-layer

Exam tip: If a question asks "why does Transformer use LayerNorm instead of BatchNorm?" → answer: (1) NLP has variable sequence lengths → batch stats are unstable, (2) LayerNorm is independent of batch size, works well with small batch sizes and inference.

5. Model Families

5.1 Encoder-only: BERT

BERT (Bidirectional Encoder Representations from Transformers) uses only the Encoder. Pre-trained using:

  • MLM (Masked Language Modeling) — mask 15% of tokens, predict the masked tokens
  • NSP (Next Sentence Prediction) — are two sentences consecutive?

Since BERT can see in both directions (bidirectional), it excels at understanding tasks: classification, NER, QA.

5.2 Decoder-only: GPT

GPT (Generative Pre-trained Transformer) uses only the Decoder with causal masking — each token can only attend to previous tokens. Pre-trained using next-token prediction.

Causal Attention Mask (GPT):

Token:   [The]  [cat]  [sat]  [on]
The       ✓      ✗      ✗      ✗
cat       ✓      ✓      ✗      ✗
sat       ✓      ✓      ✓      ✗
on        ✓      ✓      ✓      ✓

✓ = can attend    ✗ = masked (= -inf before softmax)

→ Each token only "sees" tokens before it
→ Suitable for generation: predict the next token

5.3 Encoder-Decoder: T5

T5 (Text-to-Text Transfer Transformer) uses the full Encoder-Decoder architecture. Every task is framed as "text-in → text-out":

  • Translation: "translate English to French: The cat sat" → "Le chat s'est assis"
  • Summarization: "summarize: {long text}" → "{summary}"
  • Classification: "classify: {text}" → "positive"

5.4 Comparison Table

FeatureBERT (Encoder)GPT (Decoder)T5 (Enc-Dec)
ArchitectureEncoder-onlyDecoder-onlyEncoder-Decoder
DirectionalityBidirectionalLeft-to-right (causal)Bidirectional enc + causal dec
Pre-trainingMLM + NSPNext-token predictionSpan corruption
Attention maskFull (sees all tokens)Causal (only sees past)Full enc + causal dec
StrengthsUnderstanding: NER, QA, classificationGeneration: text, codeSeq2seq: translation, summary
OutputContextual embeddingsNext token probabilityTarget sequence
Example modelsBERT, RoBERTa, DeBERTaGPT-2/3/4, LLaMAT5, BART, mBART
Decision Tree — Choosing a Model Family:

                ┌─ Need to generate long text?
                │   YES → Decoder-only (GPT, LLaMA)
                │
Task ───────────┤
                │   ┌─ Input→Output sequences?
                │   │   YES → Encoder-Decoder (T5, BART)
                NO ─┤
                    │   NO → Encoder-only (BERT)
                    │   (classification, NER, embedding)
                    └──────────────────────────────────

Exam tip: Common question: "Given task X, which model family should you use?" Quick rule: (1) Understanding / classification → BERT, (2) Generation → GPT, (3) Seq2seq (translation, summarization) → T5. Note: a sufficiently large GPT can also handle any task via prompting.

6. Tokenization

6.1 Token vs Word

Language models don't work with "words" but with tokens — units that are smaller than or equal to a word. Tokenization determines how text is split into tokens.

Example tokenization for "unbelievable":

Word-level:     ["unbelievable"]         → vocab too large, many OOV
Character:      ["u","n","b","e",...]    → sequence too long
Subword (BPE):  ["un", "believ", "able"] → balance vocab size and seq length

6.2 BPE, WordPiece, SentencePiece

AlgorithmUsed byHow It WorksCharacteristics
BPE (Byte-Pair Encoding)GPT-2/3/4, RoBERTaRepeatedly merge the most frequent byte pairsBottom-up, greedy merging
WordPieceBERT, DistilBERTMerge pairs that maximize likelihoodUses ## prefix for subwords
SentencePieceT5, LLaMA, mT5Unigram LM or BPE on raw textLanguage-agnostic, no pre-tokenization needed
Comparison example:

Input: "I love tokenization"

BPE (GPT-2):       ["I", " love", " token", "ization"]
WordPiece (BERT):   ["I", "love", "token", "##ization"]
SentencePiece (T5): ["▁I", "▁love", "▁token", "ization"]

WordPiece uses ## for continuation
SentencePiece uses ▁ for word start

6.3 Code: Tokenization with HuggingFace

from transformers import AutoTokenizer

# Load tokenizers for each model family
bert_tok = AutoTokenizer.from_pretrained("bert-base-uncased")
gpt2_tok = AutoTokenizer.from_pretrained("gpt2")
t5_tok = AutoTokenizer.from_pretrained("t5-small")

text = "Transformers are amazing for NLP tasks!"

# BERT (WordPiece)
bert_tokens = bert_tok.tokenize(text)
print(f"BERT:  {bert_tokens}")
# ['transformers', 'are', 'amazing', 'for', 'nl', '##p', 'tasks', '!']

# GPT-2 (BPE)
gpt2_tokens = gpt2_tok.tokenize(text)
print(f"GPT-2: {gpt2_tokens}")
# ['Trans', 'formers', 'Ġare', 'Ġamazing', 'Ġfor', 'ĠNLP', 'Ġtasks', '!']

# T5 (SentencePiece)
t5_tokens = t5_tok.tokenize(text)
print(f"T5:    {t5_tokens}")
# ['▁Transform', 'ers', '▁are', '▁amazing', '▁for', '▁NLP', '▁tasks', '!']

# Encode → token IDs
ids = bert_tok.encode(text, return_tensors="pt")
print(f"Token IDs shape: {ids.shape}")

# Decode back to text
decoded = bert_tok.decode(ids[0])
print(f"Decoded: {decoded}")

Vocab size affects embedding size and model capacity:

ModelTokenizerVocab SizeNotes
BERT-baseWordPiece30,522Lowercase English
GPT-2BPE50,257Case-sensitive
T5SentencePiece32,100Multilingual capable
LLaMA-2SentencePiece32,000BPE variant
GPT-4BPE (cl100k)100,256Optimized for code + multilingual

7. NLP Tasks Mapping

7.1 Task → Model → Output

NLP TaskTask TypeBest Model FamilyOutput
Text ClassificationSequence classificationEncoder (BERT)Single label
Sentiment AnalysisSequence classificationEncoder (BERT)Positive/Negative
Named Entity RecognitionToken classificationEncoder (BERT)Label per token
Question AnsweringExtractive / GenerativeEncoder or Enc-DecSpan or text
SummarizationSeq2seq generationEnc-Dec (T5, BART)Summary text
TranslationSeq2seq generationEnc-Dec (T5, mBART)Translated text
Text GenerationAutoregressiveDecoder (GPT)Continuation
Code GenerationAutoregressiveDecoder (CodeGen)Code

7.2 Code: Fine-tune BERT for Text Classification

import torch
import torch.nn as nn
from transformers import BertModel, BertTokenizer

class BertClassifier(nn.Module):
    def __init__(self, num_classes, model_name='bert-base-uncased'):
        super().__init__()
        self.bert = BertModel.from_pretrained(model_name)
        self.dropout = nn.Dropout(0.1)
        self.classifier = nn.Linear(self.bert.config.hidden_size, num_classes)

    def forward(self, input_ids, attention_mask):
        # BERT output: last_hidden_state, pooler_output
        outputs = self.bert(input_ids=input_ids,
                           attention_mask=attention_mask)

        # Use [CLS] token representation for classification
        cls_output = outputs.pooler_output  # (batch, hidden_size)
        cls_output = self.dropout(cls_output)
        logits = self.classifier(cls_output)  # (batch, num_classes)
        return logits

# Usage
tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')
model = BertClassifier(num_classes=3)

# Tokenize input
text = "This movie was absolutely wonderful!"
encoded = tokenizer(text, return_tensors='pt', padding=True,
                    truncation=True, max_length=128)

# Forward pass
logits = model(encoded['input_ids'], encoded['attention_mask'])
prediction = torch.argmax(logits, dim=-1)
print(f"Predicted class: {prediction.item()}")
# Training loop for BERT classifier
from torch.utils.data import DataLoader

optimizer = torch.optim.AdamW(model.parameters(), lr=2e-5)
loss_fn = nn.CrossEntropyLoss()

model.train()
for epoch in range(3):
    total_loss = 0
    for batch in train_loader:
        optimizer.zero_grad()

        logits = model(batch['input_ids'], batch['attention_mask'])
        loss = loss_fn(logits, batch['labels'])

        loss.backward()
        torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)
        optimizer.step()

        total_loss += loss.item()

    print(f"Epoch {epoch+1}, Loss: {total_loss / len(train_loader):.4f}")

Exam tip: When fine-tuning BERT, 3 important points: (1) Small learning rate (2e-5 to 5e-5) since the model is already pre-trained, (2) Gradient clipping (clip_grad_norm_) to stabilize training, (3) Use pooler_output (CLS token) for classification, last_hidden_state for token-level tasks (NER).

8. Cheat Sheet

ConceptKey Formula / PatternRemember
Scaled Dot-Productsoftmax(QK^T / √d_k) VDivide by √d_k to prevent softmax saturation
Multi-HeadSplit → h × Attention → Concat → Lineard_k = d_model / num_heads
Positional Encodingsin/cos functionsReused in Diffusion timestep embedding
Encoder (BERT)Bidirectional, MLMUnderstanding tasks: NER, QA, classification
Decoder (GPT)Causal mask, next-tokenGeneration tasks: text, code
Enc-Dec (T5)Cross-attention, seq2seqTranslation, summarization
BPEMerge frequent byte pairsGPT family
WordPieceMaximize likelihood mergeBERT family, uses ## prefix
SentencePieceLanguage-agnostic on raw textT5, LLaMA, uses ▁ prefix
Causal MaskLower-triangular matrixGPT: each token only sees preceding ones
Layer NormNormalize across featuresUsed in Transformer (not BatchNorm)
Fine-tune LR2e-5 → 5e-5Small LR because of pre-trained weights

9. Practice Questions

Coding assessment-style questions similar to NVIDIA DLI:

Q1: Implement a scaled_dot_product_attention function. The function takes Q, K, V tensors and an optional mask, and returns the output and attention weights.

Show Answer Q1
import torch
import torch.nn.functional as F
import math

def scaled_dot_product_attention(Q, K, V, mask=None):
    """
    Args:
        Q: (batch, seq_len, d_k)
        K: (batch, seq_len, d_k)
        V: (batch, seq_len, d_v)
        mask: optional (batch, 1, seq_len) or (batch, seq_len, seq_len)
    Returns:
        output: (batch, seq_len, d_v)
        attn_weights: (batch, seq_len, seq_len)
    """
    d_k = Q.size(-1)

    # Compute attention scores
    scores = torch.matmul(Q, K.transpose(-2, -1)) / math.sqrt(d_k)

    # Apply mask if provided
    if mask is not None:
        scores = scores.masked_fill(mask == 0, float('-inf'))

    # Softmax over last dimension (key dimension)
    attn_weights = F.softmax(scores, dim=-1)

    # Weighted sum of values
    output = torch.matmul(attn_weights, V)

    return output, attn_weights

# Verification:
B, S, D = 2, 4, 64
Q = torch.randn(B, S, D)
K = torch.randn(B, S, D)
V = torch.randn(B, S, D)
out, w = scaled_dot_product_attention(Q, K, V)
assert out.shape == (B, S, D)
assert w.shape == (B, S, S)
assert torch.allclose(w.sum(dim=-1), torch.ones(B, S), atol=1e-6)
print("All assertions passed!")

Explanation: The key points are (1) K.transpose(-2, -1) for correct matrix multiplication, (2) divide by math.sqrt(d_k) for scaling, (3) masked_fill with -inf before softmax, (4) softmax on dim=-1. Common error: forgetting to transpose K or applying softmax on the wrong dim.

Q2: What happens if Positional Encoding is removed from a Transformer? Write code to prove it.

Show Answer Q2
import torch

# Self-attention WITHOUT positional encoding
# → output is permutation invariant (token order doesn't matter)

def self_attention_no_pos(x):
    """x: (batch, seq_len, d_model)"""
    d_k = x.size(-1)
    scores = torch.matmul(x, x.transpose(-2, -1)) / (d_k ** 0.5)
    weights = torch.softmax(scores, dim=-1)
    return torch.matmul(weights, x)

# Create input
x = torch.randn(1, 4, 8)  # 4 tokens, d_model=8

# Original output
out1 = self_attention_no_pos(x)

# Shuffle token order: [0,1,2,3] → [2,0,3,1]
perm = [2, 0, 3, 1]
x_shuffled = x[:, perm, :]
out2 = self_attention_no_pos(x_shuffled)

# Check: output is also shuffled in the same order
inv_perm = [1, 3, 0, 2]  # inverse permutation
out2_reordered = out2[:, inv_perm, :]

print(f"Difference: {(out1 - out2_reordered).abs().max().item():.10f}")
# → near 0! Attention doesn't distinguish order
# → "The cat sat on mat" = "mat on sat cat The"
# That's why Positional Encoding is REQUIRED!

Explanation: Without Positional Encoding, self-attention is permutation equivariant — it processes "The cat sat" identically to "sat The cat". Positional Encoding breaks this symmetry, allowing the model to distinguish token order. In Diffusion Models, the same concept is used for timestep embedding.

Q3: GPT uses causal masking. Write code to create a causal mask and explain each element in the mask matrix.

Show Answer Q3
import torch

def create_causal_mask(seq_len):
    """
    Create a causal (look-ahead) mask for the decoder.
    mask[i][j] = 1 if token i can attend to token j (j <= i)
    mask[i][j] = 0 if token i CANNOT see token j (j > i)
    """
    mask = torch.tril(torch.ones(seq_len, seq_len))
    return mask

seq_len = 5
mask = create_causal_mask(seq_len)
print("Causal Mask:")
print(mask)
# tensor([[1., 0., 0., 0., 0.],
#         [1., 1., 0., 0., 0.],
#         [1., 1., 1., 0., 0.],
#         [1., 1., 1., 1., 0.],
#         [1., 1., 1., 1., 1.]])

# Apply to attention
def causal_attention(Q, K, V):
    d_k = Q.size(-1)
    seq_len = Q.size(1)
    scores = torch.matmul(Q, K.transpose(-2, -1)) / (d_k ** 0.5)

    # Apply causal mask
    mask = create_causal_mask(seq_len).unsqueeze(0)  # (1, S, S)
    scores = scores.masked_fill(mask == 0, float('-inf'))
    # Result: future positions → -inf → softmax → 0

    weights = torch.softmax(scores, dim=-1)
    print("Attention weights (causal):")
    print(weights[0].detach())
    # Row 0: [1.0, 0.0, 0.0, 0.0, 0.0]  ← token 0 only sees itself
    # Row 1: [0.4, 0.6, 0.0, 0.0, 0.0]  ← token 1 sees tokens 0, 1
    # Row 4: [0.1, 0.2, 0.3, 0.2, 0.2]  ← token 4 sees all tokens

    return torch.matmul(weights, V)

B, S, D = 1, 5, 32
Q = torch.randn(B, S, D)
K = torch.randn(B, S, D)
V = torch.randn(B, S, D)
out = causal_attention(Q, K, V)
print(f"Output shape: {out.shape}")  # (1, 5, 32)

Explanation: The causal mask is a lower-triangular matrix (torch.tril). Positions where mask[i][j]=0 (j > i) are filled with -inf before softmax, becoming 0 after softmax. This ensures each token only attends to preceding tokens — essential for autoregressive generation in GPT.

Q4: For the following use cases, choose the most appropriate model family and explain why:

  • (a) Email spam/not-spam classification
  • (b) Vietnamese to English translation
  • (c) Free-form text generation chatbot
  • (d) Extracting person names from text (NER)
Show Answer Q4
# Mapping use cases → model families

tasks = {
    "(a) Email spam classification": {
        "model_family": "Encoder-only (BERT)",
        "reason": "Classification task — needs to understand the entire email "
                  "(bidirectional). Output = 1 label (spam/not-spam). "
                  "BERT + Linear classifier head.",
        "code_hint": "BertForSequenceClassification.from_pretrained('bert-base-uncased', num_labels=2)"
    },
    "(b) Vietnamese → English translation": {
        "model_family": "Encoder-Decoder (T5, mBART)",
        "reason": "Seq2seq task — input sequence (Vietnamese) → output "
                  "sequence (English). Encoder understands input, decoder "
                  "generates output. T5 or mBART for multilingual.",
        "code_hint": "T5ForConditionalGeneration.from_pretrained('t5-base')"
    },
    "(c) Free-form chatbot": {
        "model_family": "Decoder-only (GPT, LLaMA)",
        "reason": "Autoregressive generation — generates text token-by-token, "
                  "no separate encoder needed. GPT/LLaMA with instruction "
                  "tuning for chatbot use case.",
        "code_hint": "AutoModelForCausalLM.from_pretrained('meta-llama/Llama-2-7b-chat-hf')"
    },
    "(d) Named Entity Recognition": {
        "model_family": "Encoder-only (BERT)",
        "reason": "Token classification — needs to assign a label to EACH token "
                  "(B-PER, I-PER, O, B-LOC,...). BERT's bidirectional attention "
                  "allows each token to see context on both sides.",
        "code_hint": "BertForTokenClassification.from_pretrained('bert-base-uncased', num_labels=9)"
    }
}

for task, info in tasks.items():
    print(f"\n{task}")
    print(f"  → {info['model_family']}")
    print(f"  Reason: {info['reason']}")
    print(f"  Code: {info['code_hint']}")

Explanation: General rule: (1) If output is 1 label for the entire input → Encoder (BERT), (2) If output is a label for each token → Encoder (BERT) + token classification head, (3) If output is a sequence in a different language/format → Encoder-Decoder (T5), (4) If continuous text generation is needed → Decoder (GPT). However, in practice, sufficiently large decoder-only LLMs (GPT-4, LLaMA-70B) can handle any task well via prompting.

Q5: Debug the following Transformer error. The code has a dimension bug — find and fix it:

# BUG CODE — Find and fix the error
class BrokenMultiHeadAttention(nn.Module):
    def __init__(self, d_model=512, num_heads=8):
        super().__init__()
        self.d_model = d_model
        self.num_heads = num_heads
        self.d_k = d_model // num_heads  # 64

        self.W_q = nn.Linear(d_model, d_model)
        self.W_k = nn.Linear(d_model, d_model)
        self.W_v = nn.Linear(d_model, d_model)
        self.W_o = nn.Linear(d_model, d_model)

    def forward(self, x, mask=None):
        B = x.size(0)
        Q = self.W_q(x)
        K = self.W_k(x)
        V = self.W_v(x)

        # BUG: wrong dimension order in reshape
        Q = Q.view(B, self.num_heads, -1, self.d_k)  # ← Wrong!
        K = K.view(B, self.num_heads, -1, self.d_k)
        V = V.view(B, self.num_heads, -1, self.d_k)

        scores = torch.matmul(Q, K.transpose(-2, -1)) / (self.d_k ** 0.5)
        weights = torch.softmax(scores, dim=-1)
        context = torch.matmul(weights, V)

        # BUG: missing contiguous() before view
        context = context.transpose(1, 2).view(B, -1, self.d_model)  # ← Wrong!
        return self.W_o(context)
Show Answer Q5
import torch
import torch.nn as nn

class FixedMultiHeadAttention(nn.Module):
    def __init__(self, d_model=512, num_heads=8):
        super().__init__()
        self.d_model = d_model
        self.num_heads = num_heads
        self.d_k = d_model // num_heads

        self.W_q = nn.Linear(d_model, d_model)
        self.W_k = nn.Linear(d_model, d_model)
        self.W_v = nn.Linear(d_model, d_model)
        self.W_o = nn.Linear(d_model, d_model)

    def forward(self, x, mask=None):
        B = x.size(0)
        Q = self.W_q(x)
        K = self.W_k(x)
        V = self.W_v(x)

        # FIX 1: Correct order: (B, S, d_model) → (B, S, h, d_k) → (B, h, S, d_k)
        # Must view as (B, S, h, d_k) FIRST, then transpose(1,2)
        Q = Q.view(B, -1, self.num_heads, self.d_k).transpose(1, 2)
        K = K.view(B, -1, self.num_heads, self.d_k).transpose(1, 2)
        V = V.view(B, -1, self.num_heads, self.d_k).transpose(1, 2)

        scores = torch.matmul(Q, K.transpose(-2, -1)) / (self.d_k ** 0.5)
        weights = torch.softmax(scores, dim=-1)
        context = torch.matmul(weights, V)  # (B, h, S, d_k)

        # FIX 2: Add .contiguous() after transpose before .view()
        context = context.transpose(1, 2).contiguous().view(B, -1, self.d_model)
        return self.W_o(context)

# Verify
model = FixedMultiHeadAttention(d_model=512, num_heads=8)
x = torch.randn(2, 10, 512)
out = model(x)
print(f"Output shape: {out.shape}")  # (2, 10, 512) ✓
assert out.shape == (2, 10, 512)
print("Fixed! All correct.")

Explanation: There are 2 bugs: Bug 1: view(B, num_heads, -1, d_k) is wrong because the tensor layout in memory is (B, S, d_model). You must first view as (B, S, num_heads, d_k) then transpose(1, 2) to get (B, num_heads, S, d_k). Directly viewing as (B, num_heads, S, d_k) will mix data across heads. Bug 2: After transpose(1, 2), the tensor is no longer contiguous in memory. Calling .view() on a non-contiguous tensor causes a RuntimeError. You must add .contiguous() before .view().