1. Introduction
Transformer is the foundational architecture behind every modern Generative AI model — from GPT, BERT, Stable Diffusion to LLaMA. In the NVIDIA DLI assessment, you must thoroughly understand how the Attention Mechanism works and be able to implement it in PyTorch.
This lesson covers everything from Scaled Dot-Product Attention to the full Transformer architecture, then maps to specific model families and NLP tasks.
Exam tip: The NVIDIA DLI assessment often requires you to complete attention mechanism code or debug dimension mismatch errors in Transformers. Master the tensor shapes at each step of attention — this is the key to passing the assessment.

2. Attention Mechanism
2.1 Intuition — "What's most important?"
Attention answers the question: "When processing the current token, which tokens in the input sequence are most important?" Instead of compressing the entire sequence into a fixed-size vector like RNN, Attention allows the model to "look" directly at every position in the input.
The mechanism works through 3 components:
- Query (Q) — "What am I looking for?" — the current token asks a question
- Key (K) — "What information do I contain?" — each token advertises its content
- Value (V) — "Here is the actual information" — the content returned when matched
Example: "The cat sat on the mat because it was tired"
↑
Token "it" (Query)
│
┌─────────────────────────────┤
│ Attention scores: │
│ "cat" = 0.72 ← high! │
│ "mat" = 0.11 │
│ "sat" = 0.08 │
│ "The" = 0.03 │
│ ... │
└─────────────────────────────┘
→ "it" attends mainly to "cat"
2.2 Scaled Dot-Product Attention
The core attention formula:
Attention(Q, K, V) = softmax(Q · K^T / √d_k) · V
Where:
Q: Query matrix — shape (seq_len, d_k)
K: Key matrix — shape (seq_len, d_k)
V: Value matrix — shape (seq_len, d_v)
d_k: dimension of Key vectors
√d_k: scaling factor to prevent gradient vanishing in softmax
Why do we need scaling by √d_k? When d_k is large, the dot product Q·K^T can be very large, causing softmax to saturate → gradient approaches 0. Dividing by √d_k keeps the variance stable.
Scaled Dot-Product Attention Flow:
Q ──┐
│──→ MatMul ──→ Scale (÷√d_k) ──→ Mask (opt.) ──→ Softmax ──→ MatMul ──→ Output
K ──┘ ↑
│
V ────────────────────────────────────────────────────────────────────┘
Shapes (batch_size=B, seq_len=S, d_k=D):
Q: (B, S, D)
K^T: (B, D, S)
Q·K^T: (B, S, S) ← attention score matrix
softmax: (B, S, S) ← attention weights (each row sums to 1)
× V: (B, S, D) ← weighted output
2.3 Code: Implement Attention from Scratch
import torch
import torch.nn.functional as F
import math
def scaled_dot_product_attention(Q, K, V, mask=None):
"""
Q: (batch, seq_len, d_k)
K: (batch, seq_len, d_k)
V: (batch, seq_len, d_v)
mask: (batch, 1, seq_len) or (batch, seq_len, seq_len)
"""
d_k = Q.size(-1)
# Step 1: Compute attention scores
scores = torch.matmul(Q, K.transpose(-2, -1)) / math.sqrt(d_k)
# scores shape: (batch, seq_len, seq_len)
# Step 2: Apply mask (for causal attention in decoder)
if mask is not None:
scores = scores.masked_fill(mask == 0, float('-inf'))
# Step 3: Softmax to get attention weights
attn_weights = F.softmax(scores, dim=-1)
# attn_weights shape: (batch, seq_len, seq_len)
# Step 4: Weighted sum of values
output = torch.matmul(attn_weights, V)
# output shape: (batch, seq_len, d_v)
return output, attn_weights
# Demo
batch_size, seq_len, d_k = 2, 5, 64
Q = torch.randn(batch_size, seq_len, d_k)
K = torch.randn(batch_size, seq_len, d_k)
V = torch.randn(batch_size, seq_len, d_k)
output, weights = scaled_dot_product_attention(Q, K, V)
print(f"Output shape: {output.shape}") # (2, 5, 64)
print(f"Weights shape: {weights.shape}") # (2, 5, 5)
print(f"Weights sum per row: {weights.sum(dim=-1)}") # each row = 1.0
Exam tip: The most common error when implementing attention: forgetting
K.transpose(-2, -1)or dividing by the wrong dimension. Always check the shape after each step —scoresmust have shape(batch, seq_len, seq_len).
3. Multi-Head Attention
3.1 Why multiple heads?
A single head can only learn one type of relationship. Multi-Head Attention allows the model to attend simultaneously to different representation subspaces:
- Head 1: learns syntactic relationships (subject-verb)
- Head 2: learns coreference (pronoun → noun)
- Head 3: learns positional proximity
- Head 4: learns semantic similarity
Multi-Head Attention Flow:
Input (batch, seq_len, d_model)
│
├── Linear → Q ──┐
├── Linear → K ──┼── Split into h heads
└── Linear → V ──┘
│
┌─────────────┼─────────────┐
▼ ▼ ▼
Head 1 Head 2 ... Head h
Attention() Attention() Attention()
│ │ │
└─────────────┼─────────────┘
│
Concatenate
│
Linear
│
Output (batch, seq_len, d_model)
Shapes (d_model=512, h=8, d_k = d_model/h = 64):
Input: (B, S, 512)
Per-head Q/K/V: (B, h, S, 64) ← reshape after linear
Per-head out: (B, h, S, 64)
Concat: (B, S, 512) ← h × d_k = d_model
Final output: (B, S, 512)
3.2 Code: MultiHeadAttention Module
import torch
import torch.nn as nn
import math
class MultiHeadAttention(nn.Module):
def __init__(self, d_model, num_heads):
super().__init__()
assert d_model % num_heads == 0, "d_model must be divisible by num_heads"
self.d_model = d_model
self.num_heads = num_heads
self.d_k = d_model // num_heads
# Linear projections for Q, K, V and output
self.W_q = nn.Linear(d_model, d_model)
self.W_k = nn.Linear(d_model, d_model)
self.W_v = nn.Linear(d_model, d_model)
self.W_o = nn.Linear(d_model, d_model)
def forward(self, query, key, value, mask=None):
batch_size = query.size(0)
# Step 1: Linear projections
Q = self.W_q(query) # (B, S, d_model)
K = self.W_k(key)
V = self.W_v(value)
# Step 2: Reshape to multi-head format
# (B, S, d_model) → (B, S, h, d_k) → (B, h, S, d_k)
Q = Q.view(batch_size, -1, self.num_heads, self.d_k).transpose(1, 2)
K = K.view(batch_size, -1, self.num_heads, self.d_k).transpose(1, 2)
V = V.view(batch_size, -1, self.num_heads, self.d_k).transpose(1, 2)
# Step 3: Scaled dot-product attention for each head
scores = torch.matmul(Q, K.transpose(-2, -1)) / math.sqrt(self.d_k)
if mask is not None:
scores = scores.masked_fill(mask == 0, float('-inf'))
attn_weights = torch.softmax(scores, dim=-1)
context = torch.matmul(attn_weights, V)
# context: (B, h, S, d_k)
# Step 4: Concatenate heads
# (B, h, S, d_k) → (B, S, h, d_k) → (B, S, d_model)
context = context.transpose(1, 2).contiguous().view(
batch_size, -1, self.d_model
)
# Step 5: Final linear projection
output = self.W_o(context)
return output
# Demo
d_model, num_heads = 512, 8
mha = MultiHeadAttention(d_model, num_heads)
x = torch.randn(2, 10, d_model) # batch=2, seq_len=10
output = mha(x, x, x) # self-attention: Q=K=V=x
print(f"Output shape: {output.shape}") # (2, 10, 512)
| Parameter | Common Values | Notes |
|---|---|---|
| d_model | 512, 768, 1024 | Embedding size |
| num_heads | 8, 12, 16 | d_model must be divisible by num_heads |
| d_k = d_model / h | 64, 64, 64 | Each head typically has d_k = 64 |
| Total params (MHA) | 4 × d_model² | 4 linear layers: W_q, W_k, W_v, W_o |
Exam tip: In the DLI assessment, if you encounter
RuntimeError: shape mismatchin attention, check: (1)d_model % num_heads == 0, (2)view()andtranspose()are in the correct order, (3) remember to call.contiguous()before.view()after transpose.
4. Transformer Architecture
4.1 Encoder Block
Each layer in the Transformer Encoder consists of 2 sub-layers with residual connections and layer normalization:
Transformer Encoder Block:
Input
│
▼
┌─────────────────────────────┐
│ Multi-Head Self-Attention │
└──────────────┬──────────────┘
│
┌──────────┴──────────┐
│ Add & Norm │ ← x + Sublayer(x), then LayerNorm
└──────────┬──────────┘
│
┌──────────┴──────────┐
│ Feed-Forward Net │ ← 2 linear layers + ReLU/GELU
│ FFN(x) = W₂·σ(W₁x + b₁) + b₂
└──────────┬──────────┘
│
┌──────────┴──────────┐
│ Add & Norm │
└──────────┬──────────┘
│
Output
4.2 Decoder Block
The Decoder has an additional sub-layer for Cross-Attention — attending to the encoder output:
Transformer Decoder Block:
Input (shifted right)
│
▼
┌────────────────────────────────┐
│ Masked Multi-Head Attention │ ← causal mask: only sees previous tokens
└──────────────┬─────────────────┘
│
┌──────────┴──────────┐
│ Add & Norm │
└──────────┬──────────┘
│
┌──────────┴──────────────────────────┐
│ Cross Multi-Head Attention │ ← Q from decoder, K/V from encoder
│ Q = decoder hidden states │
│ K, V = encoder output │
└──────────────────┬──────────────────┘
│
┌──────────┴──────────┐
│ Add & Norm │
└──────────┬──────────┘
│
┌──────────┴──────────┐
│ Feed-Forward Net │
└──────────┬──────────┘
│
┌──────────┴──────────┐
│ Add & Norm │
└──────────┬──────────┘
│
Output
4.3 Full Transformer Architecture
┌─────────────────────────────────────────────────────────────────┐
│ TRANSFORMER ARCHITECTURE │
├───────────────────────┬─────────────────────────────────────────┤
│ ENCODER │ DECODER │
│ │ │
│ Input Embedding │ Output Embedding │
│ + │ + │
│ Positional Enc. │ Positional Enc. │
│ │ │ │ │
│ ┌────┴────┐ │ ┌────┴──────────┐ │
│ │ MH │ │ │ Masked MH │ │
│ │ Self- │ │ │ Self- │ │
│ │ Attn │ │ │ Attention │ │
│ └────┬────┘ │ └────┬──────────┘ │
│ Add & Norm │ Add & Norm │
│ │ │ │ │
│ ┌────┴────┐ │ ┌────────┴────────────┐ │
│ │ Feed │ │ │ Cross MH Attention │ │
│ │ Forward │ ─────┼───►│ Q=dec, K/V=enc │ │
│ │ Net │ │ └─────────┬───────────┘ │
│ └────┬────┘ │ Add & Norm │
│ Add & Norm │ │ │
│ │ │ ┌────┴────┐ │
│ × N layers │ │ Feed │ │
│ │ │ │ Forward │ │
│ Encoder │ └────┬────┘ │
│ Output │ Add & Norm │
│ │ │ │
│ │ × N layers │
│ │ │ │
│ │ Linear + Softmax │
│ │ │ │
│ │ Output Probabilities │
└───────────────────────┴─────────────────────────────────────────┘
4.4 Positional Encoding
Transformers have no concept of "order" like RNNs. Positional Encoding adds position information to embeddings using sin/cos functions:
PE(pos, 2i) = sin(pos / 10000^(2i/d_model))
PE(pos, 2i+1) = cos(pos / 10000^(2i/d_model))
pos: token position in sequence (0, 1, 2, ...)
i: dimension index (0, 1, 2, ..., d_model/2)
This is the exact same concept reused in Diffusion Models — sinusoidal embeddings to encode timestep t. Mastering positional encoding here will help you greatly in the Diffusion section.
import torch
import math
class PositionalEncoding(nn.Module):
def __init__(self, d_model, max_len=5000):
super().__init__()
pe = torch.zeros(max_len, d_model)
position = torch.arange(0, max_len).unsqueeze(1).float()
div_term = torch.exp(
torch.arange(0, d_model, 2).float()
* (-math.log(10000.0) / d_model)
)
pe[:, 0::2] = torch.sin(position * div_term) # even indices
pe[:, 1::2] = torch.cos(position * div_term) # odd indices
pe = pe.unsqueeze(0) # (1, max_len, d_model)
self.register_buffer('pe', pe)
def forward(self, x):
# x: (batch, seq_len, d_model)
return x + self.pe[:, :x.size(1), :]
# Demo
pe = PositionalEncoding(d_model=512)
x = torch.randn(2, 100, 512) # batch=2, seq_len=100
output = pe(x)
print(f"Output shape: {output.shape}") # (2, 100, 512)
4.5 Layer Normalization vs Batch Normalization
| Feature | Batch Normalization | Layer Normalization |
|---|---|---|
| Normalizes across | Batch dimension | Feature dimension |
| Depends on batch size | Yes — requires large enough batch | No — operates on each sample individually |
| Used in | CNN (computer vision) | Transformer, RNN (NLP) |
| Inference behavior | Uses running statistics | Computed directly — same as training |
| Used in Transformer | No | Yes — every sub-layer |
Exam tip: If a question asks "why does Transformer use LayerNorm instead of BatchNorm?" → answer: (1) NLP has variable sequence lengths → batch stats are unstable, (2) LayerNorm is independent of batch size, works well with small batch sizes and inference.
5. Model Families
5.1 Encoder-only: BERT
BERT (Bidirectional Encoder Representations from Transformers) uses only the Encoder. Pre-trained using:
- MLM (Masked Language Modeling) — mask 15% of tokens, predict the masked tokens
- NSP (Next Sentence Prediction) — are two sentences consecutive?
Since BERT can see in both directions (bidirectional), it excels at understanding tasks: classification, NER, QA.
5.2 Decoder-only: GPT
GPT (Generative Pre-trained Transformer) uses only the Decoder with causal masking — each token can only attend to previous tokens. Pre-trained using next-token prediction.
Causal Attention Mask (GPT):
Token: [The] [cat] [sat] [on]
The ✓ ✗ ✗ ✗
cat ✓ ✓ ✗ ✗
sat ✓ ✓ ✓ ✗
on ✓ ✓ ✓ ✓
✓ = can attend ✗ = masked (= -inf before softmax)
→ Each token only "sees" tokens before it
→ Suitable for generation: predict the next token
5.3 Encoder-Decoder: T5
T5 (Text-to-Text Transfer Transformer) uses the full Encoder-Decoder architecture. Every task is framed as "text-in → text-out":
- Translation:
"translate English to French: The cat sat"→"Le chat s'est assis" - Summarization:
"summarize: {long text}"→"{summary}" - Classification:
"classify: {text}"→"positive"
5.4 Comparison Table
| Feature | BERT (Encoder) | GPT (Decoder) | T5 (Enc-Dec) |
|---|---|---|---|
| Architecture | Encoder-only | Decoder-only | Encoder-Decoder |
| Directionality | Bidirectional | Left-to-right (causal) | Bidirectional enc + causal dec |
| Pre-training | MLM + NSP | Next-token prediction | Span corruption |
| Attention mask | Full (sees all tokens) | Causal (only sees past) | Full enc + causal dec |
| Strengths | Understanding: NER, QA, classification | Generation: text, code | Seq2seq: translation, summary |
| Output | Contextual embeddings | Next token probability | Target sequence |
| Example models | BERT, RoBERTa, DeBERTa | GPT-2/3/4, LLaMA | T5, BART, mBART |
Decision Tree — Choosing a Model Family:
┌─ Need to generate long text?
│ YES → Decoder-only (GPT, LLaMA)
│
Task ───────────┤
│ ┌─ Input→Output sequences?
│ │ YES → Encoder-Decoder (T5, BART)
NO ─┤
│ NO → Encoder-only (BERT)
│ (classification, NER, embedding)
└──────────────────────────────────
Exam tip: Common question: "Given task X, which model family should you use?" Quick rule: (1) Understanding / classification → BERT, (2) Generation → GPT, (3) Seq2seq (translation, summarization) → T5. Note: a sufficiently large GPT can also handle any task via prompting.
6. Tokenization
6.1 Token vs Word
Language models don't work with "words" but with tokens — units that are smaller than or equal to a word. Tokenization determines how text is split into tokens.
Example tokenization for "unbelievable":
Word-level: ["unbelievable"] → vocab too large, many OOV
Character: ["u","n","b","e",...] → sequence too long
Subword (BPE): ["un", "believ", "able"] → balance vocab size and seq length
6.2 BPE, WordPiece, SentencePiece
| Algorithm | Used by | How It Works | Characteristics |
|---|---|---|---|
| BPE (Byte-Pair Encoding) | GPT-2/3/4, RoBERTa | Repeatedly merge the most frequent byte pairs | Bottom-up, greedy merging |
| WordPiece | BERT, DistilBERT | Merge pairs that maximize likelihood | Uses ## prefix for subwords |
| SentencePiece | T5, LLaMA, mT5 | Unigram LM or BPE on raw text | Language-agnostic, no pre-tokenization needed |
Comparison example:
Input: "I love tokenization"
BPE (GPT-2): ["I", " love", " token", "ization"]
WordPiece (BERT): ["I", "love", "token", "##ization"]
SentencePiece (T5): ["▁I", "▁love", "▁token", "ization"]
WordPiece uses ## for continuation
SentencePiece uses ▁ for word start
6.3 Code: Tokenization with HuggingFace
from transformers import AutoTokenizer
# Load tokenizers for each model family
bert_tok = AutoTokenizer.from_pretrained("bert-base-uncased")
gpt2_tok = AutoTokenizer.from_pretrained("gpt2")
t5_tok = AutoTokenizer.from_pretrained("t5-small")
text = "Transformers are amazing for NLP tasks!"
# BERT (WordPiece)
bert_tokens = bert_tok.tokenize(text)
print(f"BERT: {bert_tokens}")
# ['transformers', 'are', 'amazing', 'for', 'nl', '##p', 'tasks', '!']
# GPT-2 (BPE)
gpt2_tokens = gpt2_tok.tokenize(text)
print(f"GPT-2: {gpt2_tokens}")
# ['Trans', 'formers', 'Ġare', 'Ġamazing', 'Ġfor', 'ĠNLP', 'Ġtasks', '!']
# T5 (SentencePiece)
t5_tokens = t5_tok.tokenize(text)
print(f"T5: {t5_tokens}")
# ['▁Transform', 'ers', '▁are', '▁amazing', '▁for', '▁NLP', '▁tasks', '!']
# Encode → token IDs
ids = bert_tok.encode(text, return_tensors="pt")
print(f"Token IDs shape: {ids.shape}")
# Decode back to text
decoded = bert_tok.decode(ids[0])
print(f"Decoded: {decoded}")
Vocab size affects embedding size and model capacity:
| Model | Tokenizer | Vocab Size | Notes |
|---|---|---|---|
| BERT-base | WordPiece | 30,522 | Lowercase English |
| GPT-2 | BPE | 50,257 | Case-sensitive |
| T5 | SentencePiece | 32,100 | Multilingual capable |
| LLaMA-2 | SentencePiece | 32,000 | BPE variant |
| GPT-4 | BPE (cl100k) | 100,256 | Optimized for code + multilingual |
7. NLP Tasks Mapping
7.1 Task → Model → Output
| NLP Task | Task Type | Best Model Family | Output |
|---|---|---|---|
| Text Classification | Sequence classification | Encoder (BERT) | Single label |
| Sentiment Analysis | Sequence classification | Encoder (BERT) | Positive/Negative |
| Named Entity Recognition | Token classification | Encoder (BERT) | Label per token |
| Question Answering | Extractive / Generative | Encoder or Enc-Dec | Span or text |
| Summarization | Seq2seq generation | Enc-Dec (T5, BART) | Summary text |
| Translation | Seq2seq generation | Enc-Dec (T5, mBART) | Translated text |
| Text Generation | Autoregressive | Decoder (GPT) | Continuation |
| Code Generation | Autoregressive | Decoder (CodeGen) | Code |
7.2 Code: Fine-tune BERT for Text Classification
import torch
import torch.nn as nn
from transformers import BertModel, BertTokenizer
class BertClassifier(nn.Module):
def __init__(self, num_classes, model_name='bert-base-uncased'):
super().__init__()
self.bert = BertModel.from_pretrained(model_name)
self.dropout = nn.Dropout(0.1)
self.classifier = nn.Linear(self.bert.config.hidden_size, num_classes)
def forward(self, input_ids, attention_mask):
# BERT output: last_hidden_state, pooler_output
outputs = self.bert(input_ids=input_ids,
attention_mask=attention_mask)
# Use [CLS] token representation for classification
cls_output = outputs.pooler_output # (batch, hidden_size)
cls_output = self.dropout(cls_output)
logits = self.classifier(cls_output) # (batch, num_classes)
return logits
# Usage
tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')
model = BertClassifier(num_classes=3)
# Tokenize input
text = "This movie was absolutely wonderful!"
encoded = tokenizer(text, return_tensors='pt', padding=True,
truncation=True, max_length=128)
# Forward pass
logits = model(encoded['input_ids'], encoded['attention_mask'])
prediction = torch.argmax(logits, dim=-1)
print(f"Predicted class: {prediction.item()}")
# Training loop for BERT classifier
from torch.utils.data import DataLoader
optimizer = torch.optim.AdamW(model.parameters(), lr=2e-5)
loss_fn = nn.CrossEntropyLoss()
model.train()
for epoch in range(3):
total_loss = 0
for batch in train_loader:
optimizer.zero_grad()
logits = model(batch['input_ids'], batch['attention_mask'])
loss = loss_fn(logits, batch['labels'])
loss.backward()
torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)
optimizer.step()
total_loss += loss.item()
print(f"Epoch {epoch+1}, Loss: {total_loss / len(train_loader):.4f}")
Exam tip: When fine-tuning BERT, 3 important points: (1) Small learning rate (2e-5 to 5e-5) since the model is already pre-trained, (2) Gradient clipping (
clip_grad_norm_) to stabilize training, (3) Use pooler_output (CLS token) for classification, last_hidden_state for token-level tasks (NER).
8. Cheat Sheet
| Concept | Key Formula / Pattern | Remember |
|---|---|---|
| Scaled Dot-Product | softmax(QK^T / √d_k) V | Divide by √d_k to prevent softmax saturation |
| Multi-Head | Split → h × Attention → Concat → Linear | d_k = d_model / num_heads |
| Positional Encoding | sin/cos functions | Reused in Diffusion timestep embedding |
| Encoder (BERT) | Bidirectional, MLM | Understanding tasks: NER, QA, classification |
| Decoder (GPT) | Causal mask, next-token | Generation tasks: text, code |
| Enc-Dec (T5) | Cross-attention, seq2seq | Translation, summarization |
| BPE | Merge frequent byte pairs | GPT family |
| WordPiece | Maximize likelihood merge | BERT family, uses ## prefix |
| SentencePiece | Language-agnostic on raw text | T5, LLaMA, uses ▁ prefix |
| Causal Mask | Lower-triangular matrix | GPT: each token only sees preceding ones |
| Layer Norm | Normalize across features | Used in Transformer (not BatchNorm) |
| Fine-tune LR | 2e-5 → 5e-5 | Small LR because of pre-trained weights |
9. Practice Questions
Coding assessment-style questions similar to NVIDIA DLI:
Q1: Implement a scaled_dot_product_attention function. The function takes Q, K, V tensors and an optional mask, and returns the output and attention weights.
Show Answer Q1
import torch
import torch.nn.functional as F
import math
def scaled_dot_product_attention(Q, K, V, mask=None):
"""
Args:
Q: (batch, seq_len, d_k)
K: (batch, seq_len, d_k)
V: (batch, seq_len, d_v)
mask: optional (batch, 1, seq_len) or (batch, seq_len, seq_len)
Returns:
output: (batch, seq_len, d_v)
attn_weights: (batch, seq_len, seq_len)
"""
d_k = Q.size(-1)
# Compute attention scores
scores = torch.matmul(Q, K.transpose(-2, -1)) / math.sqrt(d_k)
# Apply mask if provided
if mask is not None:
scores = scores.masked_fill(mask == 0, float('-inf'))
# Softmax over last dimension (key dimension)
attn_weights = F.softmax(scores, dim=-1)
# Weighted sum of values
output = torch.matmul(attn_weights, V)
return output, attn_weights
# Verification:
B, S, D = 2, 4, 64
Q = torch.randn(B, S, D)
K = torch.randn(B, S, D)
V = torch.randn(B, S, D)
out, w = scaled_dot_product_attention(Q, K, V)
assert out.shape == (B, S, D)
assert w.shape == (B, S, S)
assert torch.allclose(w.sum(dim=-1), torch.ones(B, S), atol=1e-6)
print("All assertions passed!")
Explanation: The key points are (1) K.transpose(-2, -1) for correct matrix multiplication, (2) divide by math.sqrt(d_k) for scaling, (3) masked_fill with -inf before softmax, (4) softmax on dim=-1. Common error: forgetting to transpose K or applying softmax on the wrong dim.
Q2: What happens if Positional Encoding is removed from a Transformer? Write code to prove it.
Show Answer Q2
import torch
# Self-attention WITHOUT positional encoding
# → output is permutation invariant (token order doesn't matter)
def self_attention_no_pos(x):
"""x: (batch, seq_len, d_model)"""
d_k = x.size(-1)
scores = torch.matmul(x, x.transpose(-2, -1)) / (d_k ** 0.5)
weights = torch.softmax(scores, dim=-1)
return torch.matmul(weights, x)
# Create input
x = torch.randn(1, 4, 8) # 4 tokens, d_model=8
# Original output
out1 = self_attention_no_pos(x)
# Shuffle token order: [0,1,2,3] → [2,0,3,1]
perm = [2, 0, 3, 1]
x_shuffled = x[:, perm, :]
out2 = self_attention_no_pos(x_shuffled)
# Check: output is also shuffled in the same order
inv_perm = [1, 3, 0, 2] # inverse permutation
out2_reordered = out2[:, inv_perm, :]
print(f"Difference: {(out1 - out2_reordered).abs().max().item():.10f}")
# → near 0! Attention doesn't distinguish order
# → "The cat sat on mat" = "mat on sat cat The"
# That's why Positional Encoding is REQUIRED!
Explanation: Without Positional Encoding, self-attention is permutation equivariant — it processes "The cat sat" identically to "sat The cat". Positional Encoding breaks this symmetry, allowing the model to distinguish token order. In Diffusion Models, the same concept is used for timestep embedding.
Q3: GPT uses causal masking. Write code to create a causal mask and explain each element in the mask matrix.
Show Answer Q3
import torch
def create_causal_mask(seq_len):
"""
Create a causal (look-ahead) mask for the decoder.
mask[i][j] = 1 if token i can attend to token j (j <= i)
mask[i][j] = 0 if token i CANNOT see token j (j > i)
"""
mask = torch.tril(torch.ones(seq_len, seq_len))
return mask
seq_len = 5
mask = create_causal_mask(seq_len)
print("Causal Mask:")
print(mask)
# tensor([[1., 0., 0., 0., 0.],
# [1., 1., 0., 0., 0.],
# [1., 1., 1., 0., 0.],
# [1., 1., 1., 1., 0.],
# [1., 1., 1., 1., 1.]])
# Apply to attention
def causal_attention(Q, K, V):
d_k = Q.size(-1)
seq_len = Q.size(1)
scores = torch.matmul(Q, K.transpose(-2, -1)) / (d_k ** 0.5)
# Apply causal mask
mask = create_causal_mask(seq_len).unsqueeze(0) # (1, S, S)
scores = scores.masked_fill(mask == 0, float('-inf'))
# Result: future positions → -inf → softmax → 0
weights = torch.softmax(scores, dim=-1)
print("Attention weights (causal):")
print(weights[0].detach())
# Row 0: [1.0, 0.0, 0.0, 0.0, 0.0] ← token 0 only sees itself
# Row 1: [0.4, 0.6, 0.0, 0.0, 0.0] ← token 1 sees tokens 0, 1
# Row 4: [0.1, 0.2, 0.3, 0.2, 0.2] ← token 4 sees all tokens
return torch.matmul(weights, V)
B, S, D = 1, 5, 32
Q = torch.randn(B, S, D)
K = torch.randn(B, S, D)
V = torch.randn(B, S, D)
out = causal_attention(Q, K, V)
print(f"Output shape: {out.shape}") # (1, 5, 32)
Explanation: The causal mask is a lower-triangular matrix (torch.tril). Positions where mask[i][j]=0 (j > i) are filled with -inf before softmax, becoming 0 after softmax. This ensures each token only attends to preceding tokens — essential for autoregressive generation in GPT.
Q4: For the following use cases, choose the most appropriate model family and explain why:
- (a) Email spam/not-spam classification
- (b) Vietnamese to English translation
- (c) Free-form text generation chatbot
- (d) Extracting person names from text (NER)
Show Answer Q4
# Mapping use cases → model families
tasks = {
"(a) Email spam classification": {
"model_family": "Encoder-only (BERT)",
"reason": "Classification task — needs to understand the entire email "
"(bidirectional). Output = 1 label (spam/not-spam). "
"BERT + Linear classifier head.",
"code_hint": "BertForSequenceClassification.from_pretrained('bert-base-uncased', num_labels=2)"
},
"(b) Vietnamese → English translation": {
"model_family": "Encoder-Decoder (T5, mBART)",
"reason": "Seq2seq task — input sequence (Vietnamese) → output "
"sequence (English). Encoder understands input, decoder "
"generates output. T5 or mBART for multilingual.",
"code_hint": "T5ForConditionalGeneration.from_pretrained('t5-base')"
},
"(c) Free-form chatbot": {
"model_family": "Decoder-only (GPT, LLaMA)",
"reason": "Autoregressive generation — generates text token-by-token, "
"no separate encoder needed. GPT/LLaMA with instruction "
"tuning for chatbot use case.",
"code_hint": "AutoModelForCausalLM.from_pretrained('meta-llama/Llama-2-7b-chat-hf')"
},
"(d) Named Entity Recognition": {
"model_family": "Encoder-only (BERT)",
"reason": "Token classification — needs to assign a label to EACH token "
"(B-PER, I-PER, O, B-LOC,...). BERT's bidirectional attention "
"allows each token to see context on both sides.",
"code_hint": "BertForTokenClassification.from_pretrained('bert-base-uncased', num_labels=9)"
}
}
for task, info in tasks.items():
print(f"\n{task}")
print(f" → {info['model_family']}")
print(f" Reason: {info['reason']}")
print(f" Code: {info['code_hint']}")
Explanation: General rule: (1) If output is 1 label for the entire input → Encoder (BERT), (2) If output is a label for each token → Encoder (BERT) + token classification head, (3) If output is a sequence in a different language/format → Encoder-Decoder (T5), (4) If continuous text generation is needed → Decoder (GPT). However, in practice, sufficiently large decoder-only LLMs (GPT-4, LLaMA-70B) can handle any task well via prompting.
Q5: Debug the following Transformer error. The code has a dimension bug — find and fix it:
# BUG CODE — Find and fix the error
class BrokenMultiHeadAttention(nn.Module):
def __init__(self, d_model=512, num_heads=8):
super().__init__()
self.d_model = d_model
self.num_heads = num_heads
self.d_k = d_model // num_heads # 64
self.W_q = nn.Linear(d_model, d_model)
self.W_k = nn.Linear(d_model, d_model)
self.W_v = nn.Linear(d_model, d_model)
self.W_o = nn.Linear(d_model, d_model)
def forward(self, x, mask=None):
B = x.size(0)
Q = self.W_q(x)
K = self.W_k(x)
V = self.W_v(x)
# BUG: wrong dimension order in reshape
Q = Q.view(B, self.num_heads, -1, self.d_k) # ← Wrong!
K = K.view(B, self.num_heads, -1, self.d_k)
V = V.view(B, self.num_heads, -1, self.d_k)
scores = torch.matmul(Q, K.transpose(-2, -1)) / (self.d_k ** 0.5)
weights = torch.softmax(scores, dim=-1)
context = torch.matmul(weights, V)
# BUG: missing contiguous() before view
context = context.transpose(1, 2).view(B, -1, self.d_model) # ← Wrong!
return self.W_o(context)
Show Answer Q5
import torch
import torch.nn as nn
class FixedMultiHeadAttention(nn.Module):
def __init__(self, d_model=512, num_heads=8):
super().__init__()
self.d_model = d_model
self.num_heads = num_heads
self.d_k = d_model // num_heads
self.W_q = nn.Linear(d_model, d_model)
self.W_k = nn.Linear(d_model, d_model)
self.W_v = nn.Linear(d_model, d_model)
self.W_o = nn.Linear(d_model, d_model)
def forward(self, x, mask=None):
B = x.size(0)
Q = self.W_q(x)
K = self.W_k(x)
V = self.W_v(x)
# FIX 1: Correct order: (B, S, d_model) → (B, S, h, d_k) → (B, h, S, d_k)
# Must view as (B, S, h, d_k) FIRST, then transpose(1,2)
Q = Q.view(B, -1, self.num_heads, self.d_k).transpose(1, 2)
K = K.view(B, -1, self.num_heads, self.d_k).transpose(1, 2)
V = V.view(B, -1, self.num_heads, self.d_k).transpose(1, 2)
scores = torch.matmul(Q, K.transpose(-2, -1)) / (self.d_k ** 0.5)
weights = torch.softmax(scores, dim=-1)
context = torch.matmul(weights, V) # (B, h, S, d_k)
# FIX 2: Add .contiguous() after transpose before .view()
context = context.transpose(1, 2).contiguous().view(B, -1, self.d_model)
return self.W_o(context)
# Verify
model = FixedMultiHeadAttention(d_model=512, num_heads=8)
x = torch.randn(2, 10, 512)
out = model(x)
print(f"Output shape: {out.shape}") # (2, 10, 512) ✓
assert out.shape == (2, 10, 512)
print("Fixed! All correct.")
Explanation: There are 2 bugs: Bug 1: view(B, num_heads, -1, d_k) is wrong because the tensor layout in memory is (B, S, d_model). You must first view as (B, S, num_heads, d_k) then transpose(1, 2) to get (B, num_heads, S, d_k). Directly viewing as (B, num_heads, S, d_k) will mix data across heads. Bug 2: After transpose(1, 2), the tensor is no longer contiguous in memory. Calling .view() on a non-contiguous tensor causes a RuntimeError. You must add .contiguous() before .view().