Introduction
"The bottleneck of Seq2Seq: compressing an entire sentence into a single fixed-size vector."
Attention mechanism allows the model to look back at all input tokens when generating output — without needing to compress all the information into a single vector. This is the direct foundation of the Transformer.
1. The problem of Seq2Seq without Attention
Encoder: "I love natural language processing"
│
▼
[context vector] ← Toàn bộ câu nén vào 1 vector!
│
▼
Decoder: "Tôi yêu xử lý ngôn ngữ tự nhiên"
The longer the sentence → the greater the information loss.
2. Attention — Core idea
Instead of just using the final context vector, the decoder looks back at all the encoder's hidden states:
$$\text{Attention}(Q, K, V) = \text{softmax}\left(\frac{QK^T}{\sqrt{d_k}}\right)V$$
In which:
- Q (Query): "What am I looking for?" — hidden current state of the decoder
- K (Key): "What is in each input?" — hidden states of the encoder
- V (Value): "Real information" — same as K in basic attention
Scaled Dot-Product Attention
import torch
import torch.nn.functional as F
import math
def scaled_dot_product_attention(Q, K, V, mask=None):
"""
Q: (batch, seq_q, d_k)
K: (batch, seq_k, d_k)
V: (batch, seq_k, d_v)
"""
d_k = Q.size(-1)
# 1. Tính attention scores
scores = torch.matmul(Q, K.transpose(-2, -1)) / math.sqrt(d_k)
# scores: (batch, seq_q, seq_k)
# 2. Mask (optional)
if mask is not None:
scores = scores.masked_fill(mask == 0, -1e9)
# 3. Softmax → attention weights
weights = F.softmax(scores, dim=-1)
# 4. Weighted sum of Values
output = torch.matmul(weights, V)
return output, weights
3. Multi-Head Attention
Instead of 1 attention, use multiple "heads" — each head learns a different type of relationship:
class MultiHeadAttention(nn.Module):
def __init__(self, d_model, num_heads):
super().__init__()
assert d_model % num_heads == 0
self.d_k = d_model // num_heads
self.num_heads = num_heads
self.W_q = nn.Linear(d_model, d_model)
self.W_k = nn.Linear(d_model, d_model)
self.W_v = nn.Linear(d_model, d_model)
self.W_o = nn.Linear(d_model, d_model)
def forward(self, Q, K, V, mask=None):
batch_size = Q.size(0)
# Linear projections rồi split thành heads
Q = self.W_q(Q).view(batch_size, -1, self.num_heads, self.d_k).transpose(1, 2)
K = self.W_k(K).view(batch_size, -1, self.num_heads, self.d_k).transpose(1, 2)
V = self.W_v(V).view(batch_size, -1, self.num_heads, self.d_k).transpose(1, 2)
# Attention cho từng head
out, weights = scaled_dot_product_attention(Q, K, V, mask)
# Concat heads
out = out.transpose(1, 2).contiguous().view(batch_size, -1, self.num_heads * self.d_k)
return self.W_o(out)
4. Self-Attention
When Q, K, V all come from same sequence → Self-attention. Each token "sees" all the other tokens in the sentence.
Input: "The cat sat on the mat"
Self-attention cho từ "sat":
→ "The": 0.05 (ít liên quan)
→ "cat": 0.45 (chủ ngữ, rất liên quan!)
→ "sat": 0.20 (chính nó)
→ "on": 0.15
→ "the": 0.05
→ "mat": 0.10
5. Compare types of Attention
| Type | Q | K, V | Use |
|---|---|---|---|
| Bahdanau | Decoder hidden | Encoder hidden | Seq2Seq translation |
| Salary | Decoder hidden | Encoder hidden | Seq2Seq (simpler) |
| Self-attention | Same sequence | Same sequence | Transformer encoder |
| Cross-attention | Decoder | Encoders | Transformer decoder |
| Causal self-attention | Same + mask future | Similar | GPT (autoregressive) |
Summary
| Concept | Meaning |
|---|---|
| Attention | Allows the model to "look back" at all inputs |
| Q, K, V | Query searches in Keys, gets the corresponding Values |
| Scaling | Divide $\sqrt{d_k}$ to stabilize the gradient |
| Multi-head | Many parallel "perspectives" |
| Self-attention | Each token attends to all other tokens |
Next article
Lesson 9: Transformer — "Attention Is All You Need" — Complete architecture combining self-attention, positional encoding, layer norm — the foundation of every modern LLM.