Chuyển đến nội dung chính

Lesson 8: Attention Mechanism — The turning point of NLP

Intuition: why do we need attention? Bahdanau attention vs Luong attention. Self-attention. Scaled dot-product attention. Multi-head attention. Visualize attention weights. From Seq2Seq with attention comes the Transformer platform.

🧠 AI & ML — Lesson 7 Lesson 8: Attention Mechanism — The turning point of NLP

NLP from Basics to Advanced: Mastering Natural Language Processing

Part 3: Deep Learning for NLP — RNN, LSTM, to Transformer

xdev.asia

Introduction

"The bottleneck of Seq2Seq: compressing an entire sentence into a single fixed-size vector."

Attention mechanism allows the model to look back at all input tokens when generating output — without needing to compress all the information into a single vector. This is the direct foundation of the Transformer.


1. The problem of Seq2Seq without Attention

Encoder:  "I love natural language processing"
              │
              ▼
         [context vector]  ← Toàn bộ câu nén vào 1 vector!
              │
              ▼
Decoder:  "Tôi yêu xử lý ngôn ngữ tự nhiên"

The longer the sentence → the greater the information loss.


2. Attention — Core idea

Instead of just using the final context vector, the decoder looks back at all the encoder's hidden states:

$$\text{Attention}(Q, K, V) = \text{softmax}\left(\frac{QK^T}{\sqrt{d_k}}\right)V$$

In which:

  • Q (Query): "What am I looking for?" — hidden current state of the decoder
  • K (Key): "What is in each input?" — hidden states of the encoder
  • V (Value): "Real information" — same as K in basic attention

Scaled Dot-Product Attention

import torch
import torch.nn.functional as F
import math

def scaled_dot_product_attention(Q, K, V, mask=None):
    """
    Q: (batch, seq_q, d_k)
    K: (batch, seq_k, d_k)
    V: (batch, seq_k, d_v)
    """
    d_k = Q.size(-1)

    # 1. Tính attention scores
    scores = torch.matmul(Q, K.transpose(-2, -1)) / math.sqrt(d_k)
    # scores: (batch, seq_q, seq_k)

    # 2. Mask (optional)
    if mask is not None:
        scores = scores.masked_fill(mask == 0, -1e9)

    # 3. Softmax → attention weights
    weights = F.softmax(scores, dim=-1)

    # 4. Weighted sum of Values
    output = torch.matmul(weights, V)
    return output, weights

3. Multi-Head Attention

Instead of 1 attention, use multiple "heads" — each head learns a different type of relationship:

class MultiHeadAttention(nn.Module):
    def __init__(self, d_model, num_heads):
        super().__init__()
        assert d_model % num_heads == 0
        self.d_k = d_model // num_heads
        self.num_heads = num_heads

        self.W_q = nn.Linear(d_model, d_model)
        self.W_k = nn.Linear(d_model, d_model)
        self.W_v = nn.Linear(d_model, d_model)
        self.W_o = nn.Linear(d_model, d_model)

    def forward(self, Q, K, V, mask=None):
        batch_size = Q.size(0)

        # Linear projections rồi split thành heads
        Q = self.W_q(Q).view(batch_size, -1, self.num_heads, self.d_k).transpose(1, 2)
        K = self.W_k(K).view(batch_size, -1, self.num_heads, self.d_k).transpose(1, 2)
        V = self.W_v(V).view(batch_size, -1, self.num_heads, self.d_k).transpose(1, 2)

        # Attention cho từng head
        out, weights = scaled_dot_product_attention(Q, K, V, mask)

        # Concat heads
        out = out.transpose(1, 2).contiguous().view(batch_size, -1, self.num_heads * self.d_k)
        return self.W_o(out)

4. Self-Attention

When Q, K, V all come from same sequence → Self-attention. Each token "sees" all the other tokens in the sentence.

Input: "The cat sat on the mat"

Self-attention cho từ "sat":
  → "The": 0.05 (ít liên quan)
  → "cat": 0.45 (chủ ngữ, rất liên quan!)
  → "sat": 0.20 (chính nó)
  → "on":  0.15
  → "the": 0.05
  → "mat": 0.10

5. Compare types of Attention

TypeQK, VUse
BahdanauDecoder hiddenEncoder hiddenSeq2Seq translation
SalaryDecoder hiddenEncoder hiddenSeq2Seq (simpler)
Self-attentionSame sequenceSame sequenceTransformer encoder
Cross-attentionDecoderEncodersTransformer decoder
Causal self-attentionSame + mask futureSimilarGPT (autoregressive)

Summary

ConceptMeaning
AttentionAllows the model to "look back" at all inputs
Q, K, VQuery searches in Keys, gets the corresponding Values ​​
ScalingDivide $\sqrt{d_k}$ to stabilize the gradient
Multi-headMany parallel "perspectives"
Self-attentionEach token attends to all other tokens

Next article

Lesson 9: Transformer — "Attention Is All You Need" — Complete architecture combining self-attention, positional encoding, layer norm — the foundation of every modern LLM.