Introduction
If BERT "reads" text from two directions, GPT "writes" text from left to right — autoregressive generation. GPT is the foundation of ChatGPT, Claude, Gemini — LLMs that are changing the world.
1. GPT Architecture: Decoder-only
Input: "Once upon a time"
│ │ │ │
▼ ▼ ▼ ▼
┌──────────────────────────┐
│ Transformer Decoder │
│ (Causal Self-Attention)│
│ Chỉ nhìn bên trái! │
└──────────────────────────┘
│ │ │ │
▼ ▼ ▼ ▼
"upon" "a" "time" ","
Causal Language Modeling
$$P(x_1, x_2, ..., x_n) = \prod_{i=1}^{n} P(x_i | x_1, ..., x_{i-1})$$
The model predicts the next token based on all previous tokens (not looking into the future).
2. Evolution: GPT-1 → GPT-4
| Model | Year | Parameters | Training Data | Breakthrough |
|---|---|---|---|---|
| GPT-1 | 2018 | 117M | BookCorpus | Generative pre-training works |
| GPT-2 | 2019 | 1.5B | WebText (40GB) | "Too dangerous to release" |
| GPT-3 | 2020 | 175B | 570GB text | In-context learning, few-shot |
| GPT-4 | 2023 | ~1.8T (rumored) | Internet-scale | Multimodal, reasoning |
| GPT-4o | 2024 | Undisclosed | + Images, Audio | Native multimodal |
3. Decoding Strategies
Temperature, Top-k, Top-p
from transformers import GPT2LMHeadModel, GPT2Tokenizer
tokenizer = GPT2Tokenizer.from_pretrained("gpt2")
model = GPT2LMHeadModel.from_pretrained("gpt2")
input_text = "Artificial intelligence will"
input_ids = tokenizer.encode(input_text, return_tensors="pt")
# Greedy (deterministic, boring)
greedy = model.generate(input_ids, max_length=50, do_sample=False)
# Temperature sampling (creativity control)
creative = model.generate(
input_ids, max_length=50,
do_sample=True,
temperature=0.8, # < 1: focused, > 1: creative
top_k=50, # Chỉ xét 50 tokens có probability cao nhất
top_p=0.9, # Nucleus sampling: 90% probability mass
)
print(tokenizer.decode(creative[0]))
| Parameters | Low | Cao |
|---|---|---|
| Temperature | Accurate, repeatable | Creative, random |
| Top-k | Few options, safe | Many, diverse options |
| Top-p | Focus on tokens definitely | Consider more tokens |
4. In-Context Learning (ICL)
GPT-3 discovers: no need for fine-tuning, just put an example in the prompt!
prompt = """
Classify the sentiment:
Text: "This movie is amazing!" → Positive
Text: "Terrible experience" → Negative
Text: "The food was okay" → Neutral
Text: "I absolutely love this product!" →"""
# GPT sẽ trả lời: "Positive"
# Không cần fine-tune! Chỉ cần prompt engineering.
| Paradigm | Example | Fine-tune? |
|---|---|---|
| Zero-shot | There are no examples | No |
| One-shot | 1 example | No |
| Few-shot | 3-10 examples | No |
| Fine-tuning | Thousands of examples | Yes |
5. BERT vs GPT vs T5
| Features | BERT (Encoder) | GPT (Decoder) | T5 (Enc-Dec) |
|---|---|---|---|
| Direction | Bidirectional | Left-to-right | Both |
| Pre-training | MLM + NSP | Causal LM | Denoising |
| Good for | Classification, NER, QA | Generation, chat | Everything (text-to-text) |
| Example | PhoBERT, RoBERTa | GPT-4, LLaMA | T5, mT5, ViT5 |
Summary
| Concept | Details |
|---|---|
| GPT | Decoder-only, causal LM, autoregressive |
| Scaling laws | Bigger model + more data = better performance |
| Decoding | Temperature, top-k, top-p controls output diversity |
| ICL | Few-shot learning without fine-tuning |
| BERT vs GPT | Understanding (BERT) vs Generation (GPT) |
Next article
Lesson 12: Hugging Face Ecosystem — Practice modern NLP with the most used libraries: Transformers, Datasets, Tokenizers.