Chuyển đến nội dung chính

Lesson 11: GPT & Autoregressive Models — Generative Pre-trained Transformer

GPT architecture: causal language modeling. GPT-1 → GPT-2 → GPT-3 → GPT-4 Evolution. Autoregressive generation: temperature, top-k, top-p sampling. Emergent abilities. In-context learning. Compare BERT (encoder) vs GPT (decoder) vs T5 (encoder-decoder).

🧠 AI & ML — Lesson 10 Lesson 11: GPT & Autoregressive Models — Generative Pre-trained Transformer

NLP from Basics to Advanced: Mastering Natural Language Processing

Part 4: Pre-trained Language Models — BERT, GPT & Beyond

xdev.asia

Introduction

If BERT "reads" text from two directions, GPT "writes" text from left to right — autoregressive generation. GPT is the foundation of ChatGPT, Claude, Gemini — LLMs that are changing the world.


1. GPT Architecture: Decoder-only

Input:  "Once upon a time"
         │      │     │    │
         ▼      ▼     ▼    ▼
    ┌──────────────────────────┐
    │   Transformer Decoder    │
    │   (Causal Self-Attention)│
    │    Chỉ nhìn bên trái!   │
    └──────────────────────────┘
         │      │     │    │
         ▼      ▼     ▼    ▼
       "upon"  "a"  "time" ","

Causal Language Modeling

$$P(x_1, x_2, ..., x_n) = \prod_{i=1}^{n} P(x_i | x_1, ..., x_{i-1})$$

The model predicts the next token based on all previous tokens (not looking into the future).


2. Evolution: GPT-1 → GPT-4

ModelYearParametersTraining DataBreakthrough
GPT-12018117MBookCorpusGenerative pre-training works
GPT-220191.5BWebText (40GB)"Too dangerous to release"
GPT-32020175B570GB textIn-context learning, few-shot
GPT-42023~1.8T (rumored)Internet-scaleMultimodal, reasoning
GPT-4o2024Undisclosed+ Images, AudioNative multimodal

3. Decoding Strategies

Temperature, Top-k, Top-p

from transformers import GPT2LMHeadModel, GPT2Tokenizer

tokenizer = GPT2Tokenizer.from_pretrained("gpt2")
model = GPT2LMHeadModel.from_pretrained("gpt2")

input_text = "Artificial intelligence will"
input_ids = tokenizer.encode(input_text, return_tensors="pt")

# Greedy (deterministic, boring)
greedy = model.generate(input_ids, max_length=50, do_sample=False)

# Temperature sampling (creativity control)
creative = model.generate(
    input_ids, max_length=50,
    do_sample=True,
    temperature=0.8,   # < 1: focused, > 1: creative
    top_k=50,          # Chỉ xét 50 tokens có probability cao nhất
    top_p=0.9,         # Nucleus sampling: 90% probability mass
)

print(tokenizer.decode(creative[0]))
ParametersLowCao
TemperatureAccurate, repeatableCreative, random
Top-kFew options, safeMany, diverse options
Top-pFocus on tokens definitelyConsider more tokens

4. In-Context Learning (ICL)

GPT-3 discovers: no need for fine-tuning, just put an example in the prompt!

prompt = """
Classify the sentiment:
Text: "This movie is amazing!" → Positive
Text: "Terrible experience" → Negative
Text: "The food was okay" → Neutral
Text: "I absolutely love this product!" →"""

# GPT sẽ trả lời: "Positive"
# Không cần fine-tune! Chỉ cần prompt engineering.
ParadigmExampleFine-tune?
Zero-shotThere are no examplesNo
One-shot1 exampleNo
Few-shot3-10 examplesNo
Fine-tuningThousands of examplesYes

5. BERT vs GPT vs T5

FeaturesBERT (Encoder)GPT (Decoder)T5 (Enc-Dec)
DirectionBidirectionalLeft-to-rightBoth
Pre-trainingMLM + NSPCausal LMDenoising
Good forClassification, NER, QAGeneration, chatEverything (text-to-text)
ExamplePhoBERT, RoBERTaGPT-4, LLaMAT5, mT5, ViT5

Summary

ConceptDetails
GPTDecoder-only, causal LM, autoregressive
Scaling lawsBigger model + more data = better performance
DecodingTemperature, top-k, top-p controls output diversity
ICLFew-shot learning without fine-tuning
BERT vs GPTUnderstanding (BERT) vs Generation (GPT)

Next article

Lesson 12: Hugging Face Ecosystem — Practice modern NLP with the most used libraries: Transformers, Datasets, Tokenizers.