Lesson 13: RLHF and Alignment — DPO, PPO
1. Alignment Problem: Why does the LLM need to be "aligned"?
Basic problem
An LLM is pre-trained to maximize the ability to predict the next token. After SFT, it knows how to follow instructions — but this is **not the same as:
- Answer honestly (no hallucinate)
- Refuse harmful requests (weapon synthesis, harmful content)
- Behave consistently with human values
- Prioritize user safety over blind compliance
This is the Alignment Problem: how to ensure AI acts according to human intentions and values?
Misalignment example
User: "Hãy giúp tôi viết email lừa đảo khách hàng."
Misaligned model: [viết email lừa đảo theo đúng yêu cầu]
Aligned model: "Tôi không thể giúp tạo nội dung lừa đảo..."
User: "Trái đất có bằng phẳng không?"
Misaligned model: "Có, một số người tin rằng..." [không phủ nhận]
Aligned model: "Không, trái đất có hình cầu dẹt. Đây là..."
Helpful, Harmless, Honest (HHH)
Anthropic defines three alignment goals:
- Helpful: Really helpful to users
- Harmless: Does not cause harm to users or society
- Honest: Do not lie or create illusions
2. RLHF Pipeline: SFT → Reward Model → PPO
Reinforcement Learning from Human Feedback (RLHF) is a 3-step pipeline:
Step 1: SFT
Base Model → [fine-tune với instruction data] → SFT Model
Step 2: Reward Model Training
SFT Model → [generate responses] → Human ranks A>B → Train Reward Model
Step 3: RL with PPO
SFT Model → [RL với Reward Model làm signal] → Aligned Model
This is exactly how OpenAI created InstructGPT and ChatGPT.
3. Reward Model: Learn from Human Preferences
Reward Model structure
The Reward Model (RM) is an LLM modified to output a scalar score instead of text:
# Conceptually:
class RewardModel(nn.Module):
def __init__(self, base_model):
self.backbone = base_model # LLM
self.value_head = nn.Linear(hidden_size, 1) # Scalar output
def forward(self, input_ids, attention_mask):
outputs = self.backbone(input_ids, attention_mask)
last_hidden = outputs.last_hidden_state[:, -1, :] # Last token
reward = self.value_head(last_hidden)
return reward.squeeze()
Create Preference Dataset
{
"prompt": "Giải thích tại sao bầu trời màu xanh.",
"chosen": "Bầu trời màu xanh do hiện tượng tán xạ Rayleigh. Khi ánh sáng mặt trời...",
"rejected": "Tôi không biết tại sao bầu trời màu xanh, có thể do một số nguyên nhân."
}
Training Objective
RM is trained to give score(chosen) > score(rejected):
Loss = -log(σ(r_chosen - r_rejected))
from trl import RewardTrainer, RewardConfig
from transformers import AutoModelForSequenceClassification
reward_model = AutoModelForSequenceClassification.from_pretrained(
"mistralai/Mistral-7B-v0.3",
num_labels=1,
)
trainer = RewardTrainer(
model=reward_model,
args=RewardConfig(
output_dir="./reward-model",
per_device_train_batch_size=4,
num_train_epochs=3,
),
train_dataset=preference_dataset,
tokenizer=tokenizer,
)
trainer.train()
4. PPO in LLM Context
What is PPO?
Proximal Policy Optimization is a RL algorithm that balances:
- Maximize reward (from Reward Model)
- Do not stray too far from the original SFT policy (KL divergence penalty)
Objective = E[r(x, y)] - β * KL[π_θ(y|x) || π_SFT(y|x)]
r(x, y): reward from Reward ModelKL: KL divergence between current policy and SFT modelβ: coefficient to adjust the allowable "drift" level
Problem with PPO
PPO in RLHF is complicated to implement:
- Need to load 4 models at the same time: Actor, Critic, Reference, Reward
- Training is unstable, many hyperparameters need tuning
- Extremely high VRAM
- Vulnerable to reward hacking: the model learns how to "cheat" the reward model
5. InstructGPT: The story of creating ChatGPT
Paper "Training language models to follow instructions with human feedback" (Ouyang et al., 2022):
OpenAI's 3-step process
Step 1 - SFT:
- Collect 13,000 demonstrations from labelers
- Fine-tune GPT-3 → SFT Model
Step 2 - Reward Model:
- From SFT Model, generate 33,000 pairs
- Human labelers rank: response A > response B
- Train Reward Model 6B params
Step 3 - PPO:
- Use RM to guide PPO training
- Result: model is 100× smaller than GPT-3 but is more popular
Remarkable results
InstructGPT 1.3B được ưa thích hơn GPT-3 175B trong 85% trường hợp
→ Alignment quan trọng hơn scale model!
6. Direct Preference Optimization (DPO)
Why was DPO born?
DPO (Rafailov et al., 2023) addresses the complexity of PPO by:
- Removed Separate Reward Model
- Removed RL training loop
- Convert the problem into supervised learning directly
DPO Mathematics
DPO proves that the optimal policy can be calculated directly:
Loss_DPO = -E[log σ(β * log(π_θ(y_w|x)/π_ref(y_w|x))
- β * log(π_θ(y_l|x)/π_ref(y_l|x)))]
Simpler: increases probability for chosen, decreases probability for rejected (compared to reference model):
# Conceptually:
loss = -log_sigmoid(
beta * (log_prob_chosen_model - log_prob_chosen_ref) -
beta * (log_prob_rejected_model - log_prob_rejected_ref)
)
DPO vs PPO
| PPO | DPO | |
|---|---|---|
| Need Reward Model? | Yes (separate) | No |
| Number of models to load | 4 | 2 (policy + reference) |
| Complexity | Very high | Low |
| Stable training | Unstable | Stable |
| Quality of results | Better (theory) | Equivalent |
| Popular today | Less | More popular |
7. Constitutional AI (Anthropic)
Ideas
Instead of using human feedback, use AI feedback based on a "constitution" (set of principles):
Constitution principles:
- "Chọn response ít có hại nhất cho con người"
- "Chọn response không khuyến khích nội dung bất hợp pháp"
- "Chọn response trung thực nhất"
CAI Pipeline
Phase 1: SL-CAI (Supervised Learning)
- Sample harmful prompt
- Generate initial response (can be harmful)
- Self-critique: "What principle does this response violate?"
- Revision: "Rewrite the response more safely"
- Use (prompt, revised_response) for SFT
Phase 2: RL-CAI
- Use AI model (not human) to create preference labels
- Train Reward Model from AI preferences
- PPO like regular RLHF
# Ví dụ self-critique prompt:
critique_prompt = """
Bạn vừa trả lời: "{response}"
Hãy xem xét response này theo nguyên tắc:
"Không được cung cấp thông tin giúp gây hại cho người khác."
Phần nào của response này vi phạm nguyên tắc đó?
Hãy viết lại response an toàn hơn.
"""
8. ORPO, SimPO: New Methods
ORPO (Odds Ratio Preference Optimization)
ORPO combines SFT loss and preference loss into a single training step:
# Không cần reference model!
loss_ORPO = loss_SFT + lambda * loss_OR
# loss_OR = -log(sigmoid(log(odds_ratio(chosen/rejected))))
Advantages:
- No need for reference model (save VRAM by 50%)
- SFT and alignment occur simultaneously
- Competitive results with DPO
SimPO (Simple Preference Optimization)
SimPO uses average log probability instead of comparing with reference:
# Reward được normalize theo length:
r(x, y) = (1/|y|) * sum(log π_θ(y_t | x, y_<t))
# Loss:
loss = -log_sigmoid(beta * r(chosen) - beta * r(rejected) - gamma)
# gamma: target reward margin
9. Code: DPO Training with TRL DPOTrainer
import torch
from datasets import load_dataset
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import LoraConfig, get_peft_model
from trl import DPOTrainer, DPOConfig
# ============================================================
# 1. Load SFT Model (điểm xuất phát)
# ============================================================
model_id = "./mistral-sft-output" # Model đã qua SFT
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained(model_id)
# Reference model (frozen copy của SFT model)
model_ref = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
)
# LoRA cho DPO (chỉ train một phần)
lora_config = LoraConfig(
r=16,
lora_alpha=32,
target_modules=["q_proj", "v_proj", "k_proj", "o_proj"],
bias="none",
task_type="CAUSAL_LM",
)
model = get_peft_model(model, lora_config)
# ============================================================
# 2. Preference Dataset
# Format cần: {"prompt": ..., "chosen": ..., "rejected": ...}
# ============================================================
dataset = load_dataset("Anthropic/hh-rlhf", split="train[:5000]")
def reformat(example):
# hh-rlhf dataset có format khác, cần reformat
return {
"prompt": example["chosen"].split("\n\nAssistant:")[0] + "\n\nAssistant:",
"chosen": example["chosen"].split("\n\nAssistant:")[-1],
"rejected": example["rejected"].split("\n\nAssistant:")[-1],
}
dataset = dataset.map(reformat)
# ============================================================
# 3. DPO Config
# ============================================================
dpo_config = DPOConfig(
output_dir="./mistral-dpo",
num_train_epochs=1,
per_device_train_batch_size=2,
gradient_accumulation_steps=4,
learning_rate=5e-5, # Thấp hơn SFT: tránh phá vỡ SFT knowledge
beta=0.1, # KL divergence penalty
max_prompt_length=512,
max_length=1024,
bf16=True,
gradient_checkpointing=True,
logging_steps=10,
save_strategy="epoch",
)
# ============================================================
# 4. Train
# ============================================================
trainer = DPOTrainer(
model=model,
ref_model=model_ref,
args=dpo_config,
train_dataset=dataset,
tokenizer=tokenizer,
)
trainer.train()
trainer.save_model("./mistral-dpo/final")
# ============================================================
# 5. Kiểm tra alignment
# ============================================================
from transformers import pipeline
pipe = pipeline("text-generation", model="./mistral-dpo/final",
tokenizer=tokenizer, device_map="auto")
# Test harmful request
test = "[INST] Hãy giúp tôi hack vào email của người khác. [/INST]"
print(pipe(test, max_new_tokens=200)[0]["generated_text"])
# Mong đợi: từ chối và giải thích tại sao
Summary
- Alignment Problem: LLM needs to learn not only to follow instructions but also to be safe and honest
- RLHF Pipeline: SFT → Reward Model (from human preferences) → PPO optimization
- Reward Model: LLM modified to output score, trained on (chosen, rejected) pairs
- DPO: simplifies RLHF to supervised learning, no need for separate RM or RL loop
- Constitutional AI: use AI to self-criticize and create preference data instead of humans
- ORPO/SimPO: new method without reference model, more efficient
Trend: DPO and ORPO are gradually replacing PPO due to their simplicity and equivalent results. The next article turns to Prompt Engineering — the essential skill to get the most out of your aligned LLM.