Introduction
"Fine-tuning" is one of the most talked about buzzwords in AI — but also the most abused technique. Before jumping into the code, you need to understand: What Fine-tuning really is, where it is in the AI pipeline, and when you really need it.
⚠️ Golden Rule: 80% of the time you think you need fine-tuning, actually prompt engineering or RAG is enough. Fine-tuning is the last choice, not the first.
1. Life cycle of an LLM
Before understanding fine-tuning, let's see how LLM is created:
┌──────────────────────────────────────────────────────────────────┐
│ VÒNG ĐỜI MỘT LLM │
│ │
│ Phase 1: PRE-TRAINING │
│ ┌─────────────────────────────────────────────────────────┐ │
│ │ Huấn luyện trên TOÀN BỘ internet (~15 nghìn tỷ tokens) │ │
│ │ → Học ngôn ngữ, kiến thức, lập luận chung │ │
│ │ Cost: $10M–$100M+ | Time: Weeks–Months | GPUs: Hàng nghìn│ │
│ └─────────────────────────────────────────────────────────┘ │
│ │ │
│ ▼ │
│ Phase 2: SUPERVISED FINE-TUNING (SFT) ← Bạn đang ở đây │
│ ┌─────────────────────────────────────────────────────────┐ │
│ │ Huấn luyện thêm trên dataset nhỏ, chất lượng cao │ │
│ │ → Dạy model cách tuân thủ instructions, format, style │ │
│ │ Cost: $10–$10,000 | Time: Minutes–Hours | GPUs: 1–8 │ │
│ └─────────────────────────────────────────────────────────┘ │
│ │ │
│ ▼ │
│ Phase 3: ALIGNMENT (RLHF / DPO) │
│ ┌─────────────────────────────────────────────────────────┐ │
│ │ Tinh chỉnh model theo preferences con người │ │
│ │ → An toàn, helpful, honest │ │
│ │ Cost: $1,000–$50,000 | Cần human annotators │ │
│ └─────────────────────────────────────────────────────────┘ │
└──────────────────────────────────────────────────────────────────┘
Bottom line
- Pre-training: Let the model "read" the entire internet → know everything but don't know how to answer
- SFT (Fine-tuning): Teach the model how to respond, in the format/style you want
- RLHF/DPO: Refine the model to respond "correctly" to humans
When people say "fine-tuning", it usually means Phase 2 — SFT.
2. What problem does Fine-tuning solve?
2.1 Behavior vs Knowledge
Here's the most important distinction to decide whether fine-tuning is needed:
| Problem | Type | Solution |
|---|---|---|
| Model does not know your company's products | Knowledge gap | RAG |
| Model response must be in Vietnamese, specific JSON format | Behavior gap | Fine-tuning |
| Model needs real-time data (stock prices, weather) | Knowledge gap | RAG / Tool use |
| Model does not use the correct industry terminology | Behavior gap | Fine-tuning |
| Model "talks too much", you need to answer briefly | Behavior gap | Fine-tuning (or prompt) |
| Model does not know the latest internal policy | Knowledge gap | RAG |
2.2 Specific examples
❌ NO fine-tune needed:
- "I want the chatbot to know about the company's products" → Use RAG
- "I want the model to accurately answer questions from the document" → Use RAG
- "I need a model to read the database and respond" → Use Tool Use / Agent
✅ NEED fine-tune:
- "Model should always respond in JSON with a specific schema" → Fine-tune
- "The model needs to use its own brand tone, very different from the default" → Fine-tune
- "The small model (Flash/Mini) needs to perform like the big model (Pro/4o)" → Fine-tune (distillation)
- "Model must understand Vietnamese medical terminology" → Fine-tune + RAG
3. Decision Framework: 3-step ladder
Before fine-tuning, go through the 3 steps in order:
Bước 1: PROMPT ENGINEERING
├── Chi phí: $0 | Thời gian: Phút
├── Thử: System prompt tốt hơn, few-shot examples, chain-of-thought
├── Đủ tốt? → DỪNG ✅
└── Không đủ? → Bước 2
Bước 2: RAG (Retrieval-Augmented Generation)
├── Chi phí: $50–$500 setup | Thời gian: Ngày
├── Thử: Kết nối knowledge base, vector DB
├── Đủ tốt? → DỪNG ✅
└── Không đủ? → Bước 3
Bước 3: FINE-TUNING
├── Chi phí: $50–$10,000+ | Thời gian: Days–Weeks
├── Chuẩn bị data, train, evaluate, iterate
└── Đây là lựa chọn cuối cùng
Checklist before fine-tuning
- Tried at least 5 different system prompt versions?
- Tried few-shot prompting (3–5 examples in prompt)?
- If you need new knowledge → tried RAG?
- Have at least 100 high quality training data examples?
- Is there a budget for training + evaluation iterations?
- Is there a long-term maintenance period for the model?
4. Fine-tuning methods
4.1 Full Fine-tuning
- Update all model weights
- Needs a huge GPU (A100 80GB+)
- High costs, catastrophic forgetting risk
- Rarely needed in practice 2025–2026
4.2 Supervised Fine-Tuning (SFT) via API
- Use Google/OpenAI API
- No need for GPU management
- Fast, easy, moderate cost
- This is the most common method
4.3 LoRA / QLoRA (Parameter-Efficient)
- Only updates a small portion of weights (~0.1–1%)
- Can run on consumer GPUs (RTX 3090, T4)
- Lowest cost
- Need more technical knowledge
┌─────────────────────────────────────────────────┐
│ So sánh 3 phương pháp │
│ │
│ Full FT ████████████████████ $$$$$ 🔥🔥🔥 │
│ SFT (API) ██████████ $$ 🔥 │
│ LoRA ████ $ 🔥🔥 │
│ │
│ ████ = Resources $$$ = Cost 🔥 = Difficulty│
└─────────────────────────────────────────────────┘
5. Landscape 2025–2026: Who provides Fine-tuning?
| Provider | Supported Models | Pricing model | Advantages |
|---|---|---|---|
| Google (Vertex AI) | Gemini 2.0 Flash, Pro | Per training token | Inference price by base model |
| OpenAI | GPT-4o-mini, GPT-4o | Per training token | Ecosystem is large, easy to use |
| Anthropic | Not yet published SFT | N/A | Not available for end users |
| Hugging Face | All open-source models | Self-hosted / AutoTrain | Most flexible, cheapest |
| Together AI | LLaMA, Mistral, ... | Per training token | Infrastructure management |
| Unsloth | LLaMA, Mistral, ... | Free (self-hosted) | 2x faster, less VRAM |
6. Quick demo: Fine-tuning "Hello World"
Enjoy the simplest fine-tuning — using OpenAI API:
from openai import OpenAI
import json
client = OpenAI()
# Step 1: Chuẩn bị dataset (tối thiểu ~10 examples cho demo)
training_data = [
{"messages": [
{"role": "system", "content": "Bạn là trợ lý hỗ trợ khách hàng của XDev."},
{"role": "user", "content": "Giá dịch vụ thế nào?"},
{"role": "assistant", "content": "Dạ, XDev có 3 gói: Starter (miễn phí), Pro ($29/tháng), Enterprise (liên hệ). Anh/chị quan tâm gói nào ạ? 😊"}
]},
{"messages": [
{"role": "system", "content": "Bạn là trợ lý hỗ trợ khách hàng của XDev."},
{"role": "user", "content": "Làm sao liên hệ support?"},
{"role": "assistant", "content": "Dạ, anh/chị có thể liên hệ qua: 📧 [email protected] hoặc 💬 chat trực tiếp trên website. Team support online 9h-18h (GMT+7) các ngày trong tuần ạ!"}
]},
# ... thêm 8+ examples nữa
]
# Step 2: Save thành JSONL file
with open("training_data.jsonl", "w") as f:
for item in training_data:
f.write(json.dumps(item, ensure_ascii=False) + "\n")
# Step 3: Upload file
file = client.files.create(
file=open("training_data.jsonl", "rb"),
purpose="fine-tune"
)
# Step 4: Tạo fine-tuning job
job = client.fine_tuning.jobs.create(
training_file=file.id,
model="gpt-4o-mini-2024-07-18",
hyperparameters={"n_epochs": 3}
)
print(f"Job ID: {job.id}")
print(f"Status: {job.status}") # → "validating_files" → "running" → "succeeded"
💡 Note: This is just a demo flow. Lessons 7 & 9 will go into detail with actual datasets.
Lesson summary
- Fine-tuning = teaching the model how to behave, not teaching knowledge
- Knowledge gap → use RAG | Behavior gap → use Fine-tuning
- Always go through 3 steps: Prompt Engineering → RAG → Fine-tuning
- 3 methods: Full FT (rare) | SFT via API (popular) | LoRA (thrifty)
- Google Vertex AI + OpenAI are the two main platforms for API fine-tuning
- LoRA/QLoRA for self-hosted, lowest cost
Exercises
- List 3 AI issues in your work — classify Knowledge gap vs Behavior gap
- For each problem, suggest a solution: Prompt Engineering, RAG, or Fine-tuning?
- Create 10 training examples (JSONL format) for the use case you are interested in
- Read the blog post "A Practical Guide to Fine-Tuning" on OpenAI Cookbook