Introduction
Fine-tuning is easy to start but easy to get wrong. This article covers the 10 most common pitfalls and how to fix them.
Top 10 Pitfalls
1. Catastrophic Forgetting
Symptom: Model is good at new tasks but "forgets" old tasks Fix: Reduce learning rate, fewer epochs, add general examples to the dataset
2. Overfitting
Symptom: Training loss reduces but validation loss increases Fix: Add data, regularization, early stopping, reduce epochs
3. Data Leakage
Symptom: Eval scores are very high but production quality is poor Fix: Ensure no overlap between train/test, use temporal split
4. Bad Data Quality
Symptom: The model "learns" incorrectly because the examples are wrong Fix: Manual review random samples, quality scoring pipeline
5. Wrong Granularity
Symptom: Fine-tune for a task is too broad or too narrow Fix: Focus on specific behaviors, don't try to teach "everything"
6. Insufficient Evaluation
Symptom: "Looks good" but there are no specific metrics Fix: Multi-layer evaluation pipeline (lesson 13)
7. Ignoring Base Model Capability
Symptom: Fine-tune for the base model that worked well Fix: Always benchmark the base model first
8. Too Many Epochs
Symptom: Model answers "cliché", repeating training examples Fix: Monitor loss validation, stop when loss plateaus
9. Cost Surprise
Symptom: Unexpectedly high training/inference bill Fix: Calculate costs first (lesson 3), set budget alerts
10. No Versioning
Symptom: "Which model version is best?" — don't know Fix: Version control datasets + models + eval results
Summary
- Fine-tuning is easy to make mistakes if not systematic
- Always benchmark the base model before fine-tuning
- Multi-layer evaluation prevents most pitfalls
- Version control EVERYTHING: data, models, configs, eval results
Exercises
- Intentionally create overfitting (20 epochs) → observe symptoms
- Create a pre-flight checklist for each training job
- Implement model versioning system
- Document 3 pitfalls you have encountered (if any)