Introduction
Evaluation is not a “one-time run” — it is a continuous pipeline. This article builds an end-to-end eval system.
1. Multi-layer Evaluation
Layer 1: Automated Metrics (ROUGE, BERTScore)
→ Chạy mỗi training iteration
→ Fast, reproducible
Layer 2: LLM-as-Judge
→ Chạy mỗi version release candidate
→ Quality, relevance, format
Layer 3: Golden Test Set
→ 50+ curated test cases
→ Domain-specific validation
Layer 4: Catastrophic Forgetting Check
→ Test model vẫn giỏi tasks chung
→ Không "quên" kiến thức cũ
Layer 5: Red Teaming
→ Adversarial testing
→ Safety, edge cases, prompt injection
2. Catastrophic Forgetting Detection
GENERAL_KNOWLEDGE_TESTS = [
{"q": "Thủ đô Việt Nam là gì?", "a": "Hà Nội"},
{"q": "1 + 1 = ?", "a": "2"},
{"q": "Ai viết Romeo and Juliet?", "a": "Shakespeare"},
]
def check_forgetting(model, threshold=0.8):
correct = 0
for test in GENERAL_KNOWLEDGE_TESTS:
response = call_model(model, test["q"])
if test["a"].lower() in response.lower():
correct += 1
score = correct / len(GENERAL_KNOWLEDGE_TESTS)
if score < threshold:
print(f"⚠️ CATASTROPHIC FORGETTING DETECTED: {score:.0%}")
return score
3. CI/CD for Model Evaluation
# .github/workflows/model-eval.yml
on:
push:
paths: ['training_data/**']
jobs:
evaluate:
runs-on: ubuntu-latest
steps:
- run: python eval/run_metrics.py
- run: python eval/run_llm_judge.py
- run: python eval/check_forgetting.py
- run: python eval/generate_report.py
Summary
- 5 layers evaluation: metrics → LLM-judge → golden set → forgetting → red team
- Catastrophic forgetting: check the model does not "forget" general knowledge
- CI/CD evaluation: automatically runs when data changes
- Red teaming: test adversarial inputs before production
Exercises
- Build a 5-layer evaluation pipeline for your model
- Create catastrophic forgetting test suite (20+ general questions)
- Design red teaming scenarios (10+ adversarial prompts)
- Write an evaluation report comparing base vs fine-tuned