簡介
評估不是「一次性運行」——它是一個連續的管道。本文建構了一個端到端的評估系統。
1. 多層評估
Layer 1: Automated Metrics (ROUGE, BERTScore)
→ Chạy mỗi training iteration
→ Fast, reproducible
Layer 2: LLM-as-Judge
→ Chạy mỗi version release candidate
→ Quality, relevance, format
Layer 3: Golden Test Set
→ 50+ curated test cases
→ Domain-specific validation
Layer 4: Catastrophic Forgetting Check
→ Test model vẫn giỏi tasks chung
→ Không "quên" kiến thức cũ
Layer 5: Red Teaming
→ Adversarial testing
→ Safety, edge cases, prompt injection
2. 災難性遺忘偵測
GENERAL_KNOWLEDGE_TESTS = [
{"q": "Thủ đô Việt Nam là gì?", "a": "Hà Nội"},
{"q": "1 + 1 = ?", "a": "2"},
{"q": "Ai viết Romeo and Juliet?", "a": "Shakespeare"},
]
def check_forgetting(model, threshold=0.8):
correct = 0
for test in GENERAL_KNOWLEDGE_TESTS:
response = call_model(model, test["q"])
if test["a"].lower() in response.lower():
correct += 1
score = correct / len(GENERAL_KNOWLEDGE_TESTS)
if score < threshold:
print(f"⚠️ CATASTROPHIC FORGETTING DETECTED: {score:.0%}")
return score
3. 用於模型評估的 CI/CD
# .github/workflows/model-eval.yml
on:
push:
paths: ['training_data/**']
jobs:
evaluate:
runs-on: ubuntu-latest
steps:
- run: python eval/run_metrics.py
- run: python eval/run_llm_judge.py
- run: python eval/check_forgetting.py
- run: python eval/generate_report.py
總結
- 5層評估:指標→LLM法官→黃金組→遺忘→紅隊
- 災難性遺忘:檢查模型是否「忘記」常識
- CI/CD評估:資料變化時自動運行
- 紅隊:在生產前測試對抗性輸入
練習
- 為您的模型建立 5 層評估管道
- 創建災難性遺忘測試套件(20+一般問題) 3.設計紅隊場景(10+對抗性提示)
- 撰寫一份比較基礎與微調的評估報告