簡介
「模型好點了嗎?」—你需要具體的指標來回答,而不是感覺。本文涵蓋了所有重要指標。
1. 指標分類
┌──────────────────────────────────────────────────┐
│ LLM EVALUATION METRICS │
├──────────────────────────────────────────────────┤
│ │
│ Statistical (Lexical) Semantic │
│ ├── BLEU ├── BERTScore │
│ ├── ROUGE ├── Embedding cosine │
│ ├── METEOR └── BLEURT │
│ └── Exact Match │
│ │
│ Training-focused Task-specific │
│ ├── Perplexity ├── Accuracy │
│ └── Loss curve ├── F1 Score │
│ └── Pass@k (code) │
│ │
│ Model-based │
│ ├── LLM-as-a-Judge │
│ └── Pairwise comparison │
└──────────────────────────────────────────────────┘
2. 實施每個指標
2.1 ROUGE(面對回憶的替補)
from rouge_score import rouge_scorer
scorer = rouge_scorer.RougeScorer(['rouge1','rouge2','rougeL'])
scores = scorer.score("reference text here", "generated text here")
print(f"ROUGE-L: {scores['rougeL'].fmeasure:.3f}")
2.2 BERTcore(語意相似度)
from bert_score import score
P, R, F1 = score(["generated"], ["reference"], lang="vi")
print(f"BERTScore F1: {F1.mean():.3f}")
2.3 困惑
import math
# Perplexity = exp(average negative log-likelihood)
# Thấp hơn = model tốt hơn
perplexity = math.exp(avg_loss)
3. 何時使用哪一個指標?
| 任務 | 主要指標 | 中學 |
|---|---|---|
| 文本生成 | ROUGE-L、BERTcore | 法學碩士法官 |
| 分類 | 準確度,F1 | 混淆矩陣 |
| 總結 | ROUGE-1/2/L | BERT 評分 |
| 翻譯 | 藍色、流星 | 人類評估 |
| 程式碼產生 | Pass@k,精確比對 | 測試案例 |
| 對話 | 法學碩士法官 | 人類偏好 |
總結
- BLEU/ROUGE:快速、可重複,但僅測量詞彙重疊
- BERTScore:測量語意相似性-更適合創意任務
- 困惑:訓練診斷,而不是品質指標
- 法學碩士為法官:與人類偏好的相關性最高
- 使用組合-沒有一個單一的指標夠好
練習
- 實作評估套件:ROUGE + BERTScore + Perplexity
- 在基本模型與微調模型上運行 — 比較分數
- 創建可視化:比較指標的雷達圖
- 分析:哪個指標最能反映「真實品質」?