Introduction
"Is the model better?" — you need specific metrics to answer, not feelings. This article covers all important metrics.
1. Taxonomy of Metrics
┌──────────────────────────────────────────────────┐
│ LLM EVALUATION METRICS │
├──────────────────────────────────────────────────┤
│ │
│ Statistical (Lexical) Semantic │
│ ├── BLEU ├── BERTScore │
│ ├── ROUGE ├── Embedding cosine │
│ ├── METEOR └── BLEURT │
│ └── Exact Match │
│ │
│ Training-focused Task-specific │
│ ├── Perplexity ├── Accuracy │
│ └── Loss curve ├── F1 Score │
│ └── Pass@k (code) │
│ │
│ Model-based │
│ ├── LLM-as-a-Judge │
│ └── Pairwise comparison │
└──────────────────────────────────────────────────┘
2. Implement each Metric
2.1 ROUGE (Recall-Oriented Understudy)
from rouge_score import rouge_scorer
scorer = rouge_scorer.RougeScorer(['rouge1','rouge2','rougeL'])
scores = scorer.score("reference text here", "generated text here")
print(f"ROUGE-L: {scores['rougeL'].fmeasure:.3f}")
2.2 BERTScore (Semantic Similarity)
from bert_score import score
P, R, F1 = score(["generated"], ["reference"], lang="vi")
print(f"BERTScore F1: {F1.mean():.3f}")
2.3 Perplexity
import math
# Perplexity = exp(average negative log-likelihood)
# Thấp hơn = model tốt hơn
perplexity = math.exp(avg_loss)
3. When to use which Metric?
| Tasks | Primary Metric | Secondary |
|---|---|---|
| Text generation | ROUGE-L, BERTScore | LLM-Judge |
| Classification | Accuracy, F1 | Confusion matrix |
| Summarization | ROUGE-1/2/L | BERTScore |
| Translation | BLEU, METEOR | Human eval |
| Code generation | Pass@k, Exact Match | Test cases |
| Conversational | LLM-as-Judge | Human preference |
Summary
- BLEU/ROUGE: Fast, reproducible, but only measures lexical overlap
- BERTScore: Measures semantic similarity — better for creative tasks
- Perplexity: Training diagnostic, not quality metric
- LLM-as-Judge: Highest correlation with human preference
- Use combination — no single metric is good enough
Exercises
- Implement evaluation suite: ROUGE + BERTScore + Perplexity
- Run on base model vs fine-tuned model — compare scores
- Create visualization: radar chart comparing metrics
- Analysis: which metric best reflects "true quality"?