Chuyển đến nội dung chính

Lesson 11: LLM Evaluation Metrics — From Perplexity to BERTScore

Comprehensive guide metrics: Perplexity, BLEU, ROUGE, METEOR, BERTScore, Exact Match, F1. When to use which metric? Code implements each metric.

🧠 AI & ML — Lesson 10 Lesson 11: LLM Evaluation Metrics — From Perplexity to BERTScore

Fine-tuning LLM: The Art of AI Tuning

Part 5: Model Evaluation — Methods & Metrics

xdev.asia

Introduction

"Is the model better?" — you need specific metrics to answer, not feelings. This article covers all important metrics.


1. Taxonomy of Metrics

┌──────────────────────────────────────────────────┐
│              LLM EVALUATION METRICS               │
├──────────────────────────────────────────────────┤
│                                                  │
│  Statistical (Lexical)     Semantic              │
│  ├── BLEU                  ├── BERTScore         │  
│  ├── ROUGE                 ├── Embedding cosine  │
│  ├── METEOR                └── BLEURT            │
│  └── Exact Match                                 │
│                                                  │
│  Training-focused          Task-specific         │
│  ├── Perplexity            ├── Accuracy          │
│  └── Loss curve            ├── F1 Score          │
│                            └── Pass@k (code)     │
│                                                  │
│  Model-based                                     │
│  ├── LLM-as-a-Judge                              │
│  └── Pairwise comparison                         │
└──────────────────────────────────────────────────┘

2. Implement each Metric

2.1 ROUGE (Recall-Oriented Understudy)

from rouge_score import rouge_scorer
scorer = rouge_scorer.RougeScorer(['rouge1','rouge2','rougeL'])
scores = scorer.score("reference text here", "generated text here")
print(f"ROUGE-L: {scores['rougeL'].fmeasure:.3f}")

2.2 BERTScore (Semantic Similarity)

from bert_score import score
P, R, F1 = score(["generated"], ["reference"], lang="vi")
print(f"BERTScore F1: {F1.mean():.3f}")

2.3 Perplexity

import math
# Perplexity = exp(average negative log-likelihood)
# Thấp hơn = model tốt hơn
perplexity = math.exp(avg_loss)

3. When to use which Metric?

TasksPrimary MetricSecondary
Text generationROUGE-L, BERTScoreLLM-Judge
ClassificationAccuracy, F1Confusion matrix
SummarizationROUGE-1/2/LBERTScore
TranslationBLEU, METEORHuman eval
Code generationPass@k, Exact MatchTest cases
ConversationalLLM-as-JudgeHuman preference

Summary

  • BLEU/ROUGE: Fast, reproducible, but only measures lexical overlap
  • BERTScore: Measures semantic similarity — better for creative tasks
  • Perplexity: Training diagnostic, not quality metric
  • LLM-as-Judge: Highest correlation with human preference
  • Use combination — no single metric is good enough

Exercises

  1. Implement evaluation suite: ROUGE + BERTScore + Perplexity
  2. Run on base model vs fine-tuned model — compare scores
  3. Create visualization: radar chart comparing metrics
  4. Analysis: which metric best reflects "true quality"?