Introduction
BLEU, ROUGE measures "surface" — but AI quality must ultimately be judged by humans or smarter AI. These are the two strongest methods.
1. LLM-as-a-Judge
1.1 Pointwise Scoring
JUDGE_PROMPT = """Evaluate the response on a scale of 1-5:
**Task:** {task}
**Response:** {response}
Criteria:
- Accuracy (1-5): Thông tin chính xác không?
- Completeness (1-5): Trả lời đầy đủ không?
- Format (1-5): Đúng format yêu cầu không?
- Tone (1-5): Giọng điệu phù hợp không?
Output JSON: {"accuracy": X, "completeness": X, "format": X, "tone": X, "overall": X, "reasoning": "..."}
"""
1.2 Pairwise Comparison
PAIRWISE_PROMPT = """Compare Response A vs Response B:
Task: {task}
Response A: {response_a}
Response B: {response_b}
Which is better? Output: {"winner": "A" or "B" or "tie", "reasoning": "..."}
"""
2. Human Evaluation
Golden Test Set
golden_tests = [
{
"input": "Triệu chứng COVID-19?",
"expected_elements": ["sốt", "ho", "mệt mỏi", "mất vị giác"],
"expected_format": "bullet list",
"rubric": "Phải mention ít nhất 3/4 triệu chứng chính"
},
# 50+ test cases
]
Inter-Annotator Agreement
from sklearn.metrics import cohen_kappa_score
kappa = cohen_kappa_score(annotator_1_scores, annotator_2_scores)
# kappa > 0.6 = acceptable agreement
Summary
- LLM-as-Judge: fast, scalable, good correlation with humans
- Pairwise comparison: stronger than pointwise scoring
- Human eval: gold standard but expensive and slow
- Golden test set: 50+ curated examples for your domain
- Combining LLM-Judge + Human sampling = best practice
Exercises
- Design judge prompt for your domain (5 rubric criteria)
- Run pairwise comparison: base vs fine-tuned (20 examples)
- Create a golden test set of 30+ examples
- Measure inter-annotator agreement (invite 2 reviewers)