Giới thiệu
Bạn đã xây RAG pipeline. Nhưng nó tốt đến mức nào? Không thể chỉ "cảm thấy" — cần đo lường bằng metrics cụ thể. RAG Evaluation giúp trả lời: retrieval có tìm đúng không? LLM có trả lời chính xác không? Có hallucination không?
Ví dụ: RAG trả lời "Nghỉ phép 20 ngày" nhưng tài liệu ghi "15 ngày". Không đánh giá → không phát hiện lỗi → user mất niềm tin.
Bài này cover:
- RAG Evaluation Metrics — Faithfulness, Relevancy, Precision, Recall
- RAGAS Framework — automated evaluation
- Golden Test Set — tạo bộ test chuẩn
1. RAG Evaluation Metrics
1.1 Hai nhóm metrics
RAG Pipeline: Query → Retrieve → Generate → Answer
Retrieval Metrics (retrieve có tốt không?):
├── Context Precision: bao nhiêu % context là relevant?
└── Context Recall: có miss thông tin quan trọng không?
Generation Metrics (LLM generate có tốt không?):
├── Faithfulness: answer có đúng với context không? (no hallucination)
└── Answer Relevancy: answer có trả lời đúng câu hỏi không?
1.2 Giải thích từng metric
| Metric | Câu hỏi | Range | Mục tiêu |
|---|---|---|---|
| Faithfulness | Answer có bịa thêm không? | 0-1 | > 0.9 |
| Answer Relevancy | Answer có đúng trọng tâm? | 0-1 | > 0.8 |
| Context Precision | Context retrieved có relevant? | 0-1 | > 0.8 |
| Context Recall | Có bỏ sót context quan trọng? | 0-1 | > 0.7 |
1.3 Ví dụ minh họa
Question: "Nghỉ phép bao nhiêu ngày?"
Ground truth: "15 ngày cho full-time, 8 ngày cho part-time"
Retrieved contexts:
C1: "Nhân viên full-time được 15 ngày phép" ← Relevant
C2: "Quy trình xin nghỉ phép..." ← Partially relevant
C3: "Công ty thành lập năm 2020..." ← NOT relevant
Answer: "Nhân viên được 15 ngày nghỉ phép mỗi năm."
Metrics:
├── Faithfulness: 1.0 (answer đúng với context C1)
├── Answer Relevancy: 0.8 (trả lời đúng nhưng thiếu part-time)
├── Context Precision: 0.33 (1/3 context relevant)
└── Context Recall: 0.5 (có full-time, thiếu part-time)
2. RAGAS Framework
2.1 Cài đặt và setup
"""RAGAS — RAG Assessment framework"""
# pip install ragas
from ragas import evaluate
from ragas.metrics import (
faithfulness,
answer_relevancy,
context_precision,
context_recall,
)
from datasets import Dataset
# Chuẩn bị evaluation data
eval_data = {
"question": [
"Nghỉ phép bao nhiêu ngày?",
"Quy trình tuyển dụng thế nào?",
],
"answer": [
"Nhân viên full-time được 15 ngày phép/năm.",
"Tuyển dụng gồm 3 vòng: CV screening, phỏng vấn kỹ thuật, phỏng vấn văn hóa.",
],
"contexts": [
["Nhân viên full-time được 15 ngày phép có lương mỗi năm."],
["Quy trình tuyển dụng 3 vòng: screening CV, phỏng vấn tech, phỏng vấn culture fit."],
],
"ground_truth": [
"15 ngày cho full-time, 8 ngày cho part-time.",
"3 vòng: CV screening, phỏng vấn kỹ thuật, phỏng vấn văn hóa.",
],
}
dataset = Dataset.from_dict(eval_data)
# Evaluate
result = evaluate(
dataset=dataset,
metrics=[faithfulness, answer_relevancy, context_precision, context_recall],
)
print(result)
# {'faithfulness': 0.95, 'answer_relevancy': 0.82,
# 'context_precision': 0.90, 'context_recall': 0.65}
2.2 Hiểu kết quả
Faithfulness: 0.95 ← Tốt! Answer không bịa
Answer Relevancy: 0.82 ← OK, nhưng có thể cải thiện
Context Precision: 0.90 ← Retrieval tìm đúng documents
Context Recall: 0.65 ← ⚠️ Thiếu sót! Cần cải thiện retrieval
Action plan:
├── Context Recall thấp → cải thiện chunking hoặc thêm tài liệu
├── Answer Relevancy trung bình → tune prompt generation
└── Faithfulness cao → generation prompt tốt, ít hallucination
💡 Bài tập 1: Evaluate RAG pipeline hiện tại bằng RAGAS. Tạo 10 câu test với ground truth. Tính 4 metrics. Metric nào thấp nhất?
3. Tạo Golden Test Set
3.1 Tại sao cần Golden Test Set?
Golden Test Set = bộ câu hỏi + đáp án chuẩn (ground truth)
→ Dùng để đánh giá mọi thay đổi trong RAG pipeline
Khi thay đổi chunk size: chạy golden test → so sánh metrics
Khi thay đổi embedding: chạy golden test → so sánh metrics
Khi thêm reranker: chạy golden test → so sánh metrics
Không có golden test = mù → không biết thay đổi tốt hay xấu!
3.2 Tạo từ tài liệu
"""Tự động sinh golden test set từ tài liệu"""
from ragas.testset.generator import TestsetGenerator
from ragas.testset.evolutions import simple, reasoning, multi_context
from langchain_openai import ChatOpenAI, OpenAIEmbeddings
generator_llm = ChatOpenAI(model="gpt-4o")
critic_llm = ChatOpenAI(model="gpt-4o")
embeddings = OpenAIEmbeddings()
# Tạo test set generator
generator = TestsetGenerator.from_langchain(
generator_llm=generator_llm,
critic_llm=critic_llm,
embeddings=embeddings,
)
# Sinh test set từ documents
testset = generator.generate_with_langchain_docs(
documents=chunks, # Documents đã chunk
test_size=20, # 20 câu test
distributions={
simple: 0.5, # 50% câu đơn giản
reasoning: 0.25, # 25% câu cần suy luận
multi_context: 0.25, # 25% câu cần nhiều context
},
)
test_df = testset.to_pandas()
print(test_df[["question", "ground_truth"]].head())
3.3 Đánh giá thủ công
| # | Câu hỏi | Ground Truth | Ai tạo |
|---|---|---|---|
| 1 | Nghỉ phép bao nhiêu ngày? | 15 ngày full-time, 8 ngày part-time | Expert |
| 2 | Quy trình xin nghỉ? | Gửi đơn trước 3 ngày, quản lý duyệt | Expert |
| 3 | Lương trả ngày nào? | Ngày 5 hàng tháng | Auto (RAGAS) |
Rule of thumb: Golden test set nên có ít nhất 50 câu, bao gồm: 50% simple, 30% reasoning, 20% edge cases.
4. End-to-End Evaluation Pipeline
4.1 Automated evaluation
"""Full evaluation pipeline: RAG → Evaluate → Report"""
def evaluate_rag_pipeline(pipeline, golden_testset):
questions = golden_testset["question"]
ground_truths = golden_testset["ground_truth"]
answers = []
contexts = []
for question in questions:
# Chạy RAG pipeline
result = pipeline.invoke(question)
answers.append(result["answer"])
contexts.append([doc.page_content for doc in result["source_documents"]])
# RAGAS evaluate
eval_dataset = Dataset.from_dict({
"question": questions,
"answer": answers,
"contexts": contexts,
"ground_truth": ground_truths,
})
scores = evaluate(
dataset=eval_dataset,
metrics=[faithfulness, answer_relevancy, context_precision, context_recall],
)
return scores
# Compare 2 pipelines
scores_v1 = evaluate_rag_pipeline(pipeline_v1, golden_test)
scores_v2 = evaluate_rag_pipeline(pipeline_v2, golden_test)
print("V1:", scores_v1)
print("V2:", scores_v2)
# V1: {'faithfulness': 0.85, 'context_recall': 0.60}
# V2: {'faithfulness': 0.92, 'context_recall': 0.78} ← Better!
4.2 Evaluation Dashboard
"""Streamlit dashboard cho RAG evaluation"""
import streamlit as st
import pandas as pd
st.title("RAG Evaluation Dashboard")
# Upload golden test
uploaded = st.file_uploader("Upload golden test CSV")
if uploaded:
test_df = pd.read_csv(uploaded)
# Run evaluation
if st.button("Run Evaluation"):
scores = evaluate_rag_pipeline(pipeline, test_df)
# Display metrics
col1, col2, col3, col4 = st.columns(4)
col1.metric("Faithfulness", f"{scores['faithfulness']:.2f}")
col2.metric("Answer Relevancy", f"{scores['answer_relevancy']:.2f}")
col3.metric("Context Precision", f"{scores['context_precision']:.2f}")
col4.metric("Context Recall", f"{scores['context_recall']:.2f}")
# Per-question breakdown
st.dataframe(scores.to_pandas())
💡 Bài tập 2: Tạo golden test set (20+ câu) → evaluate RAG pipeline → thay đổi 1 param (chunk size hoặc top_k) → re-evaluate → so sánh metrics trước/sau.
5. Cải thiện RAG dựa trên Metrics
5.1 Chẩn đoán và hành động
| Metric thấp | Nguyên nhân | Hành động |
|---|---|---|
| Context Recall ↓ | Retrieval miss docs | Thêm tài liệu, cải thiện chunking, thử hybrid search |
| Context Precision ↓ | Retrieve nhiều noise | Thêm reranker, giảm top_k, filter metadata |
| Faithfulness ↓ | LLM hallucinate | Improve prompt ("chỉ dựa trên context"), dùng model lớn hơn |
| Answer Relevancy ↓ | Answer lạc đề | Improve prompt, thêm few-shot examples |
Tóm tắt
| Concept | Ghi nhớ |
|---|---|
| Faithfulness | Answer có đúng với context? (no hallucination) |
| Answer Relevancy | Answer có đúng trọng tâm câu hỏi? |
| Context Precision | Retrieval có chính xác? |
| Context Recall | Retrieval có đầy đủ? |
| RAGAS | Framework đánh giá RAG tự động |
| Golden Test Set | Bộ test chuẩn với ground truth |
| Iterate | Metric thấp → chẩn đoán → cải thiện → re-evaluate |
Bài tập tổng hợp
- ✅ Hoàn thành 2 bài tập nhỏ (1, 2)
- Full Evaluation: Tạo golden test 50 câu → evaluate pipeline hiện tại → xác định bottleneck → thay đổi 1 component → re-evaluate → viết report so sánh.
- CI/CD Evaluation: Tự động chạy RAGAS mỗi khi thay đổi pipeline. Script:
python evaluate.py→ output metrics → fail nếu faithfulness < 0.85. - Human-in-the-loop: Tạo Streamlit app: hiển thị question + answer + context + ground_truth. Người dùng rate 1-5 cho mỗi câu. So sánh human rating vs RAGAS scores.
Bài tiếp theo: Deploy RAG lên Production — API, Caching & Monitoring — đưa RAG từ notebook ra sản phẩm thực tế.