Chuyển đến nội dung chính

Bài 12: RAG Evaluation — RAGAS, Faithfulness & Relevancy

Đánh giá RAG bằng RAGAS framework. Metrics: Faithfulness, Answer Relevancy, Context Precision, Context Recall. Tạo golden test set, automated evaluation.

🧠 AI & ML — Bài 11 Bài 12: RAG Evaluation — RAGAS, Faithfulness & Relevancy

RAG Thực Chiến: Từ Basic đến Advanced

Phần 4: Advanced RAG Patterns

xdev.asia

Giới thiệu

Bạn đã xây RAG pipeline. Nhưng nó tốt đến mức nào? Không thể chỉ "cảm thấy" — cần đo lường bằng metrics cụ thể. RAG Evaluation giúp trả lời: retrieval có tìm đúng không? LLM có trả lời chính xác không? Có hallucination không?

Ví dụ: RAG trả lời "Nghỉ phép 20 ngày" nhưng tài liệu ghi "15 ngày". Không đánh giá → không phát hiện lỗi → user mất niềm tin.

Bài này cover:

  1. RAG Evaluation Metrics — Faithfulness, Relevancy, Precision, Recall
  2. RAGAS Framework — automated evaluation
  3. Golden Test Set — tạo bộ test chuẩn

1. RAG Evaluation Metrics

1.1 Hai nhóm metrics

RAG Pipeline:  Query → Retrieve → Generate → Answer

Retrieval Metrics (retrieve có tốt không?):
├── Context Precision: bao nhiêu % context là relevant?
└── Context Recall: có miss thông tin quan trọng không?

Generation Metrics (LLM generate có tốt không?):
├── Faithfulness: answer có đúng với context không? (no hallucination)
└── Answer Relevancy: answer có trả lời đúng câu hỏi không?

1.2 Giải thích từng metric

MetricCâu hỏiRangeMục tiêu
FaithfulnessAnswer có bịa thêm không?0-1> 0.9
Answer RelevancyAnswer có đúng trọng tâm?0-1> 0.8
Context PrecisionContext retrieved có relevant?0-1> 0.8
Context RecallCó bỏ sót context quan trọng?0-1> 0.7

1.3 Ví dụ minh họa

Question: "Nghỉ phép bao nhiêu ngày?"
Ground truth: "15 ngày cho full-time, 8 ngày cho part-time"

Retrieved contexts:
  C1: "Nhân viên full-time được 15 ngày phép" ← Relevant
  C2: "Quy trình xin nghỉ phép..."           ← Partially relevant
  C3: "Công ty thành lập năm 2020..."         ← NOT relevant

Answer: "Nhân viên được 15 ngày nghỉ phép mỗi năm."

Metrics:
├── Faithfulness: 1.0 (answer đúng với context C1)
├── Answer Relevancy: 0.8 (trả lời đúng nhưng thiếu part-time)
├── Context Precision: 0.33 (1/3 context relevant)
└── Context Recall: 0.5 (có full-time, thiếu part-time)

2. RAGAS Framework

2.1 Cài đặt và setup

"""RAGAS — RAG Assessment framework"""
# pip install ragas

from ragas import evaluate
from ragas.metrics import (
    faithfulness,
    answer_relevancy,
    context_precision,
    context_recall,
)
from datasets import Dataset

# Chuẩn bị evaluation data
eval_data = {
    "question": [
        "Nghỉ phép bao nhiêu ngày?",
        "Quy trình tuyển dụng thế nào?",
    ],
    "answer": [
        "Nhân viên full-time được 15 ngày phép/năm.",
        "Tuyển dụng gồm 3 vòng: CV screening, phỏng vấn kỹ thuật, phỏng vấn văn hóa.",
    ],
    "contexts": [
        ["Nhân viên full-time được 15 ngày phép có lương mỗi năm."],
        ["Quy trình tuyển dụng 3 vòng: screening CV, phỏng vấn tech, phỏng vấn culture fit."],
    ],
    "ground_truth": [
        "15 ngày cho full-time, 8 ngày cho part-time.",
        "3 vòng: CV screening, phỏng vấn kỹ thuật, phỏng vấn văn hóa.",
    ],
}

dataset = Dataset.from_dict(eval_data)

# Evaluate
result = evaluate(
    dataset=dataset,
    metrics=[faithfulness, answer_relevancy, context_precision, context_recall],
)

print(result)
# {'faithfulness': 0.95, 'answer_relevancy': 0.82,
#  'context_precision': 0.90, 'context_recall': 0.65}

2.2 Hiểu kết quả

Faithfulness: 0.95  ← Tốt! Answer không bịa
Answer Relevancy: 0.82  ← OK, nhưng có thể cải thiện
Context Precision: 0.90  ← Retrieval tìm đúng documents
Context Recall: 0.65  ← ⚠️ Thiếu sót! Cần cải thiện retrieval

Action plan:
├── Context Recall thấp → cải thiện chunking hoặc thêm tài liệu
├── Answer Relevancy trung bình → tune prompt generation
└── Faithfulness cao → generation prompt tốt, ít hallucination

💡 Bài tập 1: Evaluate RAG pipeline hiện tại bằng RAGAS. Tạo 10 câu test với ground truth. Tính 4 metrics. Metric nào thấp nhất?


3. Tạo Golden Test Set

3.1 Tại sao cần Golden Test Set?

Golden Test Set = bộ câu hỏi + đáp án chuẩn (ground truth)
→ Dùng để đánh giá mọi thay đổi trong RAG pipeline

Khi thay đổi chunk size:  chạy golden test → so sánh metrics
Khi thay đổi embedding:   chạy golden test → so sánh metrics
Khi thêm reranker:        chạy golden test → so sánh metrics

Không có golden test = mù → không biết thay đổi tốt hay xấu!

3.2 Tạo từ tài liệu

"""Tự động sinh golden test set từ tài liệu"""
from ragas.testset.generator import TestsetGenerator
from ragas.testset.evolutions import simple, reasoning, multi_context
from langchain_openai import ChatOpenAI, OpenAIEmbeddings

generator_llm = ChatOpenAI(model="gpt-4o")
critic_llm = ChatOpenAI(model="gpt-4o")
embeddings = OpenAIEmbeddings()

# Tạo test set generator
generator = TestsetGenerator.from_langchain(
    generator_llm=generator_llm,
    critic_llm=critic_llm,
    embeddings=embeddings,
)

# Sinh test set từ documents
testset = generator.generate_with_langchain_docs(
    documents=chunks,   # Documents đã chunk
    test_size=20,        # 20 câu test
    distributions={
        simple: 0.5,       # 50% câu đơn giản
        reasoning: 0.25,   # 25% câu cần suy luận
        multi_context: 0.25,  # 25% câu cần nhiều context
    },
)

test_df = testset.to_pandas()
print(test_df[["question", "ground_truth"]].head())

3.3 Đánh giá thủ công

#Câu hỏiGround TruthAi tạo
1Nghỉ phép bao nhiêu ngày?15 ngày full-time, 8 ngày part-timeExpert
2Quy trình xin nghỉ?Gửi đơn trước 3 ngày, quản lý duyệtExpert
3Lương trả ngày nào?Ngày 5 hàng thángAuto (RAGAS)

Rule of thumb: Golden test set nên có ít nhất 50 câu, bao gồm: 50% simple, 30% reasoning, 20% edge cases.


4. End-to-End Evaluation Pipeline

4.1 Automated evaluation

"""Full evaluation pipeline: RAG → Evaluate → Report"""
def evaluate_rag_pipeline(pipeline, golden_testset):
    questions = golden_testset["question"]
    ground_truths = golden_testset["ground_truth"]
    
    answers = []
    contexts = []
    
    for question in questions:
        # Chạy RAG pipeline
        result = pipeline.invoke(question)
        answers.append(result["answer"])
        contexts.append([doc.page_content for doc in result["source_documents"]])
    
    # RAGAS evaluate
    eval_dataset = Dataset.from_dict({
        "question": questions,
        "answer": answers,
        "contexts": contexts,
        "ground_truth": ground_truths,
    })
    
    scores = evaluate(
        dataset=eval_dataset,
        metrics=[faithfulness, answer_relevancy, context_precision, context_recall],
    )
    
    return scores

# Compare 2 pipelines
scores_v1 = evaluate_rag_pipeline(pipeline_v1, golden_test)
scores_v2 = evaluate_rag_pipeline(pipeline_v2, golden_test)

print("V1:", scores_v1)
print("V2:", scores_v2)
# V1: {'faithfulness': 0.85, 'context_recall': 0.60}
# V2: {'faithfulness': 0.92, 'context_recall': 0.78}  ← Better!

4.2 Evaluation Dashboard

"""Streamlit dashboard cho RAG evaluation"""
import streamlit as st
import pandas as pd

st.title("RAG Evaluation Dashboard")

# Upload golden test
uploaded = st.file_uploader("Upload golden test CSV")

if uploaded:
    test_df = pd.read_csv(uploaded)
    
    # Run evaluation
    if st.button("Run Evaluation"):
        scores = evaluate_rag_pipeline(pipeline, test_df)
        
        # Display metrics
        col1, col2, col3, col4 = st.columns(4)
        col1.metric("Faithfulness", f"{scores['faithfulness']:.2f}")
        col2.metric("Answer Relevancy", f"{scores['answer_relevancy']:.2f}")
        col3.metric("Context Precision", f"{scores['context_precision']:.2f}")
        col4.metric("Context Recall", f"{scores['context_recall']:.2f}")
        
        # Per-question breakdown
        st.dataframe(scores.to_pandas())

💡 Bài tập 2: Tạo golden test set (20+ câu) → evaluate RAG pipeline → thay đổi 1 param (chunk size hoặc top_k) → re-evaluate → so sánh metrics trước/sau.


5. Cải thiện RAG dựa trên Metrics

5.1 Chẩn đoán và hành động

Metric thấpNguyên nhânHành động
Context Recall ↓Retrieval miss docsThêm tài liệu, cải thiện chunking, thử hybrid search
Context Precision ↓Retrieve nhiều noiseThêm reranker, giảm top_k, filter metadata
Faithfulness ↓LLM hallucinateImprove prompt ("chỉ dựa trên context"), dùng model lớn hơn
Answer Relevancy ↓Answer lạc đềImprove prompt, thêm few-shot examples

Tóm tắt

ConceptGhi nhớ
FaithfulnessAnswer có đúng với context? (no hallucination)
Answer RelevancyAnswer có đúng trọng tâm câu hỏi?
Context PrecisionRetrieval có chính xác?
Context RecallRetrieval có đầy đủ?
RAGASFramework đánh giá RAG tự động
Golden Test SetBộ test chuẩn với ground truth
IterateMetric thấp → chẩn đoán → cải thiện → re-evaluate

Bài tập tổng hợp

  1. ✅ Hoàn thành 2 bài tập nhỏ (1, 2)
  2. Full Evaluation: Tạo golden test 50 câu → evaluate pipeline hiện tại → xác định bottleneck → thay đổi 1 component → re-evaluate → viết report so sánh.
  3. CI/CD Evaluation: Tự động chạy RAGAS mỗi khi thay đổi pipeline. Script: python evaluate.py → output metrics → fail nếu faithfulness < 0.85.
  4. Human-in-the-loop: Tạo Streamlit app: hiển thị question + answer + context + ground_truth. Người dùng rate 1-5 cho mỗi câu. So sánh human rating vs RAGAS scores.

Bài tiếp theo: Deploy RAG lên Production — API, Caching & Monitoring — đưa RAG từ notebook ra sản phẩm thực tế.