Chuyển đến nội dung chính

Lesson 12: RAG Evaluation — RAGAS, Faithfulness & Relevancy

Evaluate RAG using the RAGAS framework. Metrics: Faithfulness, Answer Relevancy, Context Precision, Context Recall. Create golden test set, automated evaluation.

🧠 AI & ML — Lesson 11 Lesson 12: RAG Evaluation — RAGAS, Faithfulness & Relevancy

Real Battle RAG: From Basic to Advanced

Part 4: Advanced RAG Patterns

xdev.asia

Introduction

You have built the RAG pipeline. But how good is it? It can't just be "felt" — it needs to be measured with specific metrics. RAG Evaluation helps answer: is retrieval correct? Does LLM answer correctly? Is there hallucination?

Example: RAG replies "20 days leave" but the document says "15 days". Not evaluating → not detecting errors → users lose trust.

This article covers:

  1. RAG Evaluation Metrics — Faithfulness, Relevancy, Precision, Recall
  2. RAGAS Framework — automated evaluation
  3. Golden Test Set — create a standard test set

1. RAG Evaluation Metrics

1.1 Two groups of metrics

RAG Pipeline:  Query → Retrieve → Generate → Answer

Retrieval Metrics (retrieve có tốt không?):
├── Context Precision: bao nhiêu % context là relevant?
└── Context Recall: có miss thông tin quan trọng không?

Generation Metrics (LLM generate có tốt không?):
├── Faithfulness: answer có đúng với context không? (no hallucination)
└── Answer Relevancy: answer có trả lời đúng câu hỏi không?

1.2 Explain each metric

MetricsQuestionRangeGoal
FaithfulnessIs Answer making up more?0-1> 0.9
Answer RelevancyIs the answer on point?0-1> 0.8
Context PrecisionIs Context retrieval relevant?0-1> 0.8
Context RecallMissing important context?0-1> 0.7

1.3 Illustrative example

Question: "Nghỉ phép bao nhiêu ngày?"
Ground truth: "15 ngày cho full-time, 8 ngày cho part-time"

Retrieved contexts:
  C1: "Nhân viên full-time được 15 ngày phép" ← Relevant
  C2: "Quy trình xin nghỉ phép..."           ← Partially relevant
  C3: "Công ty thành lập năm 2020..."         ← NOT relevant

Answer: "Nhân viên được 15 ngày nghỉ phép mỗi năm."

Metrics:
├── Faithfulness: 1.0 (answer đúng với context C1)
├── Answer Relevancy: 0.8 (trả lời đúng nhưng thiếu part-time)
├── Context Precision: 0.33 (1/3 context relevant)
└── Context Recall: 0.5 (có full-time, thiếu part-time)

2. RAGAS Framework

2.1 Installation and setup

"""RAGAS — RAG Assessment framework"""
# pip install ragas

from ragas import evaluate
from ragas.metrics import (
    faithfulness,
    answer_relevancy,
    context_precision,
    context_recall,
)
from datasets import Dataset

# Chuẩn bị evaluation data
eval_data = {
    "question": [
        "Nghỉ phép bao nhiêu ngày?",
        "Quy trình tuyển dụng thế nào?",
    ],
    "answer": [
        "Nhân viên full-time được 15 ngày phép/năm.",
        "Tuyển dụng gồm 3 vòng: CV screening, phỏng vấn kỹ thuật, phỏng vấn văn hóa.",
    ],
    "contexts": [
        ["Nhân viên full-time được 15 ngày phép có lương mỗi năm."],
        ["Quy trình tuyển dụng 3 vòng: screening CV, phỏng vấn tech, phỏng vấn culture fit."],
    ],
    "ground_truth": [
        "15 ngày cho full-time, 8 ngày cho part-time.",
        "3 vòng: CV screening, phỏng vấn kỹ thuật, phỏng vấn văn hóa.",
    ],
}

dataset = Dataset.from_dict(eval_data)

# Evaluate
result = evaluate(
    dataset=dataset,
    metrics=[faithfulness, answer_relevancy, context_precision, context_recall],
)

print(result)
# {'faithfulness': 0.95, 'answer_relevancy': 0.82,
#  'context_precision': 0.90, 'context_recall': 0.65}

2.2 Understand the results

Faithfulness: 0.95  ← Tốt! Answer không bịa
Answer Relevancy: 0.82  ← OK, nhưng có thể cải thiện
Context Precision: 0.90  ← Retrieval tìm đúng documents
Context Recall: 0.65  ← ⚠️ Thiếu sót! Cần cải thiện retrieval

Action plan:
├── Context Recall thấp → cải thiện chunking hoặc thêm tài liệu
├── Answer Relevancy trung bình → tune prompt generation
└── Faithfulness cao → generation prompt tốt, ít hallucination

💡 Exercise 1: Evaluate the current RAG pipeline using RAGAS. Create 10 test questions with ground truth. Calculate 4 metrics. Which metric is the lowest?


3. Create Golden Test Set

3.1 Why do we need the Golden Test Set?

Golden Test Set = bộ câu hỏi + đáp án chuẩn (ground truth)
→ Dùng để đánh giá mọi thay đổi trong RAG pipeline

Khi thay đổi chunk size:  chạy golden test → so sánh metrics
Khi thay đổi embedding:   chạy golden test → so sánh metrics
Khi thêm reranker:        chạy golden test → so sánh metrics

Không có golden test = mù → không biết thay đổi tốt hay xấu!

3.2 Create from document

"""Tự động sinh golden test set từ tài liệu"""
from ragas.testset.generator import TestsetGenerator
from ragas.testset.evolutions import simple, reasoning, multi_context
from langchain_openai import ChatOpenAI, OpenAIEmbeddings

generator_llm = ChatOpenAI(model="gpt-4o")
critic_llm = ChatOpenAI(model="gpt-4o")
embeddings = OpenAIEmbeddings()

# Tạo test set generator
generator = TestsetGenerator.from_langchain(
    generator_llm=generator_llm,
    critic_llm=critic_llm,
    embeddings=embeddings,
)

# Sinh test set từ documents
testset = generator.generate_with_langchain_docs(
    documents=chunks,   # Documents đã chunk
    test_size=20,        # 20 câu test
    distributions={
        simple: 0.5,       # 50% câu đơn giản
        reasoning: 0.25,   # 25% câu cần suy luận
        multi_context: 0.25,  # 25% câu cần nhiều context
    },
)

test_df = testset.to_pandas()
print(test_df[["question", "ground_truth"]].head())

3.3 Manual review

#QuestionGround TruthWho created
1How many days off?15 days full-time, 8 days part-timeExpert
2Procedure for applying for leave?Submit application 3 days in advance, manager approvesExpert
3What day is salary paid?5th day of every monthAuto (RAGAS)

Rule of thumb: Golden test set should have at least 50 questions, including: 50% simple, 30% reasoning, 20% edge cases.


4. End-to-End Evaluation Pipeline

4.1 Automated evaluation

"""Full evaluation pipeline: RAG → Evaluate → Report"""
def evaluate_rag_pipeline(pipeline, golden_testset):
    questions = golden_testset["question"]
    ground_truths = golden_testset["ground_truth"]
    
    answers = []
    contexts = []
    
    for question in questions:
        # Chạy RAG pipeline
        result = pipeline.invoke(question)
        answers.append(result["answer"])
        contexts.append([doc.page_content for doc in result["source_documents"]])
    
    # RAGAS evaluate
    eval_dataset = Dataset.from_dict({
        "question": questions,
        "answer": answers,
        "contexts": contexts,
        "ground_truth": ground_truths,
    })
    
    scores = evaluate(
        dataset=eval_dataset,
        metrics=[faithfulness, answer_relevancy, context_precision, context_recall],
    )
    
    return scores

# Compare 2 pipelines
scores_v1 = evaluate_rag_pipeline(pipeline_v1, golden_test)
scores_v2 = evaluate_rag_pipeline(pipeline_v2, golden_test)

print("V1:", scores_v1)
print("V2:", scores_v2)
# V1: {'faithfulness': 0.85, 'context_recall': 0.60}
# V2: {'faithfulness': 0.92, 'context_recall': 0.78}  ← Better!

4.2 Evaluation Dashboard

"""Streamlit dashboard cho RAG evaluation"""
import streamlit as st
import pandas as pd

st.title("RAG Evaluation Dashboard")

# Upload golden test
uploaded = st.file_uploader("Upload golden test CSV")

if uploaded:
    test_df = pd.read_csv(uploaded)
    
    # Run evaluation
    if st.button("Run Evaluation"):
        scores = evaluate_rag_pipeline(pipeline, test_df)
        
        # Display metrics
        col1, col2, col3, col4 = st.columns(4)
        col1.metric("Faithfulness", f"{scores['faithfulness']:.2f}")
        col2.metric("Answer Relevancy", f"{scores['answer_relevancy']:.2f}")
        col3.metric("Context Precision", f"{scores['context_precision']:.2f}")
        col4.metric("Context Recall", f"{scores['context_recall']:.2f}")
        
        # Per-question breakdown
        st.dataframe(scores.to_pandas())

💡 Exercise 2: Create golden test set (20+ sentences) → evaluate RAG pipeline → change 1 param (chunk size or top_k) → re-evaluate → compare metrics before/after.


5. Improve RAG based on Metrics

5.1 Diagnosis and actions

Low MetricCauseAction
Context Recall ↓Retrieval miss docsAdd documentation, improve chunking, try hybrid search
Context Precision ↓Retrieve a lot of noiseAdd reranker, reduce top_k, filter metadata
Faithfulness ↓LLM hallucinateImprove prompt ("based on context only"), using a larger model
Answer Relevancy ↓Answer off topicImprove prompt, add few-shot examples

Summary

ConceptsRemember
FaithfulnessIs the answer correct for the context? (no hallucination)
Answer RelevancyIs the answer correct to the point of the question?
Context PrecisionIs Retrieval accurate?
Context RecallIs the retrieval complete?
RAGASAutomated RAG assessment framework
Golden Test SetStandard test suite with ground truth
IterateLow metric → diagnose → improve → re-evaluate

General exercises

  1. ✅ Complete 2 small exercises (1, 2)
  2. Full Evaluation: Create a golden test with 50 questions → evaluate the current pipeline → identify bottlenecks → change a component → re-evaluate → write a comparison report.
  3. CI/CD Evaluation: Automatically run RAGAS every time the pipeline changes. Script: python evaluate.py → output metrics → fail if faithfulness < 0.85.
  4. Human-in-the-loop: Create Streamlit app: display question + answer + context + ground_truth. Users rate 1-5 for each sentence. Compare human rating vs RAGAS scores.

Next article: Deploy RAG to Production — API, Caching & Monitoring — bring RAG from notebook to real product.