Lesson 16: RAG — Retrieval Augmented Generation from A to Z
1. The problem of pure LLM
LLMs like GPT-4 or Claude are extremely powerful models, but they carry three core limitations when applied in practice:
Hallucination: LLMs don't "know" in the sense of looking — they generate text based on statistical probability. In the absence of solid information, they tend to create answers that sound right but are actually wrong.
Knowledge Cutoff (Time limit): Training data has a cutoff date. GPT-4o may not be aware of events occurring after April 2024. This is a serious problem in fast-changing fields such as finance, law, and healthcare.
Private Data (Internal Data): Company documents, internal codebase, email, private database — all are not included in training data. LLM is completely "blind" to this information.
RAG was born to solve all three of the above problems.
2. What is RAG and why is it effective?
Retrieval-Augmented Generation (RAG) is an architecture that combines information retrieval and text generation. Instead of relying entirely on parametric memory (knowledge within weights), RAG provides LLM with non-parametric memory — an external document store that can be continuously updated.
Why is RAG more effective than fine-tuning for many use cases?
| Criteria | RAG | Fine-tuning |
|---|---|---|
| Update data | Almost real-time | Need to retrain |
| Cost | Low (embedding + inference only) | High (GPU hours) |
| Cite source | Natural | Difficult |
| Content Control | Easy (edit corpus) | Complex |
| Suitable for | Q&A, search, enterprise chatbot | Tone, style, domain-specific tasks |
3. RAG Pipeline: Two main Phases
RAG consists of two completely separate phases:
Indexing Phase (Offline — runs once or periodically)
Tài liệu thô → Load → Clean → Chunk → Embed → Lưu vào Vector Store
Query Phase (Online — runs every time the user asks)
User query → Embed query → Tìm top-k chunks → Re-rank → Ghép vào prompt → LLM → Response
4. Document Loading
The first step is to bring documents into the system. LangChain provides more than 100 document loaders:
from langchain_community.document_loaders import (
PyPDFLoader,
Docx2txtLoader,
WebBaseLoader,
BSHTMLLoader,
JSONLoader,
)
# Load PDF
pdf_loader = PyPDFLoader("annual_report.pdf")
pdf_docs = pdf_loader.load() # List[Document]
# Load Word
word_loader = Docx2txtLoader("policy.docx")
word_docs = word_loader.load()
# Web scraping
web_loader = WebBaseLoader(
web_paths=["https://docs.python.org/3/library/functions.html"],
bs_kwargs={"parse_only": SoupStrainer(class_="body")}, # chỉ lấy phần body
)
web_docs = web_loader.load()
# Mỗi Document có: page_content (str) và metadata (dict)
print(pdf_docs[0].metadata)
# {'source': 'annual_report.pdf', 'page': 0}
5. Text Chunking Strategies
Chunking is the step that has the biggest impact on RAG quality. Chunks that are too small lose context, chunks that are too large cause noise.
Fixed-size Chunking
from langchain_text_splitters import CharacterTextSplitter
splitter = CharacterTextSplitter(
chunk_size=1000, # ký tự mỗi chunk
chunk_overlap=200, # overlap để giữ ngữ cảnh
separator="\n\n",
)
chunks = splitter.split_documents(docs)
Recursive Character Splitter (recommended for regular text)
from langchain_text_splitters import RecursiveCharacterTextSplitter
# Thử split theo: ["\n\n", "\n", " ", ""] theo thứ tự
splitter = RecursiveCharacterTextSplitter(
chunk_size=1000,
chunk_overlap=200,
length_function=len,
)
chunks = splitter.split_documents(docs)
Semantic Chunking (smartest, most expensive)
from langchain_experimental.text_splitter import SemanticChunker
from langchain_openai import OpenAIEmbeddings
# Split dựa trên sự thay đổi ngữ nghĩa
semantic_splitter = SemanticChunker(
embeddings=OpenAIEmbeddings(),
breakpoint_threshold_type="percentile",
breakpoint_threshold_amount=95,
)
chunks = semantic_splitter.split_documents(docs)
6. Embedding Models
Embedding converts text into an arithmetic vector that captures the semantics.
from langchain_openai import OpenAIEmbeddings
from langchain_huggingface import HuggingFaceEmbeddings
# OpenAI — chất lượng cao, có phí
openai_embeddings = OpenAIEmbeddings(
model="text-embedding-3-small", # 1536 dims, rẻ hơn large
# model="text-embedding-3-large", # 3072 dims, tốt hơn
)
# Sentence Transformers — miễn phí, chạy local
local_embeddings = HuggingFaceEmbeddings(
model_name="BAAI/bge-m3", # đa ngôn ngữ, hỗ trợ tiếng Việt tốt
# model_name="intfloat/multilingual-e5-large",
model_kwargs={"device": "cpu"},
encode_kwargs={"normalize_embeddings": True},
)
# Test embedding
vector = openai_embeddings.embed_query("RAG là gì?")
print(f"Dimension: {len(vector)}") # 1536
7. Vector Stores
Vector store is a specialized database for storing and searching embeddings.
from langchain_chroma import Chroma
from langchain_openai import OpenAIEmbeddings
embeddings = OpenAIEmbeddings(model="text-embedding-3-small")
# Tạo vector store từ documents
vectorstore = Chroma.from_documents(
documents=chunks,
embedding=embeddings,
persist_directory="./chroma_db", # lưu xuống disk
collection_name="my_rag_collection",
)
# Load lại từ disk
vectorstore = Chroma(
persist_directory="./chroma_db",
embedding_function=embeddings,
collection_name="my_rag_collection",
)
8. Retrieval: Cosine Similarity and MMR
# Similarity search thuần (top-4 chunks giống nhất)
retriever_basic = vectorstore.as_retriever(
search_type="similarity",
search_kwargs={"k": 4},
)
# MMR — Max Marginal Relevance: cân bằng relevance + diversity
# Tránh trả về 4 chunks gần giống nhau
retriever_mmr = vectorstore.as_retriever(
search_type="mmr",
search_kwargs={
"k": 4, # số chunks trả về
"fetch_k": 20, # fetch 20, rồi chọn 4 đa dạng nhất
"lambda_mult": 0.5, # 0=max diversity, 1=max relevance
},
)
# Score threshold — chỉ lấy chunks đủ liên quan
retriever_threshold = vectorstore.as_retriever(
search_type="similarity_score_threshold",
search_kwargs={"score_threshold": 0.7, "k": 6},
)
9. Re-ranking with Cross-Encoder
Bi-encoder (used for embed) is fast but less accurate. Cross-encoder compares the query with each document directly — slower but much more accurate. Combining both is best practice.
from langchain.retrievers import ContextualCompressionRetriever
from langchain.retrievers.document_compressors import CrossEncoderReranker
from langchain_community.cross_encoders import HuggingFaceCrossEncoder
# Cross-encoder model cho re-ranking
reranker_model = HuggingFaceCrossEncoder(
model_name="BAAI/bge-reranker-v2-m3"
)
compressor = CrossEncoderReranker(model=reranker_model, top_n=3)
# Pipeline: lấy 10 chunks, re-rank, giữ top 3
reranking_retriever = ContextualCompressionRetriever(
base_compressor=compressor,
base_retriever=vectorstore.as_retriever(search_kwargs={"k": 10}),
)
10. Generation: Insert Context into Prompt
from langchain_openai import ChatOpenAI
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser
from langchain_core.runnables import RunnablePassthrough
llm = ChatOpenAI(model="gpt-4o-mini", temperature=0)
prompt = ChatPromptTemplate.from_template("""
Bạn là trợ lý AI hữu ích. Dựa vào ngữ cảnh dưới đây để trả lời câu hỏi.
Nếu ngữ cảnh không đủ thông tin, hãy nói rõ bạn không biết.
Đừng bịa đặt thông tin không có trong ngữ cảnh.
Ngữ cảnh:
{context}
Câu hỏi: {question}
Trả lời:""")
def format_docs(docs):
return "\n\n---\n\n".join(
f"[Nguồn: {doc.metadata.get('source', 'N/A')}]\n{doc.page_content}"
for doc in docs
)
# LCEL chain
rag_chain = (
{"context": retriever_mmr | format_docs, "question": RunnablePassthrough()}
| prompt
| llm
| StrOutputParser()
)
response = rag_chain.invoke("RAG pipeline hoạt động như thế nào?")
print(response)
11. Advanced RAG Techniques
HyDE — Hypothetical Document Embeddings
Instead of embed query directly (short, little information), use LLM to generate hypothetical document and then embed that document:
from langchain.retrievers import HyDERetriever
from langchain_openai import ChatOpenAI
hyde_retriever = HyDERetriever.from_llm(
retriever=vectorstore.as_retriever(),
llm=ChatOpenAI(model="gpt-4o-mini"),
prompt_key="web_search",
)
docs = hyde_retriever.invoke("Tại sao RAG tốt hơn fine-tuning?")
Corrective RAG (CRAG)
After retrieval, use a small LLM to evaluate the relevance of each chunk. If all chunks are less relevant, fallback to web search:
from langgraph.graph import StateGraph, END
from typing import TypedDict, List
class GraphState(TypedDict):
question: str
documents: List
generation: str
web_search_needed: bool
def grade_documents(state):
"""Dùng LLM judge để chấm từng document"""
grader_prompt = "Tài liệu này có liên quan đến câu hỏi không? Trả lời 'yes' hoặc 'no'."
# ... implement grading logic
return state
# Build CRAG graph với LangGraph
workflow = StateGraph(GraphState)
# Thêm các nodes: retrieve → grade → (web_search nếu cần) → generate
RAPTOR — Recursive Abstractive Processing
Build a hierarchical tree: cluster documents → summarize each cluster → cluster summaries → summarize again. Allows you to answer both detailed and general questions.
12. Full Code: RAG Pipeline Complete
# pip install langchain langchain-openai langchain-chroma chromadb pypdf
import os
from langchain_community.document_loaders import PyPDFDirectoryLoader
from langchain_text_splitters import RecursiveCharacterTextSplitter
from langchain_openai import OpenAIEmbeddings, ChatOpenAI
from langchain_chroma import Chroma
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser
from langchain_core.runnables import RunnablePassthrough
os.environ["OPENAI_API_KEY"] = "your-api-key"
# ── INDEXING PHASE ──────────────────────────────────────────────
# 1. Load tài liệu
loader = PyPDFDirectoryLoader("./documents/")
raw_docs = loader.load()
print(f"Loaded {len(raw_docs)} pages")
# 2. Chunking
splitter = RecursiveCharacterTextSplitter(
chunk_size=1000,
chunk_overlap=200,
add_start_index=True, # ghi lại vị trí trong doc gốc
)
chunks = splitter.split_documents(raw_docs)
print(f"Created {len(chunks)} chunks")
# 3. Embedding + lưu vào ChromaDB
embeddings = OpenAIEmbeddings(model="text-embedding-3-small")
vectorstore = Chroma.from_documents(
documents=chunks,
embedding=embeddings,
persist_directory="./chroma_db",
)
print("Vectorstore ready!")
# ── QUERY PHASE ─────────────────────────────────────────────────
# 4. Retriever với MMR
retriever = vectorstore.as_retriever(
search_type="mmr",
search_kwargs={"k": 5, "fetch_k": 20},
)
# 5. Prompt template
prompt = ChatPromptTemplate.from_messages([
("system", """Bạn là chuyên gia phân tích tài liệu.
Chỉ trả lời dựa trên ngữ cảnh được cung cấp.
Luôn trích dẫn nguồn tài liệu (tên file và trang).
Ngữ cảnh:
{context}"""),
("human", "{question}"),
])
# 6. LLM
llm = ChatOpenAI(model="gpt-4o", temperature=0)
def format_docs_with_sources(docs):
formatted = []
for doc in docs:
src = doc.metadata.get("source", "unknown")
page = doc.metadata.get("page", "?")
formatted.append(f"[{src}, trang {page}]\n{doc.page_content}")
return "\n\n---\n\n".join(formatted)
# 7. RAG Chain
rag_chain = (
{
"context": retriever | format_docs_with_sources,
"question": RunnablePassthrough(),
}
| prompt
| llm
| StrOutputParser()
)
# 8. Sử dụng
if __name__ == "__main__":
questions = [
"Chính sách bảo hành sản phẩm là bao lâu?",
"Quy trình hoàn tiền như thế nào?",
]
for q in questions:
print(f"\nQ: {q}")
print(f"A: {rag_chain.invoke(q)}")
print("-" * 60)
Summary
RAG is an indispensable technique for building practical LLM applications. The standard pipeline includes: Load → Chunk → Embed → Store (offline) and Retrieve → Re-rank → Generate (online). With advanced techniques like HyDE, RAPTOR, and Corrective RAG, you can achieve production-grade accuracy. The next article will dive deeper into Vector Databases — the core component of every RAG system.