Introduction
Agents are stronger when they have their own knowledge base. RAG (Retrieval-Augmented Generation) allows agents to search for information from documents, databases, or knowledge bases before responding — reducing hallucination and increasing accuracy.
1. RAG Pipeline Overview
Documents → Chunking → Embedding → Vector Store → Retrieval → LLM → Answer
2. Implementation with ChromaDB
import chromadb
from openai import OpenAI
client = OpenAI()
chroma = chromadb.PersistentClient(path="./agent_knowledge")
collection = chroma.get_or_create_collection("docs")
def add_documents(texts, metadatas=None):
embeddings = get_embeddings(texts)
collection.add(
documents=texts,
embeddings=embeddings,
ids=[f"doc_{i}" for i in range(len(texts))],
metadatas=metadatas,
)
def search_knowledge(query, n_results=5):
query_embedding = get_embeddings([query])[0]
results = collection.query(
query_embeddings=[query_embedding],
n_results=n_results,
)
return results["documents"][0]
3. RAG as Agent Tool
@registry.register("search_knowledge", "Tìm kiếm trong knowledge base", {...})
def search_knowledge_tool(query: str) -> str:
results = search_knowledge(query, n_results=3)
return "\n---\n".join(results)
Summary
- RAG = allows agents to access private knowledge base
- Chunking strategy affects retrieval quality
- ChromaDB and Qdrant are the two most popular DB vectors
- Hybrid search (semantic + keyword) gives the best results
Exercises
- Build RAG pipeline with ChromaDB for 100+ documents
- Compare chunking strategies: fixed-size vs recursive vs semantic
- Implement hybrid search (semantic + BM25)
- Integrate RAG tool into SimpleAgent from lesson 6