RAG Architecture — Indexing Phase và Query Phase với Amazon Bedrock Knowledge Bases
1. What is RAG?
Retrieval-Augmented Generation (RAG) là kỹ thuật kết hợp FM với external knowledge sources để trả lời chính xác hơn, giảm hallucination, và cập nhật thông tin mà model chưa biết.
1.1. Why RAG?
| Problem | RAG Solution |
|---|---|
| Knowledge cutoff date | Retrieve latest documents |
| Hallucination | Ground responses in real data |
| No domain knowledge | Add company-specific documents |
| Generic answers | Cite specific sources |
| Privacy — can't send data to FM training | Keep data in your own vector DB |
1.2. RAG Architecture
RAG Pipeline:
┌─────────────────────────────────────────────────────────────┐
│ INDEXING (Done once / periodically) │
│ │
│ Documents → Chunking → Embedding Model → Vector Database │
│ (PDF, web, (split (Amazon Titan (OpenSearch, │
│ S3, etc.) text) Embeddings) Aurora pgvector)│
└─────────────────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────────────────┐
│ RETRIEVAL & GENERATION (Per query) │
│ │
│ User Query → Embed Query → Search Vector DB → Top-K docs │
│ │
│ Augmented Prompt = System Prompt + Retrieved Docs + Query │
│ │
│ Augmented Prompt → Foundation Model → Answer with sources │
└─────────────────────────────────────────────────────────────┘
2. Chunking Strategies
Trước khi tạo embeddings, documents phải được chia nhỏ (chunked) thành các đoạn phù hợp.
| Strategy | Description | Best For |
|---|---|---|
| Fixed-size | Split every N characters/tokens | Simple, uniform documents |
| Sentence-based | Split at sentence boundaries | Narrative text |
| Paragraph-based | Split at paragraph breaks | Well-structured documents |
| Semantic | Split based on topic changes | Complex documents |
| Hierarchical | Parent-child chunk relationships | Long documents with sections |
Chunk Size Trade-offs:
Small chunks (100-200 tokens):
✓ More precise retrieval
✗ May lose context
✗ More chunks to search
Large chunks (500-1000 tokens):
✓ More context preserved
✗ May include irrelevant info
✗ Fewer chunks, less granular
Overlap (e.g., 20% between chunks):
✓ Prevents information loss at boundaries
✗ Increases storage and compute
Exam tip: "How to improve RAG retrieval accuracy?" → Adjust chunk size, add overlap, use semantic chunking, improve embedding model.
3. Embeddings for RAG
3.1. AWS Embedding Models
| Model | Modality | Dimensions | Use Case |
|---|---|---|---|
| Amazon Titan Text Embeddings V2 | Text | 256/512/1024 | Semantic search, RAG |
| Amazon Titan Multimodal Embeddings | Text + Image | 256/384/1024 | Cross-modal search |
| Cohere Embed | Text | 1024 | Multilingual search |
3.2. Vector Databases on AWS
| Service | Type | Key Feature |
|---|---|---|
| Amazon OpenSearch Serverless | Managed | Vector search collection type, serverless |
| Amazon Aurora PostgreSQL | RDB + Vector | pgvector extension |
| Amazon Neptune | Graph + Vector | Knowledge graphs with vector search |
| Amazon DocumentDB | Document + Vector | MongoDB-compatible with vector search |
| Amazon MemoryDB | In-memory + Vector | Redis-compatible, ultra-low latency |
| Pinecone (3rd party) | Dedicated vector DB | Popular, integrates with Bedrock |
4. Amazon Bedrock Knowledge Bases
Bedrock Knowledge Bases là fully managed RAG solution. AWS handles chunking, embedding, indexing, retrieval — bạn chỉ cần point to data sources.
4.1. How It Works
Setup:
┌───────────┐ ┌───────────────┐ ┌─────────────────┐
│ S3 Bucket │────→│ Bedrock │────→│ Vector Store │
│ (docs) │ │ Knowledge Base│ │ (OpenSearch/ │
│ │ │ (auto-chunk, │ │ Aurora/Pinecone) │
│ │ │ auto-embed) │ │ │
└───────────┘ └───────────────┘ └─────────────────┘
Query:
┌───────────┐ ┌───────────────┐ ┌─────────────────┐
│ User │────→│ Knowledge Base│────→│ FM (Claude, │
│ "What is │ │ retrieves │ │ Titan, etc.) │
│ the..." │ │ relevant docs │ │ generates answer │
└───────────┘ └───────────────┘ └─────────────────┘
4.2. Supported Data Sources
- Amazon S3: PDF, TXT, MD, HTML, DOC, CSV
- Web Crawler: Crawl websites automatically
- Confluence: Atlassian Confluence pages
- SharePoint: Microsoft SharePoint documents
- Salesforce: Salesforce knowledge articles
4.3. Key Features
| Feature | Benefit |
|---|---|
| Managed chunking | Auto-splits documents (fixed, semantic, hierarchical) |
| Auto-sync | Periodically re-indexes when data changes |
| Source attribution | Returns source documents with answers |
| Metadata filtering | Filter chunks by custom metadata fields |
| Hybrid search | Combines semantic + keyword search |
| Guardrails integration | Apply safety filters to RAG responses |
Exam tip: "A company wants to build a chatbot that answers questions from internal documents stored in S3, with minimal custom code" → Amazon Bedrock Knowledge Bases.
5. RAG vs Fine-tuning
| Factor | RAG | Fine-tuning |
|---|---|---|
| Purpose | Access external/current data | Teach new skills/domain patterns |
| Data freshness | Always up-to-date | Fixed at training time |
| Training required? | No model training | Yes, needs labeled data + compute |
| Cost | Vector DB + retrieval costs | Training compute + storage |
| Hallucination | Reduced (grounded in data) | May still hallucinate |
| Latency | Slightly higher (retrieval step) | Same as base model |
| Best for | Q&A, search, knowledge bases | Style, tone, domain-specific patterns |
| Data privacy | Data stays in your vector DB | Data used in training process |
Decision Matrix:
"Need to answer from company docs?" → RAG
"Need real-time/latest information?" → RAG
"Need to change model's writing style?" → Fine-tuning
"Need model to follow specific format?" → Try prompting first → then fine-tuning
"Need domain-specific terminology?" → RAG (if in docs) or Fine-tuning (if patterns)
"Minimum effort/cost?" → RAG > Prompt Engineering > Fine-tuning
6. Evaluating RAG Quality
| Metric | What it measures |
|---|---|
| Faithfulness | Is the answer grounded in retrieved docs? (no hallucination) |
| Relevance | Are retrieved documents relevant to the query? |
| Answer correctness | Is the final answer factually correct? |
| Context precision | What % of retrieved chunks are actually relevant? |
| Context recall | Did we retrieve all relevant chunks? |
7. Practice Questions
Q1: A healthcare company wants an AI assistant that answers questions from their latest medical research papers stored in Amazon S3. The information changes weekly. Which approach is MOST suitable?
- A) Fine-tune a foundation model on the papers
- B) Use RAG with Amazon Bedrock Knowledge Bases ✓
- C) Use zero-shot prompting with a large context window
- D) Pre-train a custom model on medical data
Explanation: RAG with Bedrock Knowledge Bases is ideal — it automatically indexes S3 documents, retrieves relevant information per query, and keeps responses current without retraining. Weekly updates are handled by auto-sync.
Q2: What is the PRIMARY purpose of chunking documents in a RAG pipeline?
- A) To reduce storage costs
- B) To split documents into manageable pieces for embedding and retrieval ✓
- C) To encrypt sensitive data
- D) To convert documents to a different file format
Explanation: Chunking splits large documents into smaller, semantically meaningful pieces that can be individually embedded and retrieved. This enables precise retrieval of relevant information rather than processing entire documents.
Q3: A company built a RAG application, but it sometimes returns answers not supported by the retrieved documents. Which metric should they focus on improving?
- A) Context recall
- B) Answer length
- C) Faithfulness ✓
- D) Response latency
Explanation: Faithfulness measures whether the generated answer is grounded in the retrieved documents. Low faithfulness means the model is generating information beyond what the retrieved context supports (hallucination in RAG context).