Chuyển đến nội dung chính

Bài 6: RAG, Vector Databases & Bedrock Knowledge Bases

Retrieval-Augmented Generation (RAG) architecture. Vector databases, embeddings, chunking strategies. Amazon Bedrock Knowledge Bases. So sánh RAG vs Fine-tuning.

RAG Architecture

RAG Architecture — Indexing Phase và Query Phase với Amazon Bedrock Knowledge Bases

1. What is RAG?

Retrieval-Augmented Generation (RAG) là kỹ thuật kết hợp FM với external knowledge sources để trả lời chính xác hơn, giảm hallucination, và cập nhật thông tin mà model chưa biết.

1.1. Why RAG?

ProblemRAG Solution
Knowledge cutoff dateRetrieve latest documents
HallucinationGround responses in real data
No domain knowledgeAdd company-specific documents
Generic answersCite specific sources
Privacy — can't send data to FM trainingKeep data in your own vector DB

1.2. RAG Architecture

RAG Pipeline:

┌─────────────────────────────────────────────────────────────┐
│  INDEXING (Done once / periodically)                        │
│                                                             │
│  Documents → Chunking → Embedding Model → Vector Database   │
│  (PDF, web,    (split     (Amazon Titan     (OpenSearch,    │
│   S3, etc.)    text)       Embeddings)       Aurora pgvector)│
└─────────────────────────────────────────────────────────────┘

┌─────────────────────────────────────────────────────────────┐
│  RETRIEVAL & GENERATION (Per query)                         │
│                                                             │
│  User Query → Embed Query → Search Vector DB → Top-K docs   │
│                                                             │
│  Augmented Prompt = System Prompt + Retrieved Docs + Query  │
│                                                             │
│  Augmented Prompt → Foundation Model → Answer with sources  │
└─────────────────────────────────────────────────────────────┘

2. Chunking Strategies

Trước khi tạo embeddings, documents phải được chia nhỏ (chunked) thành các đoạn phù hợp.

StrategyDescriptionBest For
Fixed-sizeSplit every N characters/tokensSimple, uniform documents
Sentence-basedSplit at sentence boundariesNarrative text
Paragraph-basedSplit at paragraph breaksWell-structured documents
SemanticSplit based on topic changesComplex documents
HierarchicalParent-child chunk relationshipsLong documents with sections

Chunk Size Trade-offs:

Small chunks (100-200 tokens):
  ✓ More precise retrieval
  ✗ May lose context
  ✗ More chunks to search

Large chunks (500-1000 tokens):
  ✓ More context preserved
  ✗ May include irrelevant info
  ✗ Fewer chunks, less granular

Overlap (e.g., 20% between chunks):
  ✓ Prevents information loss at boundaries
  ✗ Increases storage and compute

Exam tip: "How to improve RAG retrieval accuracy?" → Adjust chunk size, add overlap, use semantic chunking, improve embedding model.

3. Embeddings for RAG

3.1. AWS Embedding Models

ModelModalityDimensionsUse Case
Amazon Titan Text Embeddings V2Text256/512/1024Semantic search, RAG
Amazon Titan Multimodal EmbeddingsText + Image256/384/1024Cross-modal search
Cohere EmbedText1024Multilingual search

3.2. Vector Databases on AWS

ServiceTypeKey Feature
Amazon OpenSearch ServerlessManagedVector search collection type, serverless
Amazon Aurora PostgreSQLRDB + Vectorpgvector extension
Amazon NeptuneGraph + VectorKnowledge graphs with vector search
Amazon DocumentDBDocument + VectorMongoDB-compatible with vector search
Amazon MemoryDBIn-memory + VectorRedis-compatible, ultra-low latency
Pinecone (3rd party)Dedicated vector DBPopular, integrates with Bedrock

4. Amazon Bedrock Knowledge Bases

Bedrock Knowledge Bases là fully managed RAG solution. AWS handles chunking, embedding, indexing, retrieval — bạn chỉ cần point to data sources.

4.1. How It Works

Setup:
┌───────────┐     ┌───────────────┐     ┌─────────────────┐
│ S3 Bucket │────→│ Bedrock       │────→│ Vector Store     │
│ (docs)    │     │ Knowledge Base│     │ (OpenSearch/     │
│           │     │ (auto-chunk,  │     │  Aurora/Pinecone) │
│           │     │  auto-embed)  │     │                  │
└───────────┘     └───────────────┘     └─────────────────┘

Query:
┌───────────┐     ┌───────────────┐     ┌─────────────────┐
│ User      │────→│ Knowledge Base│────→│ FM (Claude,      │
│ "What is  │     │ retrieves     │     │  Titan, etc.)    │
│  the..."  │     │ relevant docs │     │ generates answer │
└───────────┘     └───────────────┘     └─────────────────┘

4.2. Supported Data Sources

  • Amazon S3: PDF, TXT, MD, HTML, DOC, CSV
  • Web Crawler: Crawl websites automatically
  • Confluence: Atlassian Confluence pages
  • SharePoint: Microsoft SharePoint documents
  • Salesforce: Salesforce knowledge articles

4.3. Key Features

FeatureBenefit
Managed chunkingAuto-splits documents (fixed, semantic, hierarchical)
Auto-syncPeriodically re-indexes when data changes
Source attributionReturns source documents with answers
Metadata filteringFilter chunks by custom metadata fields
Hybrid searchCombines semantic + keyword search
Guardrails integrationApply safety filters to RAG responses

Exam tip: "A company wants to build a chatbot that answers questions from internal documents stored in S3, with minimal custom code" → Amazon Bedrock Knowledge Bases.

5. RAG vs Fine-tuning

FactorRAGFine-tuning
PurposeAccess external/current dataTeach new skills/domain patterns
Data freshnessAlways up-to-dateFixed at training time
Training required?No model trainingYes, needs labeled data + compute
CostVector DB + retrieval costsTraining compute + storage
HallucinationReduced (grounded in data)May still hallucinate
LatencySlightly higher (retrieval step)Same as base model
Best forQ&A, search, knowledge basesStyle, tone, domain-specific patterns
Data privacyData stays in your vector DBData used in training process

Decision Matrix:

"Need to answer from company docs?"       → RAG
"Need real-time/latest information?"       → RAG
"Need to change model's writing style?"    → Fine-tuning
"Need model to follow specific format?"    → Try prompting first → then fine-tuning
"Need domain-specific terminology?"        → RAG (if in docs) or Fine-tuning (if patterns)
"Minimum effort/cost?"                     → RAG > Prompt Engineering > Fine-tuning

6. Evaluating RAG Quality

MetricWhat it measures
FaithfulnessIs the answer grounded in retrieved docs? (no hallucination)
RelevanceAre retrieved documents relevant to the query?
Answer correctnessIs the final answer factually correct?
Context precisionWhat % of retrieved chunks are actually relevant?
Context recallDid we retrieve all relevant chunks?

7. Practice Questions

Q1: A healthcare company wants an AI assistant that answers questions from their latest medical research papers stored in Amazon S3. The information changes weekly. Which approach is MOST suitable?

  • A) Fine-tune a foundation model on the papers
  • B) Use RAG with Amazon Bedrock Knowledge Bases ✓
  • C) Use zero-shot prompting with a large context window
  • D) Pre-train a custom model on medical data

Explanation: RAG with Bedrock Knowledge Bases is ideal — it automatically indexes S3 documents, retrieves relevant information per query, and keeps responses current without retraining. Weekly updates are handled by auto-sync.

Q2: What is the PRIMARY purpose of chunking documents in a RAG pipeline?

  • A) To reduce storage costs
  • B) To split documents into manageable pieces for embedding and retrieval ✓
  • C) To encrypt sensitive data
  • D) To convert documents to a different file format

Explanation: Chunking splits large documents into smaller, semantically meaningful pieces that can be individually embedded and retrieved. This enables precise retrieval of relevant information rather than processing entire documents.

Q3: A company built a RAG application, but it sometimes returns answers not supported by the retrieved documents. Which metric should they focus on improving?

  • A) Context recall
  • B) Answer length
  • C) Faithfulness ✓
  • D) Response latency

Explanation: Faithfulness measures whether the generated answer is grounded in the retrieved documents. Low faithfulness means the model is generating information beyond what the retrieved context supports (hallucination in RAG context).