Amazon Bedrock Architecture — Foundation Models, Agents, Guardrails và Knowledge Bases
1. Amazon Bedrock Overview
Amazon Bedrock là fully managed service cung cấp access đến FMs từ nhiều providers qua single API, kèm theo tools để customize, deploy, và secure AI applications.
1.1. Key Value Propositions
- Choice: Access FMs từ Amazon, Anthropic, Meta, Mistral, Cohere, Stability AI, AI21 Labs
- Customization: Fine-tuning, continued pre-training, RAG (Knowledge Bases)
- Security: Data stays in your AWS account, encrypted, not used to train models
- Serverless: No infrastructure to manage
- Integration: Native AWS service integration (IAM, CloudWatch, CloudTrail)
1.2. Foundation Model Providers on Bedrock
| Provider | Models | Strengths |
|---|---|---|
| Amazon | Titan Text, Titan Embeddings, Titan Image Generator | General purpose, embeddings, image gen |
| Anthropic | Claude 3 Haiku, Sonnet, Opus | Complex reasoning, analysis, vision |
| Meta | Llama 2, Llama 3 | Open-source, customizable |
| Mistral AI | Mistral, Mixtral | Fast, efficient, multilingual |
| Cohere | Command, Embed | Enterprise text, multilingual embeddings |
| Stability AI | Stable Diffusion XL | Image generation |
| AI21 Labs | Jurassic | Text generation, summarization |
2. Bedrock Features Deep Dive
2.1. Amazon Bedrock Agents
Agents cho phép FMs thực hiện multi-step tasks bằng cách tự động plan, execute actions, và use tools.
User: "Book a flight from Hanoi to Tokyo for next Friday"
Agent workflow:
1. PLAN: Need to search flights, check availability, book
2. ACTION: Call flight search API → find available flights
3. OBSERVE: Found 3 flights, cheapest is $450
4. ACTION: Call booking API → reserve the flight
5. RESPOND: "Booked VN flight HAN→NRT, Dec 20, $450"
Agent Components:
| Component | Purpose |
|---|---|
| Foundation Model | Brain that reasons and plans |
| Instructions | System prompt defining agent's role |
| Action Groups | APIs the agent can call (Lambda functions or OpenAPI schemas) |
| Knowledge Bases | RAG data sources for information retrieval |
| Guardrails | Safety and compliance filters |
Exam tip: "An AI assistant needs to look up order status, check inventory, and process returns" → Bedrock Agent with action groups connected to business APIs.
2.2. Amazon Bedrock Guardrails
Guardrails implement safety controls for AI applications:
| Guardrail Type | What it does | Example |
|---|---|---|
| Content filters | Block harmful content categories | Hate, violence, sexual, insults |
| Denied topics | Block specific topics | "Don't discuss competitor products" |
| Word filters | Block specific words/phrases | Profanity, banned terms |
| PII filters | Detect and redact PII | SSN, credit card numbers, emails |
| Contextual grounding | Check if response is grounded in context | Prevent hallucination in RAG |
Guardrails Flow:
User Input → [Input Guardrails] → FM Processing → [Output Guardrails] → User
Check for: Check for:
- Denied topics - Harmful content
- Harmful input - PII in response
- PII in input - Off-topic responses
- Grounding check
2.3. Model Evaluation
Compare and evaluate FMs for your specific use case:
- Automatic evaluation: BERTScore, accuracy, toxicity metrics
- Human evaluation: Custom criteria rated by human reviewers
- A/B comparison: Side-by-side model comparison
- Custom tasks: Upload your own test dataset
2.4. Bedrock Playgrounds
| Playground | Use Case |
|---|---|
| Text playground | Test text models interactively |
| Chat playground | Test conversational models |
| Image playground | Test image generation models |
3. Amazon PartyRock
PartyRock là free, no-code playground cho Bedrock — cho phép bất kỳ ai tạo GenAI apps mà không cần AWS account hay coding skills.
| Feature | Detail |
|---|---|
| No AWS account needed | Free to use with social login |
| No coding | Drag-and-drop app builder |
| Shareable | Share apps via URL |
| Use case | Learning, prototyping, experimentation |
Exam tip: "A non-technical marketing team wants to experiment with generative AI without an AWS account" → PartyRock.
4. Amazon Q
4.1. Amazon Q Developer
AI coding assistant cho developers:
- Code generation: Write code from natural language
- Code explanation: Explain existing code
- Code transformation: Upgrade Java versions, .NET migrations
- Debugging: Identify and fix bugs
- Security scanning: Find vulnerabilities in code
- IDE integration: VS Code, JetBrains, AWS Console
4.2. Amazon Q Business
AI assistant for business users:
- Connect enterprise data: S3, SharePoint, Confluence, Salesforce, etc.
- Q&A on company data: Answers based on connected data sources
- Respects access controls: ACLs from connected systems
- Plugins: Create tickets (Jira), send emails, etc.
4.3. Amazon Q vs Bedrock
| Feature | Amazon Q | Amazon Bedrock |
|---|---|---|
| Target user | End users (devs, business) | Developers building AI apps |
| Customization | Limited (connect data sources) | Full (fine-tune, RAG, agents) |
| Managed | Fully managed assistant | API/SDK access to FMs |
| Use case | Productivity tool | Building custom AI applications |
5. Bedrock Pricing Models
| Pricing Model | How it works | Best For |
|---|---|---|
| On-Demand | Pay per input/output token | Variable, unpredictable workloads |
| Provisioned Throughput | Reserved model units (hourly) | Consistent, production workloads |
| Batch Inference | Submit batch jobs (up to 50% cheaper) | Large-scale, non-real-time processing |
Exam tip: "Cost-optimize a GenAI workload with predictable traffic?" → Provisioned Throughput. "Process thousands of documents overnight?" → Batch Inference.
6. How to Choose the Right FM
Decision Framework:
┌─────────────────────────────────────────────────┐
│ 1. TASK TYPE │
│ Text? Image? Code? Multi-modal? │
├─────────────────────────────────────────────────┤
│ 2. COMPLEXITY │
│ Simple classification → smaller model │
│ Complex reasoning → larger model │
├─────────────────────────────────────────────────┤
│ 3. LATENCY REQUIREMENTS │
│ Real-time → smaller/faster model (Haiku) │
│ Batch processing → larger model (Opus) │
├─────────────────────────────────────────────────┤
│ 4. COST CONSTRAINTS │
│ Budget limited → smaller model │
│ Quality critical → larger model │
├─────────────────────────────────────────────────┤
│ 5. CUSTOMIZATION NEEDS │
│ Fine-tuning needed? Check supported models │
│ LoRA? Check compatibility │
├─────────────────────────────────────────────────┤
│ 6. EVALUATE with Model Evaluation │
│ Test candidates side-by-side │
└─────────────────────────────────────────────────┘
7. Other AWS GenAI Services
| Service | What it does |
|---|---|
| Amazon CodeWhisperer | Now part of Amazon Q Developer (code suggestions) |
| AWS App Studio | Build enterprise apps with natural language |
| Amazon SageMaker JumpStart | Deploy open-source FMs with SageMaker |
| Amazon Comprehend | NLP service (sentiment, entities, topics — pre-built) |
| Amazon Transcribe | Speech-to-text |
| Amazon Polly | Text-to-speech |
| Amazon Translate | Machine translation |
| Amazon Rekognition | Image/video analysis |
| Amazon Textract | Extract text from documents (OCR+) |
8. Practice Questions
Q1: A retail company wants to build an AI assistant that can check inventory, process returns, and answer product questions from their catalog. Which Amazon Bedrock feature should they use?
- A) Bedrock Guardrails
- B) Bedrock Knowledge Bases only
- C) Bedrock Agents with Action Groups and Knowledge Bases ✓
- D) Bedrock Model Evaluation
Explanation: Bedrock Agents can orchestrate multi-step tasks by calling APIs (action groups for inventory/returns) and retrieving information (knowledge bases for product catalog).
Q2: Which Amazon Bedrock feature should be used to prevent a generative AI application from discussing competitor products and to filter out personally identifiable information (PII)?
- A) Bedrock Knowledge Bases
- B) Bedrock Custom Models
- C) Bedrock Guardrails ✓
- D) Bedrock Agents
Explanation: Guardrails provide denied topic filtering (block competitor discussions) and PII detection/redaction. They can be applied to both input and output of FM calls.
Q3: A company wants to process 50,000 customer reviews overnight for sentiment analysis using a foundation model. Which Bedrock pricing model is MOST cost-effective?
- A) On-Demand pricing
- B) Provisioned Throughput
- C) Batch Inference ✓
- D) Free tier
Explanation: Batch Inference is designed for large-scale, non-real-time workloads and offers up to 50% cost savings compared to on-demand pricing. Ideal for overnight processing.