Amazon Bedrock Architecture — Foundation Models, Agents, Guardrails, and Knowledge Bases
1. Amazon Bedrock Overview
Amazon Bedrock is a fully managed service that provides access to FMs from multiple providers through a single API, along with tools to customize, deploy, and secure AI applications.
1.1. Key Value Propositions
- Choice: Access FMs from Amazon, Anthropic, Meta, Mistral, Cohere, Stability AI, AI21 Labs
- Customization: Fine-tuning, continued pre-training, RAG (Knowledge Bases)
- Security: Data stays in your AWS account, encrypted, not used to train models
- Serverless: No infrastructure to manage
- Integration: Native AWS service integration (IAM, CloudWatch, CloudTrail)
1.2. Foundation Model Providers on Bedrock
| Provider | Models | Strengths |
|---|---|---|
| Amazon | Titan Text, Titan Embeddings, Titan Image Generator | General purpose, embeddings, image gen |
| Anthropic | Claude 3 Haiku, Sonnet, Opus | Complex reasoning, analysis, vision |
| Meta | Llama 2, Llama 3 | Open-source, customizable |
| Mistral AI | Mistral, Mixtral | Fast, efficient, multilingual |
| Cohere | Command, Embed | Enterprise text, multilingual embeddings |
| Stability AI | Stable Diffusion XL | Image generation |
| AI21 Labs | Jurassic | Text generation, summarization |
2. Bedrock Features Deep Dive
2.1. Amazon Bedrock Agents
Agents allow FMs to perform multi-step tasks by automatically planning, executing actions, and using tools.
User: "Book a flight from Hanoi to Tokyo for next Friday"
Agent workflow:
1. PLAN: Need to search flights, check availability, book
2. ACTION: Call flight search API → find available flights
3. OBSERVE: Found 3 flights, cheapest is $450
4. ACTION: Call booking API → reserve the flight
5. RESPOND: "Booked VN flight HAN→NRT, Dec 20, $450"
Agent Components:
| Component | Purpose |
|---|---|
| Foundation Model | Brain that reasons and plans |
| Instructions | System prompt defining agent's role |
| Action Groups | APIs the agent can call (Lambda functions or OpenAPI schemas) |
| Knowledge Bases | RAG data sources for information retrieval |
| Guardrails | Safety and compliance filters |
Exam tip: "An AI assistant needs to look up order status, check inventory, and process returns" → Bedrock Agent with action groups connected to business APIs.
2.2. Amazon Bedrock Guardrails
Guardrails implement safety controls for AI applications:
| Guardrail Type | What it does | Example |
|---|---|---|
| Content filters | Block harmful content categories | Hate, violence, sexual, insults |
| Denied topics | Block specific topics | "Don't discuss competitor products" |
| Word filters | Block specific words/phrases | Profanity, banned terms |
| PII filters | Detect and redact PII | SSN, credit card numbers, emails |
| Contextual grounding | Check if response is grounded in context | Prevent hallucination in RAG |
Guardrails Flow:
User Input → [Input Guardrails] → FM Processing → [Output Guardrails] → User
Check for: Check for:
- Denied topics - Harmful content
- Harmful input - PII in response
- PII in input - Off-topic responses
- Grounding check
2.3. Model Evaluation
Compare and evaluate FMs for your specific use case:
- Automatic evaluation: BERTScore, accuracy, toxicity metrics
- Human evaluation: Custom criteria rated by human reviewers
- A/B comparison: Side-by-side model comparison
- Custom tasks: Upload your own test dataset
2.4. Bedrock Playgrounds
| Playground | Use Case |
|---|---|
| Text playground | Test text models interactively |
| Chat playground | Test conversational models |
| Image playground | Test image generation models |
3. Amazon PartyRock
PartyRock is a free, no-code playground for Bedrock — allowing anyone to create GenAI apps without needing an AWS account or coding skills.
| Feature | Detail |
|---|---|
| No AWS account needed | Free to use with social login |
| No coding | Drag-and-drop app builder |
| Shareable | Share apps via URL |
| Use case | Learning, prototyping, experimentation |
Exam tip: "A non-technical marketing team wants to experiment with generative AI without an AWS account" → PartyRock.
4. Amazon Q
4.1. Amazon Q Developer
AI coding assistant for developers:
- Code generation: Write code from natural language
- Code explanation: Explain existing code
- Code transformation: Upgrade Java versions, .NET migrations
- Debugging: Identify and fix bugs
- Security scanning: Find vulnerabilities in code
- IDE integration: VS Code, JetBrains, AWS Console
4.2. Amazon Q Business
AI assistant for business users:
- Connect enterprise data: S3, SharePoint, Confluence, Salesforce, etc.
- Q&A on company data: Answers based on connected data sources
- Respects access controls: ACLs from connected systems
- Plugins: Create tickets (Jira), send emails, etc.
4.3. Amazon Q vs Bedrock
| Feature | Amazon Q | Amazon Bedrock |
|---|---|---|
| Target user | End users (devs, business) | Developers building AI apps |
| Customization | Limited (connect data sources) | Full (fine-tune, RAG, agents) |
| Managed | Fully managed assistant | API/SDK access to FMs |
| Use case | Productivity tool | Building custom AI applications |
5. Bedrock Pricing Models
| Pricing Model | How it works | Best For |
|---|---|---|
| On-Demand | Pay per input/output token | Variable, unpredictable workloads |
| Provisioned Throughput | Reserved model units (hourly) | Consistent, production workloads |
| Batch Inference | Submit batch jobs (up to 50% cheaper) | Large-scale, non-real-time processing |
Exam tip: "Cost-optimize a GenAI workload with predictable traffic?" → Provisioned Throughput. "Process thousands of documents overnight?" → Batch Inference.
6. How to Choose the Right FM
Decision Framework:
┌─────────────────────────────────────────────────┐
│ 1. TASK TYPE │
│ Text? Image? Code? Multi-modal? │
├─────────────────────────────────────────────────┤
│ 2. COMPLEXITY │
│ Simple classification → smaller model │
│ Complex reasoning → larger model │
├─────────────────────────────────────────────────┤
│ 3. LATENCY REQUIREMENTS │
│ Real-time → smaller/faster model (Haiku) │
│ Batch processing → larger model (Opus) │
├─────────────────────────────────────────────────┤
│ 4. COST CONSTRAINTS │
│ Budget limited → smaller model │
│ Quality critical → larger model │
├─────────────────────────────────────────────────┤
│ 5. CUSTOMIZATION NEEDS │
│ Fine-tuning needed? Check supported models │
│ LoRA? Check compatibility │
├─────────────────────────────────────────────────┤
│ 6. EVALUATE with Model Evaluation │
│ Test candidates side-by-side │
└─────────────────────────────────────────────────┘
7. Other AWS GenAI Services
| Service | What it does |
|---|---|
| Amazon CodeWhisperer | Now part of Amazon Q Developer (code suggestions) |
| AWS App Studio | Build enterprise apps with natural language |
| Amazon SageMaker JumpStart | Deploy open-source FMs with SageMaker |
| Amazon Comprehend | NLP service (sentiment, entities, topics — pre-built) |
| Amazon Transcribe | Speech-to-text |
| Amazon Polly | Text-to-speech |
| Amazon Translate | Machine translation |
| Amazon Rekognition | Image/video analysis |
| Amazon Textract | Extract text from documents (OCR+) |
8. Practice Questions
Q1: A retail company wants to build an AI assistant that can check inventory, process returns, and answer product questions from their catalog. Which Amazon Bedrock feature should they use?
- A) Bedrock Guardrails
- B) Bedrock Knowledge Bases only
- C) Bedrock Agents with Action Groups and Knowledge Bases ✓
- D) Bedrock Model Evaluation
Explanation: Bedrock Agents can orchestrate multi-step tasks by calling APIs (action groups for inventory/returns) and retrieving information (knowledge bases for product catalog).
Q2: Which Amazon Bedrock feature should be used to prevent a generative AI application from discussing competitor products and to filter out personally identifiable information (PII)?
- A) Bedrock Knowledge Bases
- B) Bedrock Custom Models
- C) Bedrock Guardrails ✓
- D) Bedrock Agents
Explanation: Guardrails provide denied topic filtering (block competitor discussions) and PII detection/redaction. They can be applied to both input and output of FM calls.
Q3: A company wants to process 50,000 customer reviews overnight for sentiment analysis using a foundation model. Which Bedrock pricing model is MOST cost-effective?
- A) On-Demand pricing
- B) Provisioned Throughput
- C) Batch Inference ✓
- D) Free tier
Explanation: Batch Inference is designed for large-scale, non-real-time workloads and offers up to 50% cost savings compared to on-demand pricing. Ideal for overnight processing.