Foundation Model Lifecycle — Pre-training, Fine-tuning, RAG, and Prompt Engineering
Domain 2 Overview
Domain 2 accounts for 24% of the exam — this is the second-largest domain. You need to have a solid understanding of Generative AI, Foundation Models, and how they differ from traditional ML.
1. What is Generative AI?
Generative AI is a branch of AI focused on creating new content (text, images, code, audio, video) based on patterns learned from training data.
Discriminative vs Generative AI
| Aspect | Discriminative AI | Generative AI |
|---|---|---|
| What it does | Classify / predict | Create / generate |
| Output | Label, category, number | New content (text, image, code) |
| Example | "Is this email spam?" → Yes/No | "Write an email about..." → New email |
| Models | Logistic Regression, SVM, CNN classifier | GPT, Claude, Stable Diffusion, DALL-E |
Generative AI Modalities
| Input → Output | Examples | Models |
|---|---|---|
| Text → Text | Chatbot, summarization, translation | GPT-4, Claude, Llama |
| Text → Image | Image generation from description | DALL-E, Stable Diffusion, Titan Image Generator |
| Text → Code | Code generation, debugging | CodeWhisperer, Copilot |
| Text → Audio | Speech synthesis, music generation | Amazon Polly (TTS) |
| Image → Text | Image captioning, visual Q&A | Claude (multi-modal), GPT-4V |
| Audio → Text | Transcription | Amazon Transcribe, Whisper |
2. Foundation Models
A Foundation Model (FM) is a very large AI model that has been pre-trained on massive datasets and can be adapted for many different downstream tasks.
Key Characteristics
- Large-scale pre-training: Trained on billions of data points (text from internet, books, code)
- General-purpose: Can handle multiple tasks without task-specific training
- Adaptable: Can be fine-tuned or prompted for specific use cases
- Expensive to train: Requires massive compute (GPU/TPU clusters)
- Accessible via API: Users don't need to train — use through APIs (Amazon Bedrock)
Foundation Model Lifecycle
┌─────────────────┐ ┌──────────────┐ ┌──────────────┐
│ 1. Pre-training │────→│ 2. Fine- │────→│ 3. Inference │
│ (Massive data, │ │ tuning │ │ (Use model │
│ Billion params,│ │ (Adapt to │ │ via API or │
│ Very expensive)│ │ specific │ │ endpoint) │
│ │ │ domain) │ │ │
└─────────────────┘ └──────────────┘ └──────────────┘
Model Provider You/Org Users
(Anthropic, Meta, (Applications)
Amazon, etc.)
3. Tokenization
Tokenization is the process of splitting text into small units (tokens) that the model can understand.
Input: "Machine learning is amazing!"
Tokens: ["Machine", " learning", " is", " amazing", "!"]
token_1 token_2 token_3 token_4 token_5
OR (subword tokenization):
Tokens: ["Mach", "ine", " learn", "ing", " is", " amaz", "ing", "!"]
Key Concepts for Exam:
- Token ≠ word: A token can be part of a word, a whole word, or punctuation
- Context window: Maximum number of tokens a model can process at once (input + output)
- Token limit: Determines how much text the model can "see" and generate
- Pricing: API calls are typically priced per token (input tokens + output tokens)
Exam tip: Context window size matters. Larger context = can process longer documents. But costs more and may be slower.
4. Model Parameters & Inference Settings
4.1. Model Parameters (Learned during training)
- Parameters = weights and biases in the neural network
- GPT-4: ~1.7 trillion parameters, Claude: undisclosed, Llama 3: 8B/70B/405B
- More parameters → generally more capable, but more expensive
4.2. Inference Parameters (Set by user)
When calling a model, you can adjust inference parameters:
| Parameter | Range | What it controls |
|---|---|---|
| Temperature | 0.0 → 1.0+ | Randomness/creativity. Low = deterministic, focused. High = creative, diverse. |
| Top-p (Nucleus) | 0.0 → 1.0 | Cumulative probability threshold. Lower = more focused vocabulary. |
| Top-k | 1 → ∞ | Number of top tokens to consider. Lower = more predictable. |
| Max tokens | 1 → limit | Maximum length of generated output. |
| Stop sequences | strings | Text that tells model to stop generating. |
Temperature Guide for Exam
Temperature = 0 → Most deterministic (factual Q&A, code, data extraction)
Temperature = 0.3 → Slightly creative (business writing, summaries)
Temperature = 0.7 → Creative (stories, brainstorming, marketing copy)
Temperature = 1.0+ → Very random (poetry, creative writing — may hallucinate more)
Exam tip: "A company needs consistent, accurate answers for customer FAQ" → use low temperature. "A company wants creative marketing slogans" → use high temperature.
5. Hallucination
Hallucination is when a model generates output that is confident but incorrect — fabricating facts, citations, or information that doesn't exist.
Causes:
- Training data gaps or outdated information
- Model doesn't truly "know" facts — it predicts likely next tokens
- Ambiguous or too-open prompts
- High temperature settings
Mitigation Strategies:
| Strategy | How it helps |
|---|---|
| RAG (Retrieval-Augmented Generation) | Ground responses in actual data from knowledge base |
| Lower temperature | Reduce randomness in generation |
| Guardrails | Filter/validate outputs (Amazon Bedrock Guardrails) |
| Better prompts | "Only answer based on provided context" / "Say I don't know if unsure" |
| Fine-tuning | Train model on domain-specific accurate data |
| Human review | Human-in-the-loop validation |
6. Foundation Models on AWS (Amazon Bedrock)
Amazon Bedrock provides access to many Foundation Models from various providers:
| Provider | Models | Strengths |
|---|---|---|
| Anthropic | Claude 3 (Haiku, Sonnet, Opus) | Reasoning, safety, long context |
| Meta | Llama 3 | Open-source, versatile |
| Amazon | Titan (Text, Embeddings, Image) | AWS-native, embeddings for RAG |
| Mistral AI | Mistral, Mixtral | Efficient, fast inference |
| Stability AI | Stable Diffusion | Image generation |
| Cohere | Command, Embed | Enterprise NLP, embeddings |
| AI21 Labs | Jurassic | Text generation |
7. Practice Questions
Q1: What is the PRIMARY advantage of Foundation Models compared to traditional ML models?
- A) They are smaller and faster
- B) They can be adapted to multiple downstream tasks without task-specific training ✓
- C) They never produce incorrect outputs
- D) They don't require any compute resources
Explanation: Foundation Models are pre-trained on massive datasets and can be adapted (via prompting or fine-tuning) for many different tasks. They are large, can hallucinate, and still require compute.
Q2: A company uses a generative AI model and notices it sometimes generates plausible but factually incorrect information. What is this phenomenon called?
- A) Overfitting
- B) Data drift
- C) Hallucination ✓
- D) Bias
Explanation: Hallucination is when a generative AI model produces confident but factually incorrect outputs.
Q3: A developer wants to ensure their generative AI chatbot provides consistent, factual answers with minimal creativity. Which inference parameter should they adjust?
- A) Set max tokens to a very high value
- B) Set temperature close to 0 ✓
- C) Set temperature close to 1
- D) Increase the top-k value
Explanation: Low temperature makes the model more deterministic and focused, reducing creativity and randomness in responses.