Foundation Model Lifecycle — Pre-training, Fine-tuning, RAG và Prompt Engineering
Tổng quan Domain 2
Domain 2 chiếm 24% đề thi — đây là domain lớn thứ hai. Bạn cần hiểu rõ Generative AI, Foundation Models, và cách chúng khác biệt so với traditional ML.
1. What is Generative AI?
Generative AI là nhánh của AI tập trung vào việc tạo nội dung mới (text, images, code, audio, video) dựa trên patterns học được từ training data.
Discriminative vs Generative AI
| Aspect | Discriminative AI | Generative AI |
|---|---|---|
| What it does | Classify / predict | Create / generate |
| Output | Label, category, number | New content (text, image, code) |
| Example | "Is this email spam?" → Yes/No | "Write an email about..." → New email |
| Models | Logistic Regression, SVM, CNN classifier | GPT, Claude, Stable Diffusion, DALL-E |
Generative AI Modalities
| Input → Output | Examples | Models |
|---|---|---|
| Text → Text | Chatbot, summarization, translation | GPT-4, Claude, Llama |
| Text → Image | Image generation from description | DALL-E, Stable Diffusion, Titan Image Generator |
| Text → Code | Code generation, debugging | CodeWhisperer, Copilot |
| Text → Audio | Speech synthesis, music generation | Amazon Polly (TTS) |
| Image → Text | Image captioning, visual Q&A | Claude (multi-modal), GPT-4V |
| Audio → Text | Transcription | Amazon Transcribe, Whisper |
2. Foundation Models
Foundation Model (FM) là model AI cực lớn, được pre-trained trên massive datasets, có thể adapt cho nhiều downstream tasks khác nhau.
Key Characteristics
- Large-scale pre-training: Trained on billions of data points (text from internet, books, code)
- General-purpose: Can handle multiple tasks without task-specific training
- Adaptable: Can be fine-tuned or prompted for specific use cases
- Expensive to train: Requires massive compute (GPU/TPU clusters)
- Accessible via API: Users don't need to train — use through APIs (Amazon Bedrock)
Foundation Model Lifecycle
┌─────────────────┐ ┌──────────────┐ ┌──────────────┐
│ 1. Pre-training │────→│ 2. Fine- │────→│ 3. Inference │
│ (Massive data, │ │ tuning │ │ (Use model │
│ Billion params,│ │ (Adapt to │ │ via API or │
│ Very expensive)│ │ specific │ │ endpoint) │
│ │ │ domain) │ │ │
└─────────────────┘ └──────────────┘ └──────────────┘
Model Provider You/Org Users
(Anthropic, Meta, (Applications)
Amazon, etc.)
3. Tokenization
Tokenization là quá trình chia text thành các đơn vị nhỏ (tokens) mà model hiểu được.
Input: "Machine learning is amazing!"
Tokens: ["Machine", " learning", " is", " amazing", "!"]
token_1 token_2 token_3 token_4 token_5
OR (subword tokenization):
Tokens: ["Mach", "ine", " learn", "ing", " is", " amaz", "ing", "!"]
Key Concepts for Exam:
- Token ≠ word: A token can be part of a word, a whole word, or punctuation
- Context window: Maximum number of tokens a model can process at once (input + output)
- Token limit: Determines how much text the model can "see" and generate
- Pricing: API calls are typically priced per token (input tokens + output tokens)
Exam tip: Context window size matters. Larger context = can process longer documents. But costs more and may be slower.
4. Model Parameters & Inference Settings
4.1. Model Parameters (Learned during training)
- Parameters = weights and biases trong neural network
- GPT-4: ~1.7 trillion parameters, Claude: undisclosed, Llama 3: 8B/70B/405B
- More parameters → generally more capable, but more expensive
4.2. Inference Parameters (Set by user)
Khi gọi model, bạn có thể điều chỉnh các inference parameters:
| Parameter | Range | What it controls |
|---|---|---|
| Temperature | 0.0 → 1.0+ | Randomness/creativity. Low = deterministic, focused. High = creative, diverse. |
| Top-p (Nucleus) | 0.0 → 1.0 | Cumulative probability threshold. Lower = more focused vocabulary. |
| Top-k | 1 → ∞ | Number of top tokens to consider. Lower = more predictable. |
| Max tokens | 1 → limit | Maximum length of generated output. |
| Stop sequences | strings | Text that tells model to stop generating. |
Temperature Guide for Exam
Temperature = 0 → Most deterministic (factual Q&A, code, data extraction)
Temperature = 0.3 → Slightly creative (business writing, summaries)
Temperature = 0.7 → Creative (stories, brainstorming, marketing copy)
Temperature = 1.0+ → Very random (poetry, creative writing — may hallucinate more)
Exam tip: "A company needs consistent, accurate answers for customer FAQ" → use low temperature. "A company wants creative marketing slogans" → use high temperature.
5. Hallucination
Hallucination là khi model tạo ra output confident nhưng incorrect — bịa ra facts, citations, hoặc thông tin không tồn tại.
Causes:
- Training data gaps or outdated information
- Model doesn't truly "know" facts — it predicts likely next tokens
- Ambiguous or too-open prompts
- High temperature settings
Mitigation Strategies:
| Strategy | How it helps |
|---|---|
| RAG (Retrieval-Augmented Generation) | Ground responses in actual data from knowledge base |
| Lower temperature | Reduce randomness in generation |
| Guardrails | Filter/validate outputs (Amazon Bedrock Guardrails) |
| Better prompts | "Only answer based on provided context" / "Say I don't know if unsure" |
| Fine-tuning | Train model on domain-specific accurate data |
| Human review | Human-in-the-loop validation |
6. Foundation Models on AWS (Amazon Bedrock)
Amazon Bedrock cung cấp access đến nhiều Foundation Models từ các providers:
| Provider | Models | Strengths |
|---|---|---|
| Anthropic | Claude 3 (Haiku, Sonnet, Opus) | Reasoning, safety, long context |
| Meta | Llama 3 | Open-source, versatile |
| Amazon | Titan (Text, Embeddings, Image) | AWS-native, embeddings for RAG |
| Mistral AI | Mistral, Mixtral | Efficient, fast inference |
| Stability AI | Stable Diffusion | Image generation |
| Cohere | Command, Embed | Enterprise NLP, embeddings |
| AI21 Labs | Jurassic | Text generation |
7. Practice Questions
Q1: What is the PRIMARY advantage of Foundation Models compared to traditional ML models?
- A) They are smaller and faster
- B) They can be adapted to multiple downstream tasks without task-specific training ✓
- C) They never produce incorrect outputs
- D) They don't require any compute resources
Explanation: Foundation Models are pre-trained on massive datasets and can be adapted (via prompting or fine-tuning) for many different tasks. They are large, can hallucinate, and still require compute.
Q2: A company uses a generative AI model and notices it sometimes generates plausible but factually incorrect information. What is this phenomenon called?
- A) Overfitting
- B) Data drift
- C) Hallucination ✓
- D) Bias
Explanation: Hallucination is when a generative AI model produces confident but factually incorrect outputs.
Q3: A developer wants to ensure their generative AI chatbot provides consistent, factual answers with minimal creativity. Which inference parameter should they adjust?
- A) Set max tokens to a very high value
- B) Set temperature close to 0 ✓
- C) Set temperature close to 1
- D) Increase the top-k value
Explanation: Low temperature makes the model more deterministic and focused, reducing creativity and randomness in responses.