Chuyển đến nội dung chính

Lesson 3: Generative AI & Foundation Models

What is Generative AI. Foundation Models: pre-training, fine-tuning. Types: text-to-text, text-to-image, text-to-code. Tokenization. Model parameters, inference, temperature, top-p, top-k.

Foundation Model Lifecycle

Foundation Model Lifecycle — Pre-training, Fine-tuning, RAG, and Prompt Engineering

Domain 2 Overview

Domain 2 accounts for 24% of the exam — this is the second-largest domain. You need to have a solid understanding of Generative AI, Foundation Models, and how they differ from traditional ML.

1. What is Generative AI?

Generative AI is a branch of AI focused on creating new content (text, images, code, audio, video) based on patterns learned from training data.

Discriminative vs Generative AI

AspectDiscriminative AIGenerative AI
What it doesClassify / predictCreate / generate
OutputLabel, category, numberNew content (text, image, code)
Example"Is this email spam?" → Yes/No"Write an email about..." → New email
ModelsLogistic Regression, SVM, CNN classifierGPT, Claude, Stable Diffusion, DALL-E

Generative AI Modalities

Input → OutputExamplesModels
Text → TextChatbot, summarization, translationGPT-4, Claude, Llama
Text → ImageImage generation from descriptionDALL-E, Stable Diffusion, Titan Image Generator
Text → CodeCode generation, debuggingCodeWhisperer, Copilot
Text → AudioSpeech synthesis, music generationAmazon Polly (TTS)
Image → TextImage captioning, visual Q&AClaude (multi-modal), GPT-4V
Audio → TextTranscriptionAmazon Transcribe, Whisper

2. Foundation Models

A Foundation Model (FM) is a very large AI model that has been pre-trained on massive datasets and can be adapted for many different downstream tasks.

Key Characteristics

  • Large-scale pre-training: Trained on billions of data points (text from internet, books, code)
  • General-purpose: Can handle multiple tasks without task-specific training
  • Adaptable: Can be fine-tuned or prompted for specific use cases
  • Expensive to train: Requires massive compute (GPU/TPU clusters)
  • Accessible via API: Users don't need to train — use through APIs (Amazon Bedrock)

Foundation Model Lifecycle

┌─────────────────┐     ┌──────────────┐     ┌──────────────┐
│ 1. Pre-training │────→│ 2. Fine-     │────→│ 3. Inference │
│ (Massive data,  │     │ tuning       │     │ (Use model   │
│  Billion params,│     │ (Adapt to    │     │  via API or  │
│  Very expensive)│     │  specific    │     │  endpoint)   │
│                 │     │  domain)     │     │              │
└─────────────────┘     └──────────────┘     └──────────────┘
     Model Provider          You/Org              Users
   (Anthropic, Meta,                         (Applications)
    Amazon, etc.)

3. Tokenization

Tokenization is the process of splitting text into small units (tokens) that the model can understand.

Input:  "Machine learning is amazing!"
Tokens: ["Machine", " learning", " is", " amazing", "!"]
         token_1    token_2      token_3  token_4    token_5

OR (subword tokenization):
Tokens: ["Mach", "ine", " learn", "ing", " is", " amaz", "ing", "!"]

Key Concepts for Exam:

  • Token ≠ word: A token can be part of a word, a whole word, or punctuation
  • Context window: Maximum number of tokens a model can process at once (input + output)
  • Token limit: Determines how much text the model can "see" and generate
  • Pricing: API calls are typically priced per token (input tokens + output tokens)

Exam tip: Context window size matters. Larger context = can process longer documents. But costs more and may be slower.

4. Model Parameters & Inference Settings

4.1. Model Parameters (Learned during training)

  • Parameters = weights and biases in the neural network
  • GPT-4: ~1.7 trillion parameters, Claude: undisclosed, Llama 3: 8B/70B/405B
  • More parameters → generally more capable, but more expensive

4.2. Inference Parameters (Set by user)

When calling a model, you can adjust inference parameters:

ParameterRangeWhat it controls
Temperature0.0 → 1.0+Randomness/creativity. Low = deterministic, focused. High = creative, diverse.
Top-p (Nucleus)0.0 → 1.0Cumulative probability threshold. Lower = more focused vocabulary.
Top-k1 → ∞Number of top tokens to consider. Lower = more predictable.
Max tokens1 → limitMaximum length of generated output.
Stop sequencesstringsText that tells model to stop generating.

Temperature Guide for Exam

Temperature = 0  →  Most deterministic (factual Q&A, code, data extraction)
Temperature = 0.3 → Slightly creative (business writing, summaries)
Temperature = 0.7 → Creative (stories, brainstorming, marketing copy)
Temperature = 1.0+ → Very random (poetry, creative writing — may hallucinate more)

Exam tip: "A company needs consistent, accurate answers for customer FAQ" → use low temperature. "A company wants creative marketing slogans" → use high temperature.

5. Hallucination

Hallucination is when a model generates output that is confident but incorrect — fabricating facts, citations, or information that doesn't exist.

Causes:

  • Training data gaps or outdated information
  • Model doesn't truly "know" facts — it predicts likely next tokens
  • Ambiguous or too-open prompts
  • High temperature settings

Mitigation Strategies:

StrategyHow it helps
RAG (Retrieval-Augmented Generation)Ground responses in actual data from knowledge base
Lower temperatureReduce randomness in generation
GuardrailsFilter/validate outputs (Amazon Bedrock Guardrails)
Better prompts"Only answer based on provided context" / "Say I don't know if unsure"
Fine-tuningTrain model on domain-specific accurate data
Human reviewHuman-in-the-loop validation

6. Foundation Models on AWS (Amazon Bedrock)

Amazon Bedrock provides access to many Foundation Models from various providers:

ProviderModelsStrengths
AnthropicClaude 3 (Haiku, Sonnet, Opus)Reasoning, safety, long context
MetaLlama 3Open-source, versatile
AmazonTitan (Text, Embeddings, Image)AWS-native, embeddings for RAG
Mistral AIMistral, MixtralEfficient, fast inference
Stability AIStable DiffusionImage generation
CohereCommand, EmbedEnterprise NLP, embeddings
AI21 LabsJurassicText generation

7. Practice Questions

Q1: What is the PRIMARY advantage of Foundation Models compared to traditional ML models?

  • A) They are smaller and faster
  • B) They can be adapted to multiple downstream tasks without task-specific training ✓
  • C) They never produce incorrect outputs
  • D) They don't require any compute resources

Explanation: Foundation Models are pre-trained on massive datasets and can be adapted (via prompting or fine-tuning) for many different tasks. They are large, can hallucinate, and still require compute.

Q2: A company uses a generative AI model and notices it sometimes generates plausible but factually incorrect information. What is this phenomenon called?

  • A) Overfitting
  • B) Data drift
  • C) Hallucination ✓
  • D) Bias

Explanation: Hallucination is when a generative AI model produces confident but factually incorrect outputs.

Q3: A developer wants to ensure their generative AI chatbot provides consistent, factual answers with minimal creativity. Which inference parameter should they adjust?

  • A) Set max tokens to a very high value
  • B) Set temperature close to 0 ✓
  • C) Set temperature close to 1
  • D) Increase the top-k value

Explanation: Low temperature makes the model more deterministic and focused, reducing creativity and randomness in responses.