Chuyển đến nội dung chính

Bài 3: Generative AI & Foundation Models

Generative AI là gì. Foundation Models: pre-training, fine-tuning. Types: text-to-text, text-to-image, text-to-code. Tokenization. Model parameters, inference, temperature, top-p, top-k.

Foundation Model Lifecycle

Foundation Model Lifecycle — Pre-training, Fine-tuning, RAG và Prompt Engineering

Tổng quan Domain 2

Domain 2 chiếm 24% đề thi — đây là domain lớn thứ hai. Bạn cần hiểu rõ Generative AI, Foundation Models, và cách chúng khác biệt so với traditional ML.

1. What is Generative AI?

Generative AI là nhánh của AI tập trung vào việc tạo nội dung mới (text, images, code, audio, video) dựa trên patterns học được từ training data.

Discriminative vs Generative AI

AspectDiscriminative AIGenerative AI
What it doesClassify / predictCreate / generate
OutputLabel, category, numberNew content (text, image, code)
Example"Is this email spam?" → Yes/No"Write an email about..." → New email
ModelsLogistic Regression, SVM, CNN classifierGPT, Claude, Stable Diffusion, DALL-E

Generative AI Modalities

Input → OutputExamplesModels
Text → TextChatbot, summarization, translationGPT-4, Claude, Llama
Text → ImageImage generation from descriptionDALL-E, Stable Diffusion, Titan Image Generator
Text → CodeCode generation, debuggingCodeWhisperer, Copilot
Text → AudioSpeech synthesis, music generationAmazon Polly (TTS)
Image → TextImage captioning, visual Q&AClaude (multi-modal), GPT-4V
Audio → TextTranscriptionAmazon Transcribe, Whisper

2. Foundation Models

Foundation Model (FM) là model AI cực lớn, được pre-trained trên massive datasets, có thể adapt cho nhiều downstream tasks khác nhau.

Key Characteristics

  • Large-scale pre-training: Trained on billions of data points (text from internet, books, code)
  • General-purpose: Can handle multiple tasks without task-specific training
  • Adaptable: Can be fine-tuned or prompted for specific use cases
  • Expensive to train: Requires massive compute (GPU/TPU clusters)
  • Accessible via API: Users don't need to train — use through APIs (Amazon Bedrock)

Foundation Model Lifecycle

┌─────────────────┐     ┌──────────────┐     ┌──────────────┐
│ 1. Pre-training │────→│ 2. Fine-     │────→│ 3. Inference │
│ (Massive data,  │     │ tuning       │     │ (Use model   │
│  Billion params,│     │ (Adapt to    │     │  via API or  │
│  Very expensive)│     │  specific    │     │  endpoint)   │
│                 │     │  domain)     │     │              │
└─────────────────┘     └──────────────┘     └──────────────┘
     Model Provider          You/Org              Users
   (Anthropic, Meta,                         (Applications)
    Amazon, etc.)

3. Tokenization

Tokenization là quá trình chia text thành các đơn vị nhỏ (tokens) mà model hiểu được.

Input:  "Machine learning is amazing!"
Tokens: ["Machine", " learning", " is", " amazing", "!"]
         token_1    token_2      token_3  token_4    token_5

OR (subword tokenization):
Tokens: ["Mach", "ine", " learn", "ing", " is", " amaz", "ing", "!"]

Key Concepts for Exam:

  • Token ≠ word: A token can be part of a word, a whole word, or punctuation
  • Context window: Maximum number of tokens a model can process at once (input + output)
  • Token limit: Determines how much text the model can "see" and generate
  • Pricing: API calls are typically priced per token (input tokens + output tokens)

Exam tip: Context window size matters. Larger context = can process longer documents. But costs more and may be slower.

4. Model Parameters & Inference Settings

4.1. Model Parameters (Learned during training)

  • Parameters = weights and biases trong neural network
  • GPT-4: ~1.7 trillion parameters, Claude: undisclosed, Llama 3: 8B/70B/405B
  • More parameters → generally more capable, but more expensive

4.2. Inference Parameters (Set by user)

Khi gọi model, bạn có thể điều chỉnh các inference parameters:

ParameterRangeWhat it controls
Temperature0.0 → 1.0+Randomness/creativity. Low = deterministic, focused. High = creative, diverse.
Top-p (Nucleus)0.0 → 1.0Cumulative probability threshold. Lower = more focused vocabulary.
Top-k1 → ∞Number of top tokens to consider. Lower = more predictable.
Max tokens1 → limitMaximum length of generated output.
Stop sequencesstringsText that tells model to stop generating.

Temperature Guide for Exam

Temperature = 0  →  Most deterministic (factual Q&A, code, data extraction)
Temperature = 0.3 → Slightly creative (business writing, summaries)
Temperature = 0.7 → Creative (stories, brainstorming, marketing copy)
Temperature = 1.0+ → Very random (poetry, creative writing — may hallucinate more)

Exam tip: "A company needs consistent, accurate answers for customer FAQ" → use low temperature. "A company wants creative marketing slogans" → use high temperature.

5. Hallucination

Hallucination là khi model tạo ra output confident nhưng incorrect — bịa ra facts, citations, hoặc thông tin không tồn tại.

Causes:

  • Training data gaps or outdated information
  • Model doesn't truly "know" facts — it predicts likely next tokens
  • Ambiguous or too-open prompts
  • High temperature settings

Mitigation Strategies:

StrategyHow it helps
RAG (Retrieval-Augmented Generation)Ground responses in actual data from knowledge base
Lower temperatureReduce randomness in generation
GuardrailsFilter/validate outputs (Amazon Bedrock Guardrails)
Better prompts"Only answer based on provided context" / "Say I don't know if unsure"
Fine-tuningTrain model on domain-specific accurate data
Human reviewHuman-in-the-loop validation

6. Foundation Models on AWS (Amazon Bedrock)

Amazon Bedrock cung cấp access đến nhiều Foundation Models từ các providers:

ProviderModelsStrengths
AnthropicClaude 3 (Haiku, Sonnet, Opus)Reasoning, safety, long context
MetaLlama 3Open-source, versatile
AmazonTitan (Text, Embeddings, Image)AWS-native, embeddings for RAG
Mistral AIMistral, MixtralEfficient, fast inference
Stability AIStable DiffusionImage generation
CohereCommand, EmbedEnterprise NLP, embeddings
AI21 LabsJurassicText generation

7. Practice Questions

Q1: What is the PRIMARY advantage of Foundation Models compared to traditional ML models?

  • A) They are smaller and faster
  • B) They can be adapted to multiple downstream tasks without task-specific training ✓
  • C) They never produce incorrect outputs
  • D) They don't require any compute resources

Explanation: Foundation Models are pre-trained on massive datasets and can be adapted (via prompting or fine-tuning) for many different tasks. They are large, can hallucinate, and still require compute.

Q2: A company uses a generative AI model and notices it sometimes generates plausible but factually incorrect information. What is this phenomenon called?

  • A) Overfitting
  • B) Data drift
  • C) Hallucination ✓
  • D) Bias

Explanation: Hallucination is when a generative AI model produces confident but factually incorrect outputs.

Q3: A developer wants to ensure their generative AI chatbot provides consistent, factual answers with minimal creativity. Which inference parameter should they adjust?

  • A) Set max tokens to a very high value
  • B) Set temperature close to 0 ✓
  • C) Set temperature close to 1
  • D) Increase the top-k value

Explanation: Low temperature makes the model more deterministic and focused, reducing creativity and randomness in responses.