Chuyển đến nội dung chính

Bài 2: ML Development Lifecycle & AWS AI Services Overview

ML pipeline: data collection → feature engineering → training → evaluation → deployment. AWS AI/ML service stack. SageMaker, Rekognition, Comprehend, Polly, Transcribe, Translate, Textract, Lex, Personalize, Forecast, Kendra.

ML Development Lifecycle Pipeline on AWS

ML Development Lifecycle Pipeline và AWS AI/ML Service Stack

1. ML Development Lifecycle

Đề thi AIF-C01 yêu cầu bạn hiểu toàn bộ vòng đời phát triển ML — từ khi xác định bài toán đến khi deploy và monitor model.

┌─────────────┐    ┌──────────────┐    ┌──────────────┐
│ 1. Business │───→│ 2. Data      │───→│ 3. Feature   │
│ Problem     │    │ Collection & │    │ Engineering  │
│ Definition  │    │ Preparation  │    │              │
└─────────────┘    └──────────────┘    └──────────────┘
                                              │
┌─────────────┐    ┌──────────────┐    ┌──────┴───────┐
│ 6. Monitor  │←───│ 5. Deploy    │←───│ 4. Model     │
│ & Retrain   │    │ & Inference  │    │ Training &   │
│             │    │              │    │ Evaluation   │
└─────────────┘    └──────────────┘    └──────────────┘

Step 1: Business Problem Definition

  • Xác định bài toán có thực sự cần ML không (sometimes rules-based is enough)
  • Define success metrics (KPIs)
  • Determine data availability

Exam tip: "Not every problem needs ML." Nếu đề bài mô tả bài toán đơn giản, có thể rule-based hoặc lookup table là đủ.

Step 2: Data Collection & Preparation

  • Data Collection: Thu thập từ databases, APIs, IoT, logs
  • Data Cleaning: Handle missing values, remove duplicates, fix errors
  • Data Labeling: Gắn nhãn cho supervised learning → Amazon SageMaker Ground Truth
  • Exploratory Data Analysis (EDA): Visualize, understand distributions, correlations

Step 3: Feature Engineering

  • Feature selection: Chọn features quan trọng, loại bỏ noise
  • Feature transformation: Normalization, scaling, encoding
  • Feature creation: Tạo features mới từ raw data
  • AWS: SageMaker Data Wrangler, SageMaker Feature Store

Step 4: Model Training & Evaluation

  • Choose algorithm appropriate for the problem
  • Split data into training/validation/test sets
  • Train model, tune hyperparameters
  • Evaluate using appropriate metrics (accuracy, F1, RMSE...)
  • AWS: Amazon SageMaker for full ML workflow

Step 5: Deployment & Inference

  • Real-time inference: Endpoint cho instant predictions
  • Batch inference: Process large datasets offline
  • Edge deployment: Run model on edge devices
  • AWS: SageMaker Endpoints, Lambda, IoT Greengrass

Step 6: Monitoring & Retraining

  • Model drift: Performance degrades over time as data changes
  • Data drift: Input data distribution changes
  • Concept drift: Relationship between input and output changes
  • Solution: Monitor → detect drift → retrain with new data
  • AWS: SageMaker Model Monitor

2. AWS AI/ML Service Stack

AWS cung cấp 3 layers of AI/ML services — từ high-level (no ML knowledge needed) đến low-level (full control):

┌─────────────────────────────────────────────────────┐
│  Layer 3: AI Services (Pre-trained, API-based)      │
│  → Rekognition, Comprehend, Polly, Transcribe,      │
│    Translate, Textract, Lex, Personalize, Forecast   │
│  → NO ML expertise needed                           │
├─────────────────────────────────────────────────────┤
│  Layer 2: ML Services (Managed platform)            │
│  → Amazon SageMaker, SageMaker JumpStart            │
│  → Amazon Bedrock (GenAI)                           │
│  → SOME ML expertise needed                         │
├─────────────────────────────────────────────────────┤
│  Layer 1: ML Frameworks & Infrastructure            │
│  → EC2 with GPU/Inferentia, Deep Learning AMIs,     │
│    Deep Learning Containers                         │
│  → FULL ML expertise needed                         │
└─────────────────────────────────────────────────────┘

3. AWS AI Services — Bảng Tổng hợp

Đây là phần rất quan trọng cho đề thi — bạn cần biết mỗi service làm gì và khi nào dùng.

3.1. Computer Vision

ServiceWhat it doesUse Cases
Amazon RekognitionImage and video analysisFace detection, object detection, content moderation, celebrity recognition, text in images (OCR)
Amazon TextractExtract text & data from documentsInvoice processing, ID document extraction, form data, table extraction
Amazon Lookout for VisionVisual inspection for manufacturingDefect detection in products on assembly line

3.2. Natural Language Processing (NLP)

ServiceWhat it doesUse Cases
Amazon ComprehendNLP analysisSentiment analysis, entity extraction, key phrases, language detection, PII detection
Amazon TranslateNeural machine translationReal-time translation, batch document translation
Amazon KendraIntelligent enterprise searchInternal knowledge search, FAQ, document search powered by NLP

3.3. Speech

ServiceWhat it doesDirection
Amazon PollyText-to-Speech (TTS)Text → Audio
Amazon TranscribeSpeech-to-Text (STT)Audio → Text
Amazon LexConversational AI (chatbot)Build chatbots with voice & text (powers Alexa)

Exam tip: Polly = text TO speech (Polly "speaks"). Transcribe = speech TO text (Transcribe "writes down").

3.4. Predictions & Recommendations

ServiceWhat it doesUse Cases
Amazon PersonalizeReal-time personalization & recommendationsProduct recommendations, personalized content
Amazon ForecastTime-series forecastingDemand planning, financial forecasting, resource planning
Amazon Fraud DetectorDetect online fraudPayment fraud, fake accounts, account takeover

4. Amazon SageMaker Overview

SageMaker là fully managed ML platform — cung cấp mọi thứ cần thiết cho toàn bộ ML lifecycle.

Key Components:

ComponentPurpose
SageMaker StudioIDE cho ML development (Jupyter-based)
SageMaker Ground TruthData labeling service (human + ML-assisted)
SageMaker Data WranglerData preparation & transformation (no code)
SageMaker Feature StoreStore & share ML features
SageMaker TrainingManaged training jobs with built-in algorithms
SageMaker AutopilotAutoML — automatic model building
SageMaker JumpStartPre-trained models & solutions (model hub)
SageMaker EndpointsDeploy models for real-time inference
SageMaker Model MonitorMonitor deployed models for drift
SageMaker ClarifyBias detection & model explainability
SageMaker CanvasNo-code ML for business users

When to use SageMaker vs AI Services?

Need custom ML model? → SageMaker
Need pre-trained capability? → AI Services (Rekognition, Comprehend, etc.)
Need GenAI/Foundation Models? → Amazon Bedrock
Business user, no code? → SageMaker Canvas

5. Use Case → AWS Service Mapping

Đây là dạng câu hỏi rất phổ biến trong đề thi:

Use CaseAWS Service
Detect faces in photosAmazon Rekognition
Extract data from invoicesAmazon Textract
Analyze customer review sentimentAmazon Comprehend
Translate content to multiple languagesAmazon Translate
Build a customer service chatbotAmazon Lex
Convert blog posts to audioAmazon Polly
Transcribe meeting recordingsAmazon Transcribe
Product recommendationsAmazon Personalize
Demand forecastingAmazon Forecast
Search internal documentsAmazon Kendra
Detect fraudulent transactionsAmazon Fraud Detector
Label training dataSageMaker Ground Truth
Build custom ML modelAmazon SageMaker
No-code ML for business analystsSageMaker Canvas
Generate text with LLMAmazon Bedrock

6. Practice Questions

Q1: A company wants to automatically extract text and structured data from scanned invoices. Which AWS service should they use?

  • A) Amazon Comprehend
  • B) Amazon Rekognition
  • C) Amazon Textract ✓
  • D) Amazon Translate

Explanation: Textract is specifically designed to extract text, forms, and tables from scanned documents. Comprehend analyzes text meaning, not document extraction. Rekognition is for image/video analysis.

Q2: A data scientist notices that their deployed model's prediction accuracy has decreased over the past month. The input data patterns have changed. What is this called?

  • A) Overfitting
  • B) Underfitting
  • C) Data drift ✓
  • D) Feature engineering

Explanation: When the statistical properties of model input data change over time, causing performance degradation, this is called data drift.

Q3: Which AWS service allows business analysts with no ML experience to build ML models using a visual interface?

  • A) SageMaker Studio
  • B) SageMaker Autopilot
  • C) SageMaker Canvas ✓
  • D) SageMaker JumpStart

Explanation: SageMaker Canvas provides a no-code, visual point-and-click interface for business analysts. Autopilot automates model building but requires some ML knowledge. JumpStart provides pre-trained models. Studio is the full ML IDE.