Chuyển đến nội dung chính

Lesson 10: Common ML Problem Patterns

Fraud detection, recommendation systems, NLP pipeline, time series forecasting, computer vision — identify the problem and choose the right AWS service/algorithm.

AWS ML Problem Patterns

ML Problem Patterns: Fraud detection, Recommendation, NLP, Time Series, and Computer Vision on AWS

1. Quick Problem Recognition

Most MLS-C01 questions are scenario-based: read the business problem description, choose the most appropriate AWS service or algorithm. This lesson provides a framework for quick recognition.

ML Problem Recognition Framework:

READ SCENARIO
    ↓
Is output a CATEGORY?          → Classification
Is output a NUMBER?            → Regression
Is output a CLUSTER/GROUP?     → Clustering (no labels)
Is output a RANK/SCORE?        → Recommendation / Ranking
Is output a DETECTION (anomaly)? → Anomaly Detection
Is output a SEQUENCE?          → Time Series / Seq2Seq

2. Fraud Detection

Fraud detection is a classic exam problem. Key challenge: extreme class imbalance (99.9% of transactions are legitimate).

ChallengeSolution
Class imbalanceSMOTE, class_weight, precision-recall (not accuracy)
Real-time scoringSageMaker Real-time Endpoint
Unsupervised fraud (no labels)Random Cut Forest (anomaly detection)
Graph-based fraud ringsAmazon Neptune + GNN (GraphSAGE)
Supervised (labeled history)XGBoost or Linear Learner

Exam tip: When the question says "no fraud labels" or "detect unusual behavior" → Random Cut Forest. When there's labeled historical fraud data → Classification (XGBoost).

3. Recommendation Systems

TypeAlgorithmAWS Service
Collaborative FilteringFactorization Machines, Neural CFSageMaker FM, Amazon Personalize
Content-BasedTF-IDF, EmbeddingsSageMaker Knn
HybridFM + content featuresAmazon Personalize (HRNN)
Cold Start ProblemContent features, metadataPersonalize context metadata

The exam often asks: "Real-time personalized recommendations for e-commerce" → the answer is almost always Amazon Personalize (managed service).

4. NLP Pipeline

NLP Task Decision Tree:

Classify documents?
    → Text Classification → BlazingText (word2vec mode), Comprehend Custom

Extract entities from text?
    → Named Entity Recognition → Amazon Comprehend (NER)

Translate text?
    → Amazon Translate

Summarize documents?
    → Hugging Face on SageMaker (T5, BART)

Q&A / Generation?
    → SageMaker JumpStart foundation models (Llama, Falcon)

Toxic content detection?
    → Amazon Comprehend (Sentiment + custom classifier)
NLP TaskAlgorithm/Service
Document classificationBlazingText, Comprehend Custom
Topic modelingLatent Dirichlet Allocation (LDA), NTM
Sentiment analysisAmazon Comprehend
Language detectionAmazon Comprehend
Machine translationAmazon Translate
Conversational AIAmazon Lex (chatbot) + Lambda

5. Time Series Forecasting

Service/AlgorithmUse Case
DeepAR+ (SageMaker built-in)Multiple related time series, cold start
Amazon ForecastFully managed, business forecasting (demand, inventory)
ARIMASingle stationary series, classic approach
ProphetSeasonality + holidays, Facebook's library

Exam tip: "Predict retail demand for thousands of products" → DeepAR+ (multiple related series) or Amazon Forecast (managed). Key differentiator: DeepAR+ trains one model for ALL items, while classical methods need one model per series.

6. Computer Vision

CV TaskAlgorithmService
Image ClassificationImage Classification (ResNet)SageMaker built-in
Object DetectionObject Detection (SSD/YOLO)SageMaker built-in
Semantic SegmentationSemantic SegmentationSageMaker built-in
Rekognition (no ML needed)Faces, objects, text, moderationAmazon Rekognition
Custom labels in imagesFine-tuningRekognition Custom Labels

7. Business Domain → AWS Service Mapping

Industry / ScenarioAWS Service
E-commerce personalizationAmazon Personalize
Call center analyticsAmazon Transcribe + Comprehend
Medical image analysisAmazon HealthLake + SageMaker (custom model)
Document understanding (invoices, forms)Amazon Textract
Product search (semantic)Amazon OpenSearch + KNN
IoT anomaly detectionAmazon Lookout for Equipment
Manufacturing defect detectionAmazon Lookout for Vision
Financial forecasting (managed)Amazon Forecast
Churn predictionXGBoost (SageMaker)

8. Practice Questions

Q1: A retail company wants to predict next quarter's sales for 10,000 individual products, incorporating holiday seasonality and promotional events. Which solution is MOST appropriate?

  • A) Train separate ARIMA models for each of the 10,000 products
  • B) Use Amazon Forecast with related time series data ✓
  • C) Use Amazon Comprehend for trend analysis
  • D) Use Amazon Rekognition to analyze product images

Explanation: Amazon Forecast is designed for business demand forecasting at scale, can handle thousands of related time series simultaneously, and supports external factors like holidays and promotions. Training 10,000 separate ARIMA models would be impractical and miss cross-series patterns.

Q2: A bank needs to detect fraudulent transactions in real time WITHOUT having labeled historical fraud data. Which algorithm should they use?

  • A) XGBoost classifier
  • B) Linear Learner with binary classification
  • C) Random Cut Forest ✓
  • D) BlazingText

Explanation: Without labeled fraud data, this is an unsupervised anomaly detection problem. Random Cut Forest detects anomalies (transactions that deviate from normal patterns) without requiring labeled examples. XGBoost and Linear Learner are supervised and require labeled training data.

Q3: A company wants to extract key-value pairs from scanned insurance claim forms (PDFs). No custom ML training has been done. Which AWS service should they use?

  • A) Amazon Comprehend
  • B) Amazon Textract ✓
  • C) SageMaker Object Detection
  • D) Amazon Rekognition

Explanation: Amazon Textract is purpose-built for extracting text, forms (key-value pairs), and tables from scanned documents including PDFs. It requires no ML training. Comprehend analyzes text meaning; Rekognition analyzes images; neither handles structured form extraction like Textract.