Chuyển đến nội dung chính

Bài 1: Framing ML Problems — Supervised, Unsupervised, RL

Cách xác định bài toán có cần ML không. Chọn đúng loại model. Business metrics vs ML metrics. Data availability assessment. Google's ML best practices.

ML Problem Framing Framework

ML Problem Framing: xác định bài toán, chọn loại model, và định nghĩa metrics theo chuẩn Google

1. Khi Nào Cần Dùng ML?

Google ML certification thường hỏi về problem framing — tức là xác định xem bài toán có phù hợp để áp dụng ML không, và nếu có thì dùng loại ML nào. Đây là skill quan trọng của một professional ML Engineer.

Câu hỏi cần đặt raNếu "Có"Nếu "Không"
Có pattern phức tạp trong data không?ML có thể giúpRules-based logic đủ rồi
Có đủ data (labels) không?Supervised LearningUnsupervised hoặc thu thập thêm
Output có thể định nghĩa rõ ràng không?Supervised MLCần clarify với stakeholders
Bài toán có cần agent tương tác với environment không?Reinforcement LearningSupervised/Unsupervised

2. Các Loại ML và Khi Nào Dùng

Problem Framing Decision Tree:

Has labeled training data?
    YES → Supervised Learning
           ├── Output is category? → Classification
           └── Output is number? → Regression

    NO → Has examples, no labels?
           YES → Unsupervised Learning
                  ├── Find groups? → Clustering
                  └── Find patterns/anomalies? → Density estimation
           NO → Agent in environment?
                  YES → Reinforcement Learning
                  NO → Reconsider problem definition
ML TypeWhen to UseGCP Services
Supervised ClassificationEmail spam, image labels, churn predictionVertex AI AutoML, BigQuery ML
Supervised RegressionPrice prediction, demand forecastVertex AI, BigQuery ML BQML_REGRESSOR
Unsupervised ClusteringCustomer segmentation, topic discoveryVertex AI Custom Training (k-means)
Reinforcement LearningGame agents, robotics, ad biddingVertex AI + custom environment
Self-supervisedLLMs, foundation modelsVertex AI Model Garden

3. Business Metrics vs. ML Metrics

Một trong những sai lầm phổ biến là optimize nhầm metric. Mục tiêu ML phải align với mục tiêu business.

Business GoalWrong ML MetricCorrect ML Metric
Giảm doanh thu bị gian lậnAccuracy (99%!)Recall (bắt được nhiều fraud)
Giảm email spam trải nghiệm người dùngRecallPrecision (ít false positive)
Dự báo nhu cầu tồn khoMSEMAPE (scale-independent)
Ranking sản phẩm trong searchAccuracyNDCG, MRR (ranking metrics)

Exam tip: Professional ML Engineer exam thường hỏi "which metric BEST aligns with the business objective". Khi thấy fraud/medical diagnosis → Recall. Khi thấy spam/precision-critical → Precision. Khi thấy class imbalance → F1 hoặc AUC-ROC.

4. Data Availability Assessment

Data SituationML Approach
Nhiều labeled dataFully supervised, train from scratch
Ít labeled data (<1000)Transfer Learning (pre-trained + fine-tune)
Không có labelsUnsupervised hoặc thu thập labels (Vertex AI Data Labeling)
Labels tốn kémActive Learning — label uncertain samples trước
Dữ liệu không cân bằngOversampling, undersampling, class weights

5. Google's ML Best Practices

  • Start simple: Bắt đầu với model đơn giản nhất, sau đó phức tạp hóa dần
  • Establish baseline: So sánh với heuristic/rules trước khi dùng ML
  • Data quality first: 80% thời gian ML project là data preparation
  • Reproducibility: Pipeline phải reproducible với cùng data
  • Monitor in production: Model decay theo thời gian — cần continuous monitoring

6. Practice Questions

Q1: A company wants to identify which of its customers are most likely to cancel their subscription in the next 30 days. They have 3 years of historical customer behavior data with known churn events. Which ML approach should they use?

  • A) Unsupervised clustering to find customer groups
  • B) Reinforcement learning to optimize retention campaigns
  • C) Supervised binary classification with historical churn labels ✓
  • D) Anomaly detection to find unusual behavior

Explanation: This is a classic supervised classification problem (churn = yes/no). Historical data with known outcomes (churned/not churned) provides the labels needed. Clustering would not predict individual churn probability. RL is for sequential decision making, not prediction.

Q2: A medical imaging ML model achieves 98% accuracy on test data but the business team is unsatisfied. The task is detecting rare cancer cells (1% prevalence). What is the most likely issue?

  • A) The model is overfitting to training data
  • B) Accuracy is the wrong metric — the model may be predicting "no cancer" for everything ✓
  • C) The model needs more training iterations
  • D) The test dataset is too small

Explanation: With 1% prevalence, a model always predicting "no cancer" achieves 99% accuracy but has 0% recall — it misses every cancer case. For rare class problems, Recall (sensitivity) is the critical metric, not accuracy.

Q3: A startup has 500 labeled product images for a new custom classification task. Which training approach is MOST appropriate?

  • A) Train a deep learning CNN from scratch on the 500 images
  • B) Use AutoML Tabular on the image metadata
  • C) Use Transfer Learning from a pre-trained image model ✓
  • D) Apply K-Means clustering since the dataset is too small

Explanation: With only 500 labeled examples, training from scratch would overfit severely. Transfer Learning reuses features from a model pre-trained on millions of images (e.g., ImageNet), requiring far less data to achieve good accuracy on the new task.