Chuyển đến nội dung chính

Lesson 1: Framing ML Problems — Supervised, Unsupervised, RL

How to determine if a problem needs ML. Choosing the right model type. Business metrics vs ML metrics. Data availability assessment. Google's ML best practices.

ML Problem Framing Framework

ML Problem Framing: identifying the problem, choosing the model type, and defining metrics per Google standards

1. When to Use ML?

The Google ML certification often tests problem framing — determining whether a problem is suitable for ML and, if so, which type of ML to apply. This is a critical skill for a professional ML Engineer.

Question to AskIf "Yes"If "No"
Are there complex patterns in the data?ML can helpRules-based logic is sufficient
Is there enough data (labels)?Supervised LearningUnsupervised or collect more data
Can the output be clearly defined?Supervised MLClarify with stakeholders
Does the problem require an agent interacting with an environment?Reinforcement LearningSupervised/Unsupervised

2. ML Types and When to Use Them

Problem Framing Decision Tree:

Has labeled training data?
    YES → Supervised Learning
           ├── Output is category? → Classification
           └── Output is number? → Regression

    NO → Has examples, no labels?
           YES → Unsupervised Learning
                  ├── Find groups? → Clustering
                  └── Find patterns/anomalies? → Density estimation
           NO → Agent in environment?
                  YES → Reinforcement Learning
                  NO → Reconsider problem definition
ML TypeWhen to UseGCP Services
Supervised ClassificationEmail spam, image labels, churn predictionVertex AI AutoML, BigQuery ML
Supervised RegressionPrice prediction, demand forecastVertex AI, BigQuery ML BQML_REGRESSOR
Unsupervised ClusteringCustomer segmentation, topic discoveryVertex AI Custom Training (k-means)
Reinforcement LearningGame agents, robotics, ad biddingVertex AI + custom environment
Self-supervisedLLMs, foundation modelsVertex AI Model Garden

3. Business Metrics vs. ML Metrics

A common mistake is optimizing the wrong metric. ML objectives must align with business goals.

Business GoalWrong ML MetricCorrect ML Metric
Reduce fraud-related revenue lossAccuracy (99%!)Recall (catch more fraud)
Reduce spam for better user experienceRecallPrecision (fewer false positives)
Forecast inventory demandMSEMAPE (scale-independent)
Rank products in search resultsAccuracyNDCG, MRR (ranking metrics)

Exam tip: The Professional ML Engineer exam often asks "which metric BEST aligns with the business objective." For fraud/medical diagnosis → Recall. For spam/precision-critical → Precision. For class imbalance → F1 or AUC-ROC.

4. Data Availability Assessment

Data SituationML Approach
Plenty of labeled dataFully supervised, train from scratch
Little labeled data (<1000)Transfer Learning (pre-trained + fine-tune)
No labelsUnsupervised or collect labels (Vertex AI Data Labeling)
Labels are expensiveActive Learning — label uncertain samples first
Imbalanced dataOversampling, undersampling, class weights

5. Google's ML Best Practices

  • Start simple: Begin with the simplest model, then increase complexity gradually
  • Establish baseline: Compare against heuristics/rules before using ML
  • Data quality first: 80% of ML project time is data preparation
  • Reproducibility: Pipelines must be reproducible with the same data
  • Monitor in production: Models decay over time — continuous monitoring is needed

6. Practice Questions

Q1: A company wants to identify which of its customers are most likely to cancel their subscription in the next 30 days. They have 3 years of historical customer behavior data with known churn events. Which ML approach should they use?

  • A) Unsupervised clustering to find customer groups
  • B) Reinforcement learning to optimize retention campaigns
  • C) Supervised binary classification with historical churn labels ✓
  • D) Anomaly detection to find unusual behavior

Explanation: This is a classic supervised classification problem (churn = yes/no). Historical data with known outcomes (churned/not churned) provides the labels needed. Clustering would not predict individual churn probability. RL is for sequential decision making, not prediction.

Q2: A medical imaging ML model achieves 98% accuracy on test data but the business team is unsatisfied. The task is detecting rare cancer cells (1% prevalence). What is the most likely issue?

  • A) The model is overfitting to training data
  • B) Accuracy is the wrong metric — the model may be predicting "no cancer" for everything ✓
  • C) The model needs more training iterations
  • D) The test dataset is too small

Explanation: With 1% prevalence, a model always predicting "no cancer" achieves 99% accuracy but has 0% recall — it misses every cancer case. For rare class problems, Recall (sensitivity) is the critical metric, not accuracy.

Q3: A startup has 500 labeled product images for a new custom classification task. Which training approach is MOST appropriate?

  • A) Train a deep learning CNN from scratch on the 500 images
  • B) Use AutoML Tabular on the image metadata
  • C) Use Transfer Learning from a pre-trained image model ✓
  • D) Apply K-Means clustering since the dataset is too small

Explanation: With only 500 labeled examples, training from scratch would overfit severely. Transfer Learning reuses features from a model pre-trained on millions of images (e.g., ImageNet), requiring far less data to achieve good accuracy on the new task.