ML Problem Framing: xác định bài toán, chọn loại model, và định nghĩa metrics theo chuẩn Google
1. Khi Nào Cần Dùng ML?
Google ML certification thường hỏi về problem framing — tức là xác định xem bài toán có phù hợp để áp dụng ML không, và nếu có thì dùng loại ML nào. Đây là skill quan trọng của một professional ML Engineer.
| Câu hỏi cần đặt ra | Nếu "Có" | Nếu "Không" |
|---|---|---|
| Có pattern phức tạp trong data không? | ML có thể giúp | Rules-based logic đủ rồi |
| Có đủ data (labels) không? | Supervised Learning | Unsupervised hoặc thu thập thêm |
| Output có thể định nghĩa rõ ràng không? | Supervised ML | Cần clarify với stakeholders |
| Bài toán có cần agent tương tác với environment không? | Reinforcement Learning | Supervised/Unsupervised |
2. Các Loại ML và Khi Nào Dùng
Problem Framing Decision Tree:
Has labeled training data?
YES → Supervised Learning
├── Output is category? → Classification
└── Output is number? → Regression
NO → Has examples, no labels?
YES → Unsupervised Learning
├── Find groups? → Clustering
└── Find patterns/anomalies? → Density estimation
NO → Agent in environment?
YES → Reinforcement Learning
NO → Reconsider problem definition
| ML Type | When to Use | GCP Services |
|---|---|---|
| Supervised Classification | Email spam, image labels, churn prediction | Vertex AI AutoML, BigQuery ML |
| Supervised Regression | Price prediction, demand forecast | Vertex AI, BigQuery ML BQML_REGRESSOR |
| Unsupervised Clustering | Customer segmentation, topic discovery | Vertex AI Custom Training (k-means) |
| Reinforcement Learning | Game agents, robotics, ad bidding | Vertex AI + custom environment |
| Self-supervised | LLMs, foundation models | Vertex AI Model Garden |
3. Business Metrics vs. ML Metrics
Một trong những sai lầm phổ biến là optimize nhầm metric. Mục tiêu ML phải align với mục tiêu business.
| Business Goal | Wrong ML Metric | Correct ML Metric |
|---|---|---|
| Giảm doanh thu bị gian lận | Accuracy (99%!) | Recall (bắt được nhiều fraud) |
| Giảm email spam trải nghiệm người dùng | Recall | Precision (ít false positive) |
| Dự báo nhu cầu tồn kho | MSE | MAPE (scale-independent) |
| Ranking sản phẩm trong search | Accuracy | NDCG, MRR (ranking metrics) |
Exam tip: Professional ML Engineer exam thường hỏi "which metric BEST aligns with the business objective". Khi thấy fraud/medical diagnosis → Recall. Khi thấy spam/precision-critical → Precision. Khi thấy class imbalance → F1 hoặc AUC-ROC.
4. Data Availability Assessment
| Data Situation | ML Approach |
|---|---|
| Nhiều labeled data | Fully supervised, train from scratch |
| Ít labeled data (<1000) | Transfer Learning (pre-trained + fine-tune) |
| Không có labels | Unsupervised hoặc thu thập labels (Vertex AI Data Labeling) |
| Labels tốn kém | Active Learning — label uncertain samples trước |
| Dữ liệu không cân bằng | Oversampling, undersampling, class weights |
5. Google's ML Best Practices
- Start simple: Bắt đầu với model đơn giản nhất, sau đó phức tạp hóa dần
- Establish baseline: So sánh với heuristic/rules trước khi dùng ML
- Data quality first: 80% thời gian ML project là data preparation
- Reproducibility: Pipeline phải reproducible với cùng data
- Monitor in production: Model decay theo thời gian — cần continuous monitoring
6. Practice Questions
Q1: A company wants to identify which of its customers are most likely to cancel their subscription in the next 30 days. They have 3 years of historical customer behavior data with known churn events. Which ML approach should they use?
- A) Unsupervised clustering to find customer groups
- B) Reinforcement learning to optimize retention campaigns
- C) Supervised binary classification with historical churn labels ✓
- D) Anomaly detection to find unusual behavior
Explanation: This is a classic supervised classification problem (churn = yes/no). Historical data with known outcomes (churned/not churned) provides the labels needed. Clustering would not predict individual churn probability. RL is for sequential decision making, not prediction.
Q2: A medical imaging ML model achieves 98% accuracy on test data but the business team is unsatisfied. The task is detecting rare cancer cells (1% prevalence). What is the most likely issue?
- A) The model is overfitting to training data
- B) Accuracy is the wrong metric — the model may be predicting "no cancer" for everything ✓
- C) The model needs more training iterations
- D) The test dataset is too small
Explanation: With 1% prevalence, a model always predicting "no cancer" achieves 99% accuracy but has 0% recall — it misses every cancer case. For rare class problems, Recall (sensitivity) is the critical metric, not accuracy.
Q3: A startup has 500 labeled product images for a new custom classification task. Which training approach is MOST appropriate?
- A) Train a deep learning CNN from scratch on the 500 images
- B) Use AutoML Tabular on the image metadata
- C) Use Transfer Learning from a pre-trained image model ✓
- D) Apply K-Means clustering since the dataset is too small
Explanation: With only 500 labeled examples, training from scratch would overfit severely. Transfer Learning reuses features from a model pre-trained on millions of images (e.g., ImageNet), requiring far less data to achieve good accuracy on the new task.