Model Evaluation: Classification metrics (AUC-ROC, F1), Regression metrics (RMSE, MAE), and Confusion Matrix
1. Classification Metrics
Choosing the right metric is one of the most important skills for an ML Engineer. The MLS-C01 exam frequently presents a scenario and asks for the appropriate metric.
1.1. Confusion Matrix
Predicted
Positive Negative
Actual Positive │ TP │ FN │ ← Recall = TP / (TP + FN)
Negative │ FP │ TN │
Precision = TP / (TP + FP) ← of all predicted positive, how many are correct?
Recall = TP / (TP + FN) ← of all actual positive, how many did we catch?
F1 Score = 2 × (P × R) / (P + R) ← harmonic mean
Accuracy = (TP + TN) / Total
| Metric | Optimize When | Real-World Example |
|---|---|---|
| Precision | FP cost is high — don't want false alarms | Spam filter (don't block legitimate email) |
| Recall (Sensitivity) | FN cost is high — don't miss positives | Cancer detection (find all cancer patients) |
| F1 Score | Balance Precision and Recall, imbalanced data | Fraud detection |
| Accuracy | Balanced classes only | Multi-class, balanced datasets |
| AUC-ROC | Ranking quality, threshold-independent | Credit scoring, ad ranking |
| PR-AUC | Imbalanced, care about minority class | Fraud, medical diagnoses |
Exam tip: Common scenario — "Medical diagnosis, missing cancer is worse than a false positive" → optimize Recall. "Spam detector, blocking good emails is bad" → optimize Precision. Imbalanced data → use F1 or AUC-ROC, not Accuracy.
2. Regression Metrics
| Metric | Formula | Sensitivity to Outliers | Use Case |
|---|---|---|---|
| RMSE | √(mean(errors²)) | High — penalizes large errors | When large errors are unacceptable (price prediction) |
| MAE | mean(|errors|) | Low — equal weight all errors | Robust for outliers, demand forecasting |
| R² (R-squared) | 1 - SS_res/SS_tot | Medium | Proportion of variance explained (0–1) |
| MAPE | mean(|error/actual|×100) | High when actuals near 0 | Percentage error, easy business interpretation |
3. Cross-Validation
| Strategy | How It Works | Best For |
|---|---|---|
| Hold-out Split | Train/Val/Test split (e.g., 70/15/15) | Large datasets, fast evaluation |
| K-Fold CV | K subsets, train on K-1, evaluate on 1, repeat K times | Medium datasets, robust estimate |
| Stratified K-Fold | Same as K-Fold but maintains class proportions each fold | Imbalanced classification |
| Leave-One-Out (LOOCV) | N-fold (each sample is test once) | Very small datasets |
| Time-Series Split | Training window grows forward — no future data in training | Time series data |
Exam tip: Time series data MUST use time-based splits, you cannot shuffle and use regular K-Fold — that would leak future data into training.
4. SageMaker Clarify — Bias & Explainability
SageMaker Clarify detects bias in data/models and provides model explainability using SHAP values.
| Feature | What It Does | Output |
|---|---|---|
| Pre-training bias detection | Analyzes raw data before training | Bias metrics: CI, DPL, KL, JS |
| Post-training bias detection | Evaluates model predictions for bias | Metrics: DPPL, DI, DCO, RD |
| Model Explainability | SHAP values for feature importance | Feature weight contribution per prediction |
SHAP Explainability Example (Loan Approval):
Feature SHAP Value Contribution
─────────────────────────────────────────────
credit_score +0.42 ↑ approval
income +0.28 ↑ approval
debt_ratio -0.35 ↓ approval
employment_years +0.15 ↑ approval
age -0.02 minimal impact
5. A/B Testing with Production Variants
SageMaker Endpoints support Production Variants — run multiple model versions simultaneously with traffic splitting.
Endpoint with A/B Testing:
┌──────────────────────────────┐
Request ─→ SageMaker Endpoint │
│ │
│ Variant A (v1): 80% traffic │──→ Model v1 (current)
│ Variant B (v2): 20% traffic │──→ Model v2 (candidate)
└──────────────────────────────┘
↓
Compare metrics, shift traffic gradually
6. Cheat Sheet — Evaluation Metrics
| Scenario | Best Metric |
|---|---|
| Medical diagnosis (FN is critical) | Recall (Sensitivity) |
| Spam filter (FP is critical) | Precision |
| Imbalanced fraud detection | F1 Score, AUC-ROC |
| House price prediction (outliers matter) | RMSE |
| Demand forecasting (robust) | MAE |
| Explain individual prediction | SHAP (via SageMaker Clarify) |
7. Practice Questions
Q1: A hospital wants to build a model to detect early-stage cancer. Missing an actual cancer case is more dangerous than a false positive. Which metric should be OPTIMIZED?
- A) Precision
- B) Recall ✓
- C) Accuracy
- D) RMSE
Explanation: Recall = TP / (TP + FN). Optimizing Recall minimizes False Negatives (missed cancer cases), which is the critical concern here. Precision optimizes against False Positives, Accuracy is misleading for imbalanced medical data, and RMSE is for regression.
Q2: A company wants to gradually test a new model version in production while keeping the existing model as fallback. Which SageMaker feature provides this capability?
- A) SageMaker Experiments
- B) SageMaker Pipelines
- C) Production Variants on SageMaker Endpoints ✓
- D) SageMaker Model Monitor
Explanation: SageMaker Endpoints support Production Variants, allowing multiple model versions to run simultaneously with configurable traffic weights. This enables A/B testing and canary deployments without downtime.
Q3: A model for predicting house prices has RMSE=50,000 and MAE=20,000. This indicates the presence of what?
- A) High bias
- B) Data leakage
- C) Outliers driving up RMSE ✓
- D) Underfitting
Explanation: When RMSE is significantly higher than MAE, it indicates outliers — since RMSE squares errors, it penalizes large errors much more than MAE. The gap (50k vs 20k) suggests some predictions have very large errors (outliers in target variable).