Feature Engineering & Data Transformation: Glue, SageMaker Data Wrangler, và xử lý missing values
1. Data Transformation trong ML Pipeline
Trước khi train model, raw data phải qua nhiều bước transformation. Đây là nguồn gốc của câu nói nổi tiếng: "Garbage in, garbage out". Đề thi MLS-C01 thường hỏi kỹ thuật xử lý data và tools phù hợp.
2. SageMaker Processing Jobs
SageMaker Processing Jobs là managed service để chạy data processing scripts (Python, Spark) trên ephemeral compute clusters.
| Processor Type | Framework | Use Case |
|---|---|---|
| ScriptProcessor | Custom Docker container | Any custom script |
| SKLearnProcessor | scikit-learn | Classic ML preprocessing |
| PySparkProcessor | Apache Spark | Large-scale distributed processing |
| FrameworkProcessor | TensorFlow/PyTorch | Deep learning data prep |
SageMaker Processing Job Flow:
S3 (input data)
↓
┌─────────────────────┐
│ Processing Job │
│ (compute cluster) │
│ │
│ - Preprocess data │
│ - Feature engineer │
│ - Split train/test │
└─────────────────────┘
↓
S3 (output: train/, validation/, test/)
3. Xử lý Missing Values
| Strategy | Method | When to Use |
|---|---|---|
| Deletion | Drop rows/columns | MCAR, ít missing (<5%) |
| Mean/Median Imputation | Điền giá trị trung bình | Numeric, MCAR/MAR |
| Mode Imputation | Điền giá trị phổ biến nhất | Categorical |
| KNN Imputation | Dùng K neighbors gần nhất | Patterns in data, không quá lớn |
| Model-based (MICE) | Multiple imputation | Complex missingness patterns |
| Indicator Feature | Thêm cột is_missing | Khi missingness chứa thông tin |
Exam tip: Ba loại missing data: MCAR (Missing Completely At Random) — deletion an toàn; MAR (Missing At Random) — imputation phù hợp; MNAR (Missing Not At Random) — cần indicator feature hoặc domain knowledge.
4. Categorical Encoding
| Encoding | Method | When to Use | Issues |
|---|---|---|---|
| One-Hot Encoding | Binary columns mỗi category | Nominal (no order), ít categories | High cardinality → curse of dimensionality |
| Label Encoding | 0, 1, 2, 3... | Ordinal (có thứ tự) | Implies false order for nominal |
| Target Encoding | Mean of target per category | High cardinality nominal | Data leakage risk nếu không cẩn thận |
| Embeddings | Dense vector representation | Text, high cardinality | Cần đủ data để learn |
5. Normalization & Scaling
| Technique | Formula | Output Range | Best For |
|---|---|---|---|
| Min-Max Normalization | (x - min) / (max - min) | [0, 1] | Neural networks, distance-based |
| Standardization (Z-score) | (x - mean) / std | Mean=0, SD=1 | Linear models, SVM, PCA |
| Robust Scaler | (x - median) / IQR | Centered | Outliers present |
| Log Transform | log(x) | Compressed | Skewed distributions |
6. Xử lý Imbalanced Data
Class imbalance (e.g., fraud detection: 99% normal, 1% fraud) khiến model bias về majority class.
| Technique | Method | Direction |
|---|---|---|
| Oversampling | Duplicate minority class samples | ↑ minority |
| SMOTE | Synthetic Minority Oversampling Technique — generate synthetic samples | ↑ minority |
| Undersampling | Remove majority class samples | ↓ majority |
| Class Weights | Penalize misclassification of minority more | No data change |
| Ensemble Methods | BalancedBagging, EasyEnsemble | Algorithm-level |
Exam tip: Metric phù hợp cho imbalanced data: F1 Score, AUC-ROC, Precision-Recall — KHÔNG dùng Accuracy (misleading). AWS SageMaker Clarify có thể detect class imbalance.
7. SageMaker Feature Store
SageMaker Feature Store là centralized repository để store, share và reuse ML features.
Feature Store Architecture:
Feature Groups
┌──────────────────────────────┐
│ user_features │
│ ┌──────┬────────┬────────┐ │
│ │ id │ age │ recency│ │
│ └──────┴────────┴────────┘ │
└──────────────────────────────┘
↓ writes ↑ reads
┌──────────────────┐ ┌──────────────────┐
│ Offline Store │ │ Online Store │
│ (S3 - training) │ │ (DynamoDB - │
│ batch reads │ │ low-latency │
│ │ │ inference) │
└──────────────────┘ └──────────────────┘
8. Cheat Sheet — Feature Engineering
| Problem | Solution |
|---|---|
| High cardinality categorical | Target encoding hoặc embeddings |
| Missing values (numeric) | Median imputation + indicator feature |
| Skewed distribution | Log transform hoặc Box-Cox |
| Outliers | Robust Scaler hoặc clip/winsorize |
| Imbalanced classes | SMOTE + class weights + AUC metric |
| Reuse features across teams | SageMaker Feature Store |
9. Practice Questions
Q1: A dataset for fraud detection has 98% negative (non-fraud) and 2% positive (fraud) examples. Which metric is MOST appropriate to evaluate the model?
- A) Accuracy
- B) R-squared
- C) AUC-ROC ✓
- D) Mean Absolute Error
Explanation: Accuracy is misleading for imbalanced data (predicting all negative gives 98% accuracy). AUC-ROC measures the model's ability to distinguish classes across all thresholds, making it ideal for imbalanced classification.
Q2: Which technique generates SYNTHETIC samples to address class imbalance?
- A) Random undersampling
- B) SMOTE (Synthetic Minority Oversampling Technique) ✓
- C) Class weighting
- D) Feature scaling
Explanation: SMOTE creates new synthetic samples for the minority class by interpolating between existing minority class examples, rather than just duplicating them.
Q3: A company wants to share engineered features between their training pipeline and real-time inference service. Which SageMaker feature addresses this?
- A) SageMaker Processing Jobs
- B) SageMaker Experiments
- C) SageMaker Feature Store ✓
- D) SageMaker Data Wrangler
Explanation: SageMaker Feature Store provides both an offline store (S3, for batch training) and online store (DynamoDB-backed, for low-latency real-time inference), ensuring feature consistency between training and serving.