Feature Engineering & Data Transformation: Glue, SageMaker Data Wrangler, and handling missing values
1. Data Transformation in the ML Pipeline
Before training a model, raw data must go through many transformation steps. This is the origin of the famous saying: "Garbage in, garbage out". The MLS-C01 exam frequently asks about data processing techniques and appropriate tools.
2. SageMaker Processing Jobs
SageMaker Processing Jobs is a managed service for running data processing scripts (Python, Spark) on ephemeral compute clusters.
| Processor Type | Framework | Use Case |
|---|---|---|
| ScriptProcessor | Custom Docker container | Any custom script |
| SKLearnProcessor | scikit-learn | Classic ML preprocessing |
| PySparkProcessor | Apache Spark | Large-scale distributed processing |
| FrameworkProcessor | TensorFlow/PyTorch | Deep learning data prep |
SageMaker Processing Job Flow:
S3 (input data)
↓
┌─────────────────────┐
│ Processing Job │
│ (compute cluster) │
│ │
│ - Preprocess data │
│ - Feature engineer │
│ - Split train/test │
└─────────────────────┘
↓
S3 (output: train/, validation/, test/)
3. Handling Missing Values
| Strategy | Method | When to Use |
|---|---|---|
| Deletion | Drop rows/columns | MCAR, low missing rate (<5%) |
| Mean/Median Imputation | Fill with average value | Numeric, MCAR/MAR |
| Mode Imputation | Fill with most common value | Categorical |
| KNN Imputation | Use K nearest neighbors | Patterns in data, not too large |
| Model-based (MICE) | Multiple imputation | Complex missingness patterns |
| Indicator Feature | Add is_missing column | When missingness carries information |
Exam tip: Three types of missing data: MCAR (Missing Completely At Random) — deletion is safe; MAR (Missing At Random) — imputation is appropriate; MNAR (Missing Not At Random) — needs indicator feature or domain knowledge.
4. Categorical Encoding
| Encoding | Method | When to Use | Issues |
|---|---|---|---|
| One-Hot Encoding | Binary columns per category | Nominal (no order), few categories | High cardinality → curse of dimensionality |
| Label Encoding | 0, 1, 2, 3... | Ordinal (has ordering) | Implies false order for nominal |
| Target Encoding | Mean of target per category | High cardinality nominal | Data leakage risk if not careful |
| Embeddings | Dense vector representation | Text, high cardinality | Needs sufficient data to learn |
5. Normalization & Scaling
| Technique | Formula | Output Range | Best For |
|---|---|---|---|
| Min-Max Normalization | (x - min) / (max - min) | [0, 1] | Neural networks, distance-based |
| Standardization (Z-score) | (x - mean) / std | Mean=0, SD=1 | Linear models, SVM, PCA |
| Robust Scaler | (x - median) / IQR | Centered | When outliers are present |
| Log Transform | log(x) | Compressed | Skewed distributions |
6. Handling Imbalanced Data
Class imbalance (e.g., fraud detection: 99% normal, 1% fraud) causes the model to be biased toward the majority class.
| Technique | Method | Direction |
|---|---|---|
| Oversampling | Duplicate minority class samples | ↑ minority |
| SMOTE | Synthetic Minority Oversampling Technique — generate synthetic samples | ↑ minority |
| Undersampling | Remove majority class samples | ↓ majority |
| Class Weights | Penalize misclassification of minority more | No data change |
| Ensemble Methods | BalancedBagging, EasyEnsemble | Algorithm-level |
Exam tip: The right metric for imbalanced data: F1 Score, AUC-ROC, Precision-Recall — do NOT use Accuracy (misleading). AWS SageMaker Clarify can detect class imbalance.
7. SageMaker Feature Store
SageMaker Feature Store is a centralized repository for storing, sharing, and reusing ML features.
Feature Store Architecture:
Feature Groups
┌──────────────────────────────┐
│ user_features │
│ ┌──────┬────────┬────────┐ │
│ │ id │ age │ recency│ │
│ └──────┴────────┴────────┘ │
└──────────────────────────────┘
↓ writes ↑ reads
┌──────────────────┐ ┌──────────────────┐
│ Offline Store │ │ Online Store │
│ (S3 - training) │ │ (DynamoDB - │
│ batch reads │ │ low-latency │
│ │ │ inference) │
└──────────────────┘ └──────────────────┘
8. Cheat Sheet — Feature Engineering
| Problem | Solution |
|---|---|
| High cardinality categorical | Target encoding or embeddings |
| Missing values (numeric) | Median imputation + indicator feature |
| Skewed distribution | Log transform or Box-Cox |
| Outliers | Robust Scaler or clip/winsorize |
| Imbalanced classes | SMOTE + class weights + AUC metric |
| Reuse features across teams | SageMaker Feature Store |
9. Practice Questions
Q1: A dataset for fraud detection has 98% negative (non-fraud) and 2% positive (fraud) examples. Which metric is MOST appropriate to evaluate the model?
- A) Accuracy
- B) R-squared
- C) AUC-ROC ✓
- D) Mean Absolute Error
Explanation: Accuracy is misleading for imbalanced data (predicting all negative gives 98% accuracy). AUC-ROC measures the model's ability to distinguish classes across all thresholds, making it ideal for imbalanced classification.
Q2: Which technique generates SYNTHETIC samples to address class imbalance?
- A) Random undersampling
- B) SMOTE (Synthetic Minority Oversampling Technique) ✓
- C) Class weighting
- D) Feature scaling
Explanation: SMOTE creates new synthetic samples for the minority class by interpolating between existing minority class examples, rather than just duplicating them.
Q3: A company wants to share engineered features between their training pipeline and real-time inference service. Which SageMaker feature addresses this?
- A) SageMaker Processing Jobs
- B) SageMaker Experiments
- C) SageMaker Feature Store ✓
- D) SageMaker Data Wrangler
Explanation: SageMaker Feature Store provides both an offline store (S3, for batch training) and online store (DynamoDB-backed, for low-latency real-time inference), ensuring feature consistency between training and serving.