Chuyển đến nội dung chính

Lesson 2: Data Transformation & Feature Engineering

SageMaker Processing Jobs for data prep. SageMaker Feature Store. Handling missing values, encoding, normalization, scaling. Text preprocessing, imbalanced data techniques.

AWS ML Data Transformation Pipeline

Feature Engineering & Data Transformation: Glue, SageMaker Data Wrangler, and handling missing values

1. Data Transformation in the ML Pipeline

Before training a model, raw data must go through many transformation steps. This is the origin of the famous saying: "Garbage in, garbage out". The MLS-C01 exam frequently asks about data processing techniques and appropriate tools.

2. SageMaker Processing Jobs

SageMaker Processing Jobs is a managed service for running data processing scripts (Python, Spark) on ephemeral compute clusters.

Processor TypeFrameworkUse Case
ScriptProcessorCustom Docker containerAny custom script
SKLearnProcessorscikit-learnClassic ML preprocessing
PySparkProcessorApache SparkLarge-scale distributed processing
FrameworkProcessorTensorFlow/PyTorchDeep learning data prep
SageMaker Processing Job Flow:

S3 (input data)
      ↓
┌─────────────────────┐
│  Processing Job     │
│  (compute cluster)  │
│                     │
│  - Preprocess data  │
│  - Feature engineer │
│  - Split train/test │
└─────────────────────┘
      ↓
S3 (output: train/, validation/, test/)

3. Handling Missing Values

StrategyMethodWhen to Use
DeletionDrop rows/columnsMCAR, low missing rate (<5%)
Mean/Median ImputationFill with average valueNumeric, MCAR/MAR
Mode ImputationFill with most common valueCategorical
KNN ImputationUse K nearest neighborsPatterns in data, not too large
Model-based (MICE)Multiple imputationComplex missingness patterns
Indicator FeatureAdd is_missing columnWhen missingness carries information

Exam tip: Three types of missing data: MCAR (Missing Completely At Random) — deletion is safe; MAR (Missing At Random) — imputation is appropriate; MNAR (Missing Not At Random) — needs indicator feature or domain knowledge.

4. Categorical Encoding

EncodingMethodWhen to UseIssues
One-Hot EncodingBinary columns per categoryNominal (no order), few categoriesHigh cardinality → curse of dimensionality
Label Encoding0, 1, 2, 3...Ordinal (has ordering)Implies false order for nominal
Target EncodingMean of target per categoryHigh cardinality nominalData leakage risk if not careful
EmbeddingsDense vector representationText, high cardinalityNeeds sufficient data to learn

5. Normalization & Scaling

TechniqueFormulaOutput RangeBest For
Min-Max Normalization(x - min) / (max - min)[0, 1]Neural networks, distance-based
Standardization (Z-score)(x - mean) / stdMean=0, SD=1Linear models, SVM, PCA
Robust Scaler(x - median) / IQRCenteredWhen outliers are present
Log Transformlog(x)CompressedSkewed distributions

6. Handling Imbalanced Data

Class imbalance (e.g., fraud detection: 99% normal, 1% fraud) causes the model to be biased toward the majority class.

TechniqueMethodDirection
OversamplingDuplicate minority class samples↑ minority
SMOTESynthetic Minority Oversampling Technique — generate synthetic samples↑ minority
UndersamplingRemove majority class samples↓ majority
Class WeightsPenalize misclassification of minority moreNo data change
Ensemble MethodsBalancedBagging, EasyEnsembleAlgorithm-level

Exam tip: The right metric for imbalanced data: F1 Score, AUC-ROC, Precision-Recall — do NOT use Accuracy (misleading). AWS SageMaker Clarify can detect class imbalance.

7. SageMaker Feature Store

SageMaker Feature Store is a centralized repository for storing, sharing, and reusing ML features.

Feature Store Architecture:

          Feature Groups
         ┌──────────────────────────────┐
         │  user_features               │
         │  ┌──────┬────────┬────────┐  │
         │  │ id   │ age    │ recency│  │
         │  └──────┴────────┴────────┘  │
         └──────────────────────────────┘
               ↓ writes              ↑ reads
    ┌──────────────────┐   ┌──────────────────┐
    │  Offline Store   │   │  Online Store    │
    │  (S3 - training) │   │  (DynamoDB -     │
    │  batch reads     │   │  low-latency     │
    │                  │   │  inference)      │
    └──────────────────┘   └──────────────────┘

8. Cheat Sheet — Feature Engineering

ProblemSolution
High cardinality categoricalTarget encoding or embeddings
Missing values (numeric)Median imputation + indicator feature
Skewed distributionLog transform or Box-Cox
OutliersRobust Scaler or clip/winsorize
Imbalanced classesSMOTE + class weights + AUC metric
Reuse features across teamsSageMaker Feature Store

9. Practice Questions

Q1: A dataset for fraud detection has 98% negative (non-fraud) and 2% positive (fraud) examples. Which metric is MOST appropriate to evaluate the model?

  • A) Accuracy
  • B) R-squared
  • C) AUC-ROC ✓
  • D) Mean Absolute Error

Explanation: Accuracy is misleading for imbalanced data (predicting all negative gives 98% accuracy). AUC-ROC measures the model's ability to distinguish classes across all thresholds, making it ideal for imbalanced classification.

Q2: Which technique generates SYNTHETIC samples to address class imbalance?

  • A) Random undersampling
  • B) SMOTE (Synthetic Minority Oversampling Technique) ✓
  • C) Class weighting
  • D) Feature scaling

Explanation: SMOTE creates new synthetic samples for the minority class by interpolating between existing minority class examples, rather than just duplicating them.

Q3: A company wants to share engineered features between their training pipeline and real-time inference service. Which SageMaker feature addresses this?

  • A) SageMaker Processing Jobs
  • B) SageMaker Experiments
  • C) SageMaker Feature Store ✓
  • D) SageMaker Data Wrangler

Explanation: SageMaker Feature Store provides both an offline store (S3, for batch training) and online store (DynamoDB-backed, for low-latency real-time inference), ensuring feature consistency between training and serving.