Chuyển đến nội dung chính

Lesson 5: Training & Hyperparameter Tuning

SageMaker Training Jobs: instance types, Pipe Mode vs File Mode. Distributed training: data parallelism vs model parallelism. Automatic Model Tuning (HPO): Bayesian vs Random vs Grid search. Spot Instance Training to reduce costs.

SageMaker Training & Hyperparameter Tuning

SageMaker Training Jobs & Hyperparameter Tuning: distributed training, Spot Instances, and HPO strategies

1. SageMaker Training Jobs

SageMaker Training Jobs run ML training code on managed compute infrastructure. Training runs on ephemeral instances — you only pay while they're running.

Training Job Lifecycle:

  Submit Job ──→ Provision Instances ──→ Download Data
                                              ↓
                                       Run Training Code
                                              ↓
                                       Save Model to S3
                                              ↓
                                       Terminate Instances  

2. Instance Types for Training

Instance FamilyHardwareBest For
ml.c5CPU optimizedTabular ML, XGBoost, sklearn
ml.m5General purpose CPULight training, data processing
ml.p3V100 GPUDeep learning training
ml.p4dA100 GPU (8x)Large-scale DL, distributed training
ml.g4dnT4 GPU (cost-effective)Small-medium DL models
ml.trn1AWS TrainiumLLM training, cost optimization

3. Distributed Training

When the model or dataset is too large for a single instance, distributed training across multiple instances is needed.

StrategyHow It WorksWhen to Use
Data ParallelismEach instance has a copy of the model, trains on a subset of data, syncs gradientsDataset too large, model fits in one GPU
Model ParallelismModel split across instances, each holds a portionModel too large for a single GPU (LLMs)
Data Parallelism:

Instance 1 [Full Model] ──→ Train on data shard A ──→ ↓
Instance 2 [Full Model] ──→ Train on data shard B ──→ ↓  AllReduce
Instance 3 [Full Model] ──→ Train on data shard C ──→ ↓  (sync gradients)
                                                          ↓
                                              Updated Model Weights

Model Parallelism:

Instance 1 [Layers 1-4]  ──→ forward pass ──→
Instance 2 [Layers 5-8]  ──→ forward pass ──→
Instance 3 [Layers 9-12] ──→ forward pass ──→ output

Exam tip: SageMaker provides the SageMaker Distributed library with 2 modules: (1) smdistributed.dataparallel — optimized AllReduce; (2) smdistributed.modelparallel — auto pipeline parallelism. When the question asks "large model training" → model parallelism.

4. Automatic Model Tuning (HPO)

Hyperparameter Optimization (HPO) automatically finds the best hyperparameters by running multiple training jobs with different configurations.

StrategyHow It WorksTradeoff
Random SearchRandomly sample hyperparameters from rangeFast, good baseline
Grid SearchTry all combinationsExhaustive, expensive, bad for large spaces
Bayesian OptimizationProbabilistic model of outcome, suggests best next configEfficient, learns from previous trials — SageMaker default
HyperbandEarly-stop poorly performing trialsResource-efficient, fast

Exam tip: SageMaker AMT (Automatic Model Tuning) uses Bayesian Optimization by default. It EXAMINES RESULTS from previous jobs to suggest the next hyperparameter set — intelligent search, not brute force.

5. Spot Instance Training

SageMaker supports using EC2 Spot Instances for training jobs, saving up to 90% in cost compared to On-Demand.

FeatureDetail
MaxWaitTimeInSecondsMaximum time to wait for spot capacity
CheckpointingSaves model to S3 periodically — resume after interruption
use_spot_instances=TrueParameter in SageMaker Estimator

Exam tip: When the question asks "reduce training costs", the answer is usually Spot Instances with checkpointing. Checkpointing is essential to avoid losing progress when spot instances are terminated.

6. Bias-Variance Tradeoff

IssueSymptomCauseSolution
High Bias (Underfitting)High train error, high test errorModel too simpleIncrease model complexity, add features, reduce regularization
High Variance (Overfitting)Low train error, high test errorModel too complexAdd more data, dropout, regularization, feature selection
BalancedLow train error, low test error (close to each other)Good fitDeploy model

7. Practice Questions

Q1: A company is training a large deep learning model that doesn't fit on a single GPU instance. Which SageMaker distributed training strategy should they use?

  • A) Data parallelism
  • B) Model parallelism ✓
  • C) Pipeline parallelism only
  • D) Increase batch size

Explanation: Model parallelism splits the model itself across multiple GPU instances, allowing training of models too large to fit in a single GPU's memory. Data parallelism keeps a full model copy on each instance, which doesn't help when the model itself is too large.

Q2: A team wants to minimize the cost of running 500 hyperparameter tuning jobs. Training can tolerate interruptions. What is the MOST cost-effective approach?

  • A) Use larger instances to run jobs faster
  • B) Use Spot Instances with checkpointing enabled ✓
  • C) Use Grid Search instead of Bayesian Optimization
  • D) Reduce the number of epochs

Explanation: Spot Instances can save up to 90% compared to On-Demand pricing. With checkpointing enabled, interrupted jobs save their state to S3 and can resume, making Spot Instances practical for long HPO jobs.

Q3: A model achieves 95% accuracy on training data but only 62% on the test set. What problem does this indicate?

  • A) Underfitting / High bias
  • B) Overfitting / High variance ✓
  • C) Data leakage
  • D) Class imbalance

Explanation: The large gap between training accuracy (95%) and test accuracy (62%) is a classic sign of overfitting (high variance). The model memorized the training data but fails to generalize. Solutions: more data, regularization (L1/L2, dropout), reduce model complexity.