SageMaker Training Jobs & Hyperparameter Tuning: distributed training, Spot Instances, và HPO strategies
1. SageMaker Training Jobs
SageMaker Training Jobs chạy ML training code trên managed compute infrastructure. Training xảy ra trên ephemeral instances — chỉ tính phí khi chạy.
Training Job Lifecycle:
Submit Job ──→ Provision Instances ──→ Download Data
↓
Run Training Code
↓
Save Model to S3
↓
Terminate Instances
2. Instance Types cho Training
| Instance Family | Hardware | Best For |
|---|---|---|
| ml.c5 | CPU optimized | Tabular ML, XGBoost, sklearn |
| ml.m5 | General purpose CPU | Light training, data processing |
| ml.p3 | V100 GPU | Deep learning training |
| ml.p4d | A100 GPU (8x) | Large-scale DL, distributed training |
| ml.g4dn | T4 GPU (cost-effective) | Small-medium DL models |
| ml.trn1 | AWS Trainium | LLM training, cost optimization |
3. Distributed Training
Khi model hoặc dataset quá lớn cho một instance, cần distributed training trên nhiều instances.
| Strategy | How It Works | When to Use |
|---|---|---|
| Data Parallelism | Mỗi instance có copy của model, train trên subset của data, sync gradients | Dataset quá lớn, model vừa vặn trong 1 GPU |
| Model Parallelism | Model split across instances, mỗi instance chứa 1 phần | Model quá lớn cho 1 GPU (LLMs) |
Data Parallelism:
Instance 1 [Full Model] ──→ Train on data shard A ──→ ↓
Instance 2 [Full Model] ──→ Train on data shard B ──→ ↓ AllReduce
Instance 3 [Full Model] ──→ Train on data shard C ──→ ↓ (sync gradients)
↓
Updated Model Weights
Model Parallelism:
Instance 1 [Layers 1-4] ──→ forward pass ──→
Instance 2 [Layers 5-8] ──→ forward pass ──→
Instance 3 [Layers 9-12] ──→ forward pass ──→ output
Exam tip: SageMaker cung cấp SageMaker Distributed library với 2 modules: (1)
smdistributed.dataparallel— optimized AllReduce; (2)smdistributed.modelparallel— auto pipeline parallelism. Khi đề hỏi "large model training" → model parallelism.
4. Automatic Model Tuning (HPO)
Hyperparameter Optimization (HPO) tự động tìm hyperparameters tốt nhất bằng cách chạy nhiều training jobs với configs khác nhau.
| Strategy | How It Works | Tradeoff |
|---|---|---|
| Random Search | Randomly sample hyperparameters từ range | Fast, good baseline |
| Grid Search | Try all combinations | Exhaustive, expensive, bad for large spaces |
| Bayesian Optimization | Probabilistic model của outcome, suggest best next config | Efficient, learns from previous trials — SageMaker default |
| Hyperband | Early-stop poorly performing trials | Resource-efficient, fast |
Exam tip: SageMaker AMT (Automatic Model Tuning) dùng Bayesian Optimization by default. Nó XEM KẾT QUẢ từ các jobs trước để suggest next hyperparameter set — intelligent search, không phải brute force.
5. Spot Instance Training
SageMaker hỗ trợ dùng EC2 Spot Instances cho training jobs, tiết kiệm đến 90% chi phí so với On-Demand.
| Feature | Detail |
|---|---|
| MaxWaitTimeInSeconds | Maximum thời gian đợi spot capacity |
| Checkpointing | Lưu model to S3 periodically — resume sau khi bị interrupt |
| use_spot_instances=True | Parameter trong SageMaker Estimator |
Exam tip: Khi đề hỏi "reduce training costs", đáp án thường là Spot Instances với checkpointing. Checkpointing quan trọng để tránh mất progress khi spot instance bị terminate.
6. Bias-Variance Tradeoff
| Issue | Symptom | Cause | Solution |
|---|---|---|---|
| High Bias (Underfitting) | High train error, high test error | Model quá đơn giản | Tăng model complexity, thêm features, giảm regularization |
| High Variance (Overfitting) | Low train error, high test error | Model quá phức tạp | Thêm data, dropout, regularization, feature selection |
| Balanced | Low train error, low test error (gần nhau) | Good fit | Deploy model |
7. Practice Questions
Q1: A company is training a large deep learning model that doesn't fit on a single GPU instance. Which SageMaker distributed training strategy should they use?
- A) Data parallelism
- B) Model parallelism ✓
- C) Pipeline parallelism only
- D) Increase batch size
Explanation: Model parallelism splits the model itself across multiple GPU instances, allowing training of models too large to fit in a single GPU's memory. Data parallelism keeps a full model copy on each instance, which doesn't help when the model itself is too large.
Q2: A team wants to minimize the cost of running 500 hyperparameter tuning jobs. Training can tolerate interruptions. What is the MOST cost-effective approach?
- A) Use larger instances to run jobs faster
- B) Use Spot Instances with checkpointing enabled ✓
- C) Use Grid Search instead of Bayesian Optimization
- D) Reduce the number of epochs
Explanation: Spot Instances can save up to 90% compared to On-Demand pricing. With checkpointing enabled, interrupted jobs save their state to S3 and can resume, making Spot Instances practical for long HPO jobs.
Q3: A model achieves 95% accuracy on training data but only 62% on the test set. What problem does this indicate?
- A) Underfitting / High bias
- B) Overfitting / High variance ✓
- C) Data leakage
- D) Class imbalance
Explanation: The large gap between training accuracy (95%) and test accuracy (62%) is a classic sign of overfitting (high variance). The model memorized the training data but fails to generalize. Solutions: more data, regularization (L1/L2, dropout), reduce model complexity.