Chuyển đến nội dung chính

Lesson 5: Vertex AI Training — Custom & AutoML

Custom Training Jobs: pre-built containers, custom containers. Distributed training on GPU/TPU. AutoML: Tabular, Image, Text, Video. Training pipeline setup. Hyperparameter tuning service.

Vertex AI Custom Training

Vertex AI Custom Training: Training Jobs, AutoML, distributed training, and optimization

1. Vertex AI Custom Training

Custom Training lets you run your own training code on Google Cloud infrastructure. There are 2 ways to package code:

MethodDescriptionWhen to Use
Pre-built containersGCP-provided containers: TF, PyTorch, Scikit-learn, XGBoostStandard ML frameworks, fast setup
Custom containersBuild your own Docker imageCustom dependencies, special environments
Custom Training Job Structure:

training_package/ (Python package or Docker image)
│
├── trainer/
│   ├── __init__.py
│   ├── task.py        ← entry point (main training script)
│   └── model.py       ← model definition
│
└── setup.py

Arguments passed via:
  TRAINING_DATA_URI: gs://bucket/data/
  TRAINING_OUTPUT_URI: gs://bucket/model/
  Hyperparameters: --learning-rate=0.001

2. Compute Options

HardwareBest ForNotes
CPUScikit-learn, small tabularCheapest, no GPU parallelism
GPU (T4, A100, V100)Deep learning, NLP, CV10-100x faster than CPU for DL
TPU v3, v4TensorFlow large-scale trainingGoogle-specific; very fast for TF/JAX

Exam tip: TPU is Google-specific hardware optimized for TensorFlow and JAX. GPUs work with all frameworks. TPUs are most cost-effective for very large TF models; GPUs are more versatile. The exam may ask "most cost-effective for TensorFlow large-scale" → TPU.

3. Distributed Training on Vertex AI

StrategyDescriptionUse Case
Data ParallelismSplit data across workers, same modelMost DL training scenarios
Model ParallelismSplit model layers across workersModel too large for one GPU
MirroredStrategy (TF)Multi-GPU, single machineSingle node, multiple GPUs
MultiWorkerMirroredStrategyMulti-GPU, multi-machineCluster training
ParameterServerStrategyAsync updates via parameter serverVery large models (legacy)

4. Vertex AI AutoML

AutoML TypeInput DataSupported Tasks
AutoML TabularCSV, BigQuery tableClassification, Regression, Forecasting
AutoML ImageJPEG, PNG, BMPClassification (single/multi), Object Detection, Segmentation
AutoML TextText documentsClassification, Entity Extraction, Sentiment
AutoML VideoMP4, AVI, MOVClassification, Object Detection, Action Recognition

5. Vertex AI Hyperparameter Tuning

Vertex AI Hyperparameter Tuning automatically finds the best hyperparameter combinations.

Search AlgorithmDescription
Grid SearchExhaustive, expensive; small search space
Random SearchRandom sampling; often better than grid
Bayesian OptimizationSmart search using Gaussian Process; most efficient
HPT Job Setup:

hyperparameters:
  - parameter_id: learning_rate
    type: DOUBLE
    min_value: 0.0001
    max_value: 0.1
    scale: LOG  ← log scale for LR

  - parameter_id: batch_size
    type: INTEGER
    values: [32, 64, 128, 256]

metric:
  metric_id: val_accuracy
  goal: MAXIMIZE
  
max_trial_count: 50
parallel_trial_count: 5

6. Practice Questions

Q1: A team wants to train a custom TensorFlow model across multiple machines with 8 GPUs each. They want gradients synchronized across all workers without a parameter server. Which TensorFlow distribution strategy should they use?

  • A) MirroredStrategy
  • B) MultiWorkerMirroredStrategy ✓
  • C) ParameterServerStrategy
  • D) TPUStrategy

Explanation: MultiWorkerMirroredStrategy enables synchronous data-parallel training across multiple machines, each with multiple GPUs. MirroredStrategy is single-machine multi-GPU only. ParameterServerStrategy uses asynchronous updates. TPUStrategy is for TPU pods.

Q2: A company needs to train an image classification model but their team has no deep learning expertise. They have 5,000 labeled product images. Which Vertex AI option requires the LEAST ML expertise?

  • A) Vertex AI Custom Training with TensorFlow CNN
  • B) Vertex AI AutoML Image Classification ✓
  • C) Dataproc Spark ML
  • D) BigQuery ML

Explanation: AutoML Image Classification handles architecture selection, hyperparameter tuning, and training automatically. A team just needs to upload labeled images and specify the task. No code or deep learning expertise is required.

Q3: Which hyperparameter search strategy is MOST efficient when evaluating expensive-to-train deep learning models with a large search space?

  • A) Grid Search — tests all combinations
  • B) Random Search — samples uniformly
  • C) Bayesian Optimization — uses past trial results to guide search ✓
  • D) Manual tuning — expert selects parameters

Explanation: Bayesian Optimization builds a probabilistic model of the objective function using Gaussian Processes to intelligently select the next hyperparameter configuration to evaluate, based on past trial results. It finds good configurations with far fewer trials than grid or random search.