Vertex AI Custom Training: Training Jobs, AutoML, distributed training, and optimization
1. Vertex AI Custom Training
Custom Training lets you run your own training code on Google Cloud infrastructure. There are 2 ways to package code:
| Method | Description | When to Use |
|---|---|---|
| Pre-built containers | GCP-provided containers: TF, PyTorch, Scikit-learn, XGBoost | Standard ML frameworks, fast setup |
| Custom containers | Build your own Docker image | Custom dependencies, special environments |
Custom Training Job Structure:
training_package/ (Python package or Docker image)
│
├── trainer/
│ ├── __init__.py
│ ├── task.py ← entry point (main training script)
│ └── model.py ← model definition
│
└── setup.py
Arguments passed via:
TRAINING_DATA_URI: gs://bucket/data/
TRAINING_OUTPUT_URI: gs://bucket/model/
Hyperparameters: --learning-rate=0.001
2. Compute Options
| Hardware | Best For | Notes |
|---|---|---|
| CPU | Scikit-learn, small tabular | Cheapest, no GPU parallelism |
| GPU (T4, A100, V100) | Deep learning, NLP, CV | 10-100x faster than CPU for DL |
| TPU v3, v4 | TensorFlow large-scale training | Google-specific; very fast for TF/JAX |
Exam tip: TPU is Google-specific hardware optimized for TensorFlow and JAX. GPUs work with all frameworks. TPUs are most cost-effective for very large TF models; GPUs are more versatile. The exam may ask "most cost-effective for TensorFlow large-scale" → TPU.
3. Distributed Training on Vertex AI
| Strategy | Description | Use Case |
|---|---|---|
| Data Parallelism | Split data across workers, same model | Most DL training scenarios |
| Model Parallelism | Split model layers across workers | Model too large for one GPU |
| MirroredStrategy (TF) | Multi-GPU, single machine | Single node, multiple GPUs |
| MultiWorkerMirroredStrategy | Multi-GPU, multi-machine | Cluster training |
| ParameterServerStrategy | Async updates via parameter server | Very large models (legacy) |
4. Vertex AI AutoML
| AutoML Type | Input Data | Supported Tasks |
|---|---|---|
| AutoML Tabular | CSV, BigQuery table | Classification, Regression, Forecasting |
| AutoML Image | JPEG, PNG, BMP | Classification (single/multi), Object Detection, Segmentation |
| AutoML Text | Text documents | Classification, Entity Extraction, Sentiment |
| AutoML Video | MP4, AVI, MOV | Classification, Object Detection, Action Recognition |
5. Vertex AI Hyperparameter Tuning
Vertex AI Hyperparameter Tuning automatically finds the best hyperparameter combinations.
| Search Algorithm | Description |
|---|---|
| Grid Search | Exhaustive, expensive; small search space |
| Random Search | Random sampling; often better than grid |
| Bayesian Optimization | Smart search using Gaussian Process; most efficient |
HPT Job Setup:
hyperparameters:
- parameter_id: learning_rate
type: DOUBLE
min_value: 0.0001
max_value: 0.1
scale: LOG ← log scale for LR
- parameter_id: batch_size
type: INTEGER
values: [32, 64, 128, 256]
metric:
metric_id: val_accuracy
goal: MAXIMIZE
max_trial_count: 50
parallel_trial_count: 5
6. Practice Questions
Q1: A team wants to train a custom TensorFlow model across multiple machines with 8 GPUs each. They want gradients synchronized across all workers without a parameter server. Which TensorFlow distribution strategy should they use?
- A) MirroredStrategy
- B) MultiWorkerMirroredStrategy ✓
- C) ParameterServerStrategy
- D) TPUStrategy
Explanation: MultiWorkerMirroredStrategy enables synchronous data-parallel training across multiple machines, each with multiple GPUs. MirroredStrategy is single-machine multi-GPU only. ParameterServerStrategy uses asynchronous updates. TPUStrategy is for TPU pods.
Q2: A company needs to train an image classification model but their team has no deep learning expertise. They have 5,000 labeled product images. Which Vertex AI option requires the LEAST ML expertise?
- A) Vertex AI Custom Training with TensorFlow CNN
- B) Vertex AI AutoML Image Classification ✓
- C) Dataproc Spark ML
- D) BigQuery ML
Explanation: AutoML Image Classification handles architecture selection, hyperparameter tuning, and training automatically. A team just needs to upload labeled images and specify the task. No code or deep learning expertise is required.
Q3: Which hyperparameter search strategy is MOST efficient when evaluating expensive-to-train deep learning models with a large search space?
- A) Grid Search — tests all combinations
- B) Random Search — samples uniformly
- C) Bayesian Optimization — uses past trial results to guide search ✓
- D) Manual tuning — expert selects parameters
Explanation: Bayesian Optimization builds a probabilistic model of the objective function using Gaussian Processes to intelligently select the next hyperparameter configuration to evaluate, based on past trial results. It finds good configurations with far fewer trials than grid or random search.