GCP AI/ML Ecosystem: Vertex AI, AutoML, BigQuery ML, Pre-trained APIs and when to use each
1. GCP ML Landscape Overview
GCP ML Capability Spectrum:
LOW CODE ◄────────────────────────────────────► HIGH CONTROL
│ │ │
▼ ▼ ▼
Pre-trained APIs Vertex AI AutoML Custom Training
(Vision, NLP, (no code needed, (full control,
Translation) you bring data) you bring code)
│ │ │
No ML expertise Some domain ML expertise
needed expertise required
BigQuery ML ────── SQL interface for ML on warehouse data
2. Vertex AI — Unified ML Platform
Vertex AI is GCP's unified platform for the entire ML lifecycle. Understanding its components is essential for the exam.
| Component | Purpose |
|---|---|
| Vertex AI Workbench | Managed Jupyter notebooks for data scientists |
| Vertex AI Training | Custom training jobs (CPUs, GPUs, TPUs) |
| Vertex AI AutoML | No-code model training (Tabular, Image, Text, Video) |
| Vertex AI Endpoints | Deploy models for online prediction |
| Vertex AI Batch Prediction | Asynchronous batch scoring |
| Vertex AI Feature Store | Serve features consistently across training/serving |
| Vertex AI Pipelines | Kubeflow Pipelines-based ML workflow orchestration |
| Vertex AI Experiments | Track runs, compare metrics |
| Vertex AI Model Registry | Version control for models |
| Vertex AI Model Monitoring | Detect feature skew and prediction drift |
3. AutoML vs. Custom Training
| Criteria | AutoML | Custom Training |
|---|---|---|
| ML expertise needed | Minimal | Required |
| Training time | Hours (automated) | Variable (you control) |
| Model interpretability | Limited | Full control |
| Cost | Higher per model | Pay per compute used |
| Best for | Quick prototypes, standard tasks | Custom architectures, research |
| Supported data types | Tabular, Image, Text, Video | Any (you write the code) |
Exam tip: Questions with "team doesn't have ML expertise" or "fastest time to deployment" → AutoML. Questions with "custom neural architecture" or "full control over training loop" → Custom Training.
4. BigQuery ML
BigQuery ML lets you train and serve ML models using SQL — no need to export data from BigQuery.
| Model Type | SQL Keyword | Use Case |
|---|---|---|
| Linear Regression | LINEAR_REG | Price prediction |
| Logistic Regression | LOGISTIC_REG | Classification |
| K-Means Clustering | KMEANS | Customer segmentation |
| XGBoost | BOOSTED_TREE_CLASSIFIER/REGRESSOR | Tabular classification/regression |
| Deep Neural Network | DNN_CLASSIFIER/DNN_REGRESSOR | Complex patterns |
| Matrix Factorization | MATRIX_FACTORIZATION | Recommendations |
| Imported TF models | TENSORFLOW | Custom TF models |
5. Pre-trained AI APIs
| API | Capabilities | Use Case |
|---|---|---|
| Cloud Vision API | Labels, OCR, faces, logos, safe search | Image analysis without training |
| Cloud Natural Language API | Entities, sentiment, syntax, categories | Text analytics |
| Cloud Translation API | 100+ language pairs | Multi-language content |
| Cloud Speech-to-Text | Transcription, speaker diarization | Audio processing |
| Cloud Text-to-Speech | WaveNet voices, SSML | Voice UI, accessibility |
| Document AI | Form parsing, invoice extraction | Document automation |
| Recommendations AI | Real-time product recommendations | E-commerce personalization |
6. Service Selection Decision Tree
WHICH GCP ML SERVICE?
Do you have LABELED DATA?
│
├── NO → Pre-trained API sufficient for your task (Vision, NLP)?
│ YES → Use Pre-trained API
│ NO → Vertex AI Custom Training (unsupervised)
│
└── YES → Is your data already IN BigQuery?
│
├── YES → BigQuery ML (SQL-based, fast, no export)
│
└── NO → Need rapid prototyping, no ML team?
│
├── YES → Vertex AI AutoML
│
└── NO → Vertex AI Custom Training
7. Practice Questions
Q1: A data analytics team has petabytes of customer transaction data in BigQuery. They want to build a churn prediction model using their existing SQL skills without data exports. Which approach is BEST?
- A) Export to Cloud Storage, then use Vertex AI Custom Training
- B) Use Cloud Natural Language API
- C) Use BigQuery ML with CREATE MODEL LOGISTIC_REGRESSION ✓
- D) Use Vertex AI AutoML Tabular
Explanation: BigQuery ML allows training classification models directly on BigQuery data using SQL, leveraging existing data infrastructure and skills without exporting data. This is the fastest path when data is already in BigQuery.
Q2: A small startup needs to add sentiment analysis to customer reviews. They have no ML team and no labeled sentiment data. Which solution requires the LEAST effort?
- A) Vertex AI AutoML Text Sentiment
- B) Train a custom BERT model on Vertex AI
- C) Cloud Natural Language API sentiment analysis ✓
- D) BigQuery ML DNN classifier
Explanation: Cloud Natural Language API is a pre-trained, fully managed service that requires no training data, no ML expertise, and no infrastructure setup. Just call the API. AutoML requires labeled sentiment examples; custom BERT requires significantly more expertise.
Q3: Which Vertex AI component should a team use to ensure that feature values used during model training are identical to those served at prediction time?
- A) Vertex AI Experiments
- B) Vertex AI Feature Store ✓
- C) Vertex AI Model Registry
- D) Vertex AI Pipelines
Explanation: Vertex AI Feature Store provides a centralized repository for storing, serving, and sharing ML features. It ensures training-serving consistency by using the same feature definitions and values for both training and online/batch prediction, preventing training-serving skew.