Vertex AI MLOps: Pipelines, CI/CD, Model Registry, and monitoring for production ML
1. MLOps Maturity Levels
| Level | Description | Automation |
|---|---|---|
| Level 0 | Manual process, scripts only | None |
| Level 1 | ML pipeline automation, continuous training | Training pipeline |
| Level 2 | Full CI/CD for ML, automated retraining triggers | Everything |
2. Vertex AI Pipelines
Vertex AI Pipelines is a managed execution environment for Kubeflow Pipelines (KFP). Pipelines are defined using the Python SDK and compiled to YAML.
Vertex AI Pipeline Structure:
@component (preprocess_data)
↓
@component (train_model)
↓
@component (evaluate_model)
↓ (if accuracy > threshold)
@component (deploy_model)
Each component = isolated Docker container
Artifacts (data, models) stored in Cloud Storage
Metadata tracked in Vertex ML Metadata Store
| Pipeline SDK | Notes |
|---|---|
| Kubeflow Pipelines SDK v2 | Primary SDK for Vertex AI Pipelines |
| TFX | TensorFlow-specific pipeline components |
| Google Cloud Pipeline Components | Pre-built components for Vertex AI services |
3. Vertex AI Model Monitoring
| Monitoring Type | What It Detects |
|---|---|
| Feature Skew Monitoring | Serving feature distribution ≠ training baseline |
| Feature Drift Monitoring | Serving feature distribution changes over time |
| Prediction Drift | Model output distribution changes (indirect label drift) |
Model Monitoring Workflow:
Training Data Baseline (BigQuery/GCS)
↓ (establish distribution)
Deploy to Endpoint with Monitoring enabled
↓ (collect serving requests)
Periodic Analysis (hourly/daily)
↓ (compare distributions)
Alert if skew/drift > threshold
↓
Retrain trigger → new Pipeline run
4. Vertex AI Experiments & Metadata
| Component | Purpose |
|---|---|
| Vertex AI Experiments | Track hyperparameters, metrics, artifacts across runs |
| ML Metadata Store | Track lineage: data → model → endpoint |
| Vertex AI TensorBoard | Visualize training metrics (loss, accuracy curves) |
5. CI/CD for ML on GCP
ML CI/CD Pipeline on GCP:
Code Push to Cloud Source Repositories
↓
Cloud Build trigger (CI)
├── Unit tests for ML components
├── Data validation tests
└── Build Docker image → push to Artifact Registry
↓
Vertex AI Pipeline trigger (CD/CT)
├── Data preprocessing
├── Model training
├── Model evaluation
└── Conditional deployment → Vertex AI Endpoint
Exam tip: CI/CD for ML = Cloud Build (code testing + Docker build) + Vertex AI Pipelines (training + deployment orchestration). Cloud Source Repositories is GCP's Git hosting. Artifact Registry replaces Container Registry for storing Docker images.
6. Practice Questions
Q1: A production ML model's prediction distribution has shifted significantly over 3 weeks, but ground truth labels are not yet available to measure accuracy directly. Which Vertex AI monitoring type detects this?
- A) Feature Skew Monitoring
- B) Prediction Drift Monitoring ✓
- C) Training data validation
- D) Vertex AI Experiments baseline comparison
Explanation: Prediction Drift Monitoring tracks how the model's output distribution changes over time, serving as an indirect signal of model degradation even when ground truth labels are unavailable. Feature Skew compares serving vs training feature distributions (requires known training baseline).
Q2: A team is building a Vertex AI Pipeline that includes data preprocessing, model training, and deployment. They need to track all inputs, outputs, and model artifacts for auditability and reproducibility. Which service stores this lineage information?
- A) Cloud Logging
- B) Vertex AI ML Metadata Store ✓
- C) Cloud Storage versioning
- D) Vertex AI Experiments dashboard
Explanation: Vertex AI ML Metadata Store (also called Vertex ML Metadata) automatically tracks lineage: which datasets produced which models, which models were deployed to which endpoints, including hyperparameters and evaluation metrics — enabling full provenance tracking.
Q3: A company wants to automatically retrain their ML model whenever new training data is available in Cloud Storage. The retraining should run a Vertex AI Pipeline and deploy if metrics pass thresholds. Which GCP service should trigger the pipeline?
- A) Vertex AI Schedules
- B) Cloud Storage notifications + Cloud Functions/Eventarc → Vertex AI Pipelines ✓
- C) BigQuery scheduled queries
- D) Cloud Scheduler alone
Explanation: Cloud Storage object finalize notifications can trigger Cloud Functions or Eventarc, which then programmatically start a Vertex AI Pipeline run. This creates event-driven continuous training (MLOps Level 1). Cloud Scheduler triggers on time, not on data availability.