Chuyển đến nội dung chính

Bài 8: Vertex AI Pipelines & MLOps

Vertex AI Pipelines (Kubeflow Pipelines SDK). Model Registry, Experiments, Metadata Store. Vertex AI Model Monitoring: skew, drift detection. CI/CD cho ML: Cloud Build + Vertex AI.

Vertex AI Pipelines & MLOps

Vertex AI MLOps: Pipelines, CI/CD, Model Registry, và monitoring cho production ML

1. MLOps Maturity Levels

LevelDescriptionAutomation
Level 0Manual process, scripts onlyNone
Level 1ML pipeline automation, continuous trainingTraining pipeline
Level 2Full CI/CD for ML, automated retraining triggersEverything

2. Vertex AI Pipelines

Vertex AI Pipelines là managed execution environment cho Kubeflow Pipelines (KFP). Pipeline được định nghĩa bằng Python SDK và compile thành YAML.

Vertex AI Pipeline Structure:

@component (preprocess_data)
     ↓
@component (train_model)
     ↓
@component (evaluate_model)
     ↓ (if accuracy > threshold)
@component (deploy_model)

Each component = isolated Docker container
Artifacts (data, models) stored in Cloud Storage
Metadata tracked in Vertex ML Metadata Store
Pipeline SDKNotes
Kubeflow Pipelines SDK v2Primary SDK for Vertex AI Pipelines
TFXTensorFlow-specific pipeline components
Google Cloud Pipeline ComponentsPre-built components cho Vertex AI services

3. Vertex AI Model Monitoring

Monitoring TypeWhat It Detects
Feature Skew MonitoringServing feature distribution ≠ training baseline
Feature Drift MonitoringServing feature distribution changes over time
Prediction DriftModel output distribution changes (indirect label drift)
Model Monitoring Workflow:

Training Data Baseline (BigQuery/GCS)
     ↓ (establish distribution)
Deploy to Endpoint with Monitoring enabled
     ↓ (collect serving requests)
Periodic Analysis (hourly/daily)
     ↓ (compare distributions)
Alert if skew/drift > threshold
     ↓
Retrain trigger → new Pipeline run

4. Vertex AI Experiments & Metadata

ComponentPurpose
Vertex AI ExperimentsTrack hyperparameters, metrics, artifacts across runs
ML Metadata StoreTrack lineage: data → model → endpoint
Vertex AI TensorBoardVisualize training metrics (loss, accuracy curves)

5. CI/CD for ML on GCP

ML CI/CD Pipeline on GCP:

Code Push to Cloud Source Repositories
     ↓
Cloud Build trigger (CI)
     ├── Unit tests for ML components
     ├── Data validation tests
     └── Build Docker image → push to Artifact Registry
          ↓
Vertex AI Pipeline trigger (CD/CT)
     ├── Data preprocessing
     ├── Model training
     ├── Model evaluation
     └── Conditional deployment → Vertex AI Endpoint

Exam tip: CI/CD cho ML = Cloud Build (code testing + Docker build) + Vertex AI Pipelines (training + deployment orchestration). Cloud Source Repositories là GCP's Git hosting. Artifact Registry thay thế Container Registry để lưu Docker images.

6. Practice Questions

Q1: A production ML model's prediction distribution has shifted significantly over 3 weeks, but ground truth labels are not yet available to measure accuracy directly. Which Vertex AI monitoring type detects this?

  • A) Feature Skew Monitoring
  • B) Prediction Drift Monitoring ✓
  • C) Training data validation
  • D) Vertex AI Experiments baseline comparison

Explanation: Prediction Drift Monitoring tracks how the model's output distribution changes over time, serving as an indirect signal of model degradation even when ground truth labels are unavailable. Feature Skew compares serving vs training feature distributions (requires known training baseline).

Q2: A team is building a Vertex AI Pipeline that includes data preprocessing, model training, and deployment. They need to track all inputs, outputs, and model artifacts for auditability and reproducibility. Which service stores this lineage information?

  • A) Cloud Logging
  • B) Vertex AI ML Metadata Store ✓
  • C) Cloud Storage versioning
  • D) Vertex AI Experiments dashboard

Explanation: Vertex AI ML Metadata Store (also called Vertex ML Metadata) automatically tracks lineage: which datasets produced which models, which models were deployed to which endpoints, including hyperparameters and evaluation metrics — enabling full provenance tracking.

Q3: A company wants to automatically retrain their ML model whenever new training data is available in Cloud Storage. The retraining should run a Vertex AI Pipeline and deploy if metrics pass thresholds. Which GCP service should trigger the pipeline?

  • A) Vertex AI Schedules
  • B) Cloud Storage notifications + Cloud Functions/Eventarc → Vertex AI Pipelines ✓
  • C) BigQuery scheduled queries
  • D) Cloud Scheduler alone

Explanation: Cloud Storage object finalize notifications can trigger Cloud Functions or Eventarc, which then programmatically start a Vertex AI Pipeline run. This creates event-driven continuous training (MLOps Level 1). Cloud Scheduler triggers on time, not on data availability.