Vertex AI MLOps:管線、CI/CD、Model Registry與生產ML監控
1. MLOps成熟度等級
| 等級 | 說明 | 自動化程度 |
|---|---|---|
| 等級0 | 手動流程,僅有腳本 | 無 |
| 等級1 | ML管線自動化、持續訓練 | 訓練管線 |
| 等級2 | 完整的ML CI/CD、自動重新訓練觸發 | 全部 |
2. Vertex AI Pipelines
Vertex AI Pipelines是Kubeflow Pipelines(KFP)的託管執行環境。管線使用Python SDK定義,編譯為YAML。
Vertex AI Pipeline Structure:
@component (preprocess_data)
↓
@component (train_model)
↓
@component (evaluate_model)
↓ (if accuracy > threshold)
@component (deploy_model)
Each component = isolated Docker container
Artifacts (data, models) stored in Cloud Storage
Metadata tracked in Vertex ML Metadata Store
| 管線SDK | 備註 |
|---|---|
| Kubeflow Pipelines SDK v2 | Vertex AI Pipelines的主要SDK |
| TFX | TensorFlow專用管線元件 |
| Google Cloud Pipeline Components | Vertex AI服務的預建元件 |
3. Vertex AI Model Monitoring
| 監控類型 | 偵測內容 |
|---|---|
| 特徵偏移監控 | 服務端特徵分佈 ≠ 訓練基準線 |
| 特徵漂移監控 | 服務端特徵分佈隨時間變化 |
| 預測漂移 | 模型輸出分佈變化(間接標籤漂移) |
Model Monitoring Workflow:
Training Data Baseline (BigQuery/GCS)
↓ (establish distribution)
Deploy to Endpoint with Monitoring enabled
↓ (collect serving requests)
Periodic Analysis (hourly/daily)
↓ (compare distributions)
Alert if skew/drift > threshold
↓
Retrain trigger → new Pipeline run
4. Vertex AI Experiments與Metadata
| 元件 | 用途 |
|---|---|
| Vertex AI Experiments | 追蹤超參數、指標、跨執行的成品 |
| ML Metadata Store | 追蹤血統:資料 → 模型 → 端點 |
| Vertex AI TensorBoard | 視覺化訓練指標(損失、準確率曲線) |
5. GCP上ML的CI/CD
ML CI/CD Pipeline on GCP:
Code Push to Cloud Source Repositories
↓
Cloud Build trigger (CI)
├── Unit tests for ML components
├── Data validation tests
└── Build Docker image → push to Artifact Registry
↓
Vertex AI Pipeline trigger (CD/CT)
├── Data preprocessing
├── Model training
├── Model evaluation
└── Conditional deployment → Vertex AI Endpoint
考試提示: ML的CI/CD = Cloud Build(程式碼測試 + Docker建構)+ Vertex AI Pipelines(訓練 + 部署協調)。Cloud Source Repositories是GCP的Git託管服務。Artifact Registry取代Container Registry用於儲存Docker映像檔。
6. 練習題
Q1: 生產ML模型的預測分佈在3週內發生了顯著偏移,但真實標籤尚不可用,無法直接衡量準確率。哪種Vertex AI監控類型能偵測到這一點?
- A) 特徵偏移監控
- B) 預測漂移監控 ✓
- C) 訓練資料驗證
- D) Vertex AI Experiments基準線比較
解說:預測漂移監控追蹤模型輸出分佈隨時間的變化,作為模型退化的間接信號,即使在真實標籤不可用時也能使用。特徵偏移比較的是服務端與訓練端的特徵分佈(需要已知的訓練基準線)。
Q2: 團隊正在建構包含資料前處理、模型訓練和部署的Vertex AI管線。需要追蹤所有輸入、輸出和模型成品以進行稽核和可重現性。哪個服務儲存此血統資訊?
- A) Cloud Logging
- B) Vertex AI ML Metadata Store ✓
- C) Cloud Storage版本控制
- D) Vertex AI Experiments儀表板
解說:Vertex AI ML Metadata Store自動追蹤血統:哪些資料集產生了哪些模型、哪些模型部署到了哪些端點,包括超參數和評估指標——實現完整的來源追蹤。
Q3: 一家公司想在Cloud Storage中有新的訓練資料時自動重新訓練ML模型。重新訓練應執行Vertex AI管線,並在指標通過閾值時部署。哪個GCP服務應觸發管線?
- A) Vertex AI Schedules
- B) Cloud Storage通知 + Cloud Functions/Eventarc → Vertex AI Pipelines ✓
- C) BigQuery排程查詢
- D) 單獨使用Cloud Scheduler
解說:Cloud Storage物件完成通知可以觸發Cloud Functions或Eventarc,然後以程式方式啟動Vertex AI管線執行。這建立了事件驅動的持續訓練(MLOps等級1)。Cloud Scheduler按時間觸發,而非按資料可用性觸發。