Chuyển đến nội dung chính

第 1 課:什麼是 MLOps? — 機器學習生命週期與成熟度等級

MLOps 基礎:ML 生命週期、DevOps 與 MLOps、成熟度等級 (0→4)、ML 中的技術債、團隊結構、工俱生態系統概述。

🧠 人工智慧與機器學習 — 第 0 課 第 1 課:什麼是 MLOps? — 機器學習生命週期 & 成熟度級別

MLOps 和 LLMOps:將 AI 引入生產

第 1 部分:MLOps 基礎

亞洲開發網

簡介

87% 的 ML 模型從未投入生產。問題不在於模型不好,而是沒有流程將模型從筆記本轉移到現實世界。 MLOps 解決了這個問題。

🎯 MLOps = 機器學習 + DevOps + 資料工程


1. ML 生命週期 — ML 專案生命週期

┌─────────────────────────────────────────────────────────┐
│                    ML LIFECYCLE                          │
│                                                         │
│  1. Problem     2. Data        3. Feature               │
│     Definition     Collection     Engineering           │
│         │              │              │                  │
│         ▼              ▼              ▼                  │
│  4. Model       5. Training    6. Evaluation            │
│     Selection      & Tuning       & Validation          │
│         │              │              │                  │
│         ▼              ▼              ▼                  │
│  7. Deployment  8. Monitoring  9. Retraining            │
│     & Serving      & Alerts       & Updates             │
│         │              │              │                  │
│         └──────────────┴──────────────┘                  │
│                   (Continuous Loop)                      │
└─────────────────────────────────────────────────────────┘

每一步都有自己的問題:

相常見問題
資料資料變更、架構漂移、品質問題
訓練不可重複、迷失方向的實驗
評估線下指標≠線上表現
部署“在我的電腦上運行”但在伺服器上失敗
監控模型衰退,不知道何時重新訓練

2.DevOps 與 MLOps

DevOps (Software):
  Code → Build → Test → Deploy → Monitor
  ✅ Deterministic (cùng code → cùng output)

MLOps (Machine Learning):
  Data + Code + Config → Train → Evaluate → Deploy → Monitor
  ❌ Non-deterministic (cùng code, khác data → khác model)
  ❌ Data dependency (model phụ thuộc vào data quality)
  ❌ Model decay (model giảm chất lượng theo thời gian)

主要區別:

開發營運MLOps
神器二進位/容器模型+資料+配置
測試單元測試、整合+ 資料驗證、模型驗證
CI/CD代碼變更+ 資料變更、模型再訓練
監控正常運作時間、延遲+ 資料漂移、模型效能
版本控制程式碼 (Git)+ 資料 + 模型 + 管道
再現性簡單非常困難(隨機種子、GPU...)

3. MLOps 成熟度模型

0 級:無 MLOps(手動)

Đặc điểm:
  ❌ Jupyter Notebook → Manual deploy
  ❌ Không track experiments
  ❌ Không monitoring
  ❌ Retrain = "ai đó nhớ thì làm"

Team:
  1 Data Scientist làm hết

Phù hợp: POC, hackathon

第 1 級:DevOps,但還不是 MLOps

Đặc điểm:
  ✅ Code trên Git
  ✅ CI/CD pipeline
  ✅ Automated testing (unit tests)
  ❌ Chưa track data versions
  ❌ Chưa track experiments
  ❌ Manual retraining

Team:
  DS + ML Engineer

Phù hợp: Startup giai đoạn đầu

第 2 級:機器學習管道自動化

Đặc điểm:
  ✅ Automated training pipeline
  ✅ Experiment tracking (MLflow)
  ✅ Data versioning (DVC)
  ✅ Model registry
  ✅ Feature store
  ⚠️ Manual trigger retraining

Team:
  DS + ML Engineer + Data Engineer

Phù hợp: Công ty có 5-10 ML models

第 3 級:機器學習的 CI/CD

Đặc điểm:
  ✅ Automated retraining (trigger by data/schedule)
  ✅ A/B testing, canary deployment
  ✅ Model validation pipeline
  ✅ Monitoring + alerting
  ✅ Feature store shared

Team:
  DS + ML Engineer + Data Engineer + ML Platform

Phù hợp: Công ty scale (>10 models)

第 4 級:完整 MLOps

Đặc điểm:
  ✅ Self-healing pipelines
  ✅ Auto-retrain on drift detection
  ✅ Multi-model management
  ✅ Cost optimization
  ✅ Governance & compliance

Team:
  Full ML Platform team

Phù hợp: Big Tech, AI-first companies

4. ML 中的技術債務

Google's "Hidden Technical Debt in ML Systems" (NeurIPS 2015):

┌────────────────────────────────────────────┐
│              ML System                      │
│  ┌────────────────────────────────────┐    │
│  │         ML Code (~5%)              │    │
│  └────────────────────────────────────┘    │
│  ┌────┬─────┬──────┬──────┬────┬──────┐   │
│  │Data│Data │Feat. │Config│Serv│Monit.│   │
│  │Col.│Veri.│Extr. │     │ing │oring │   │
│  └────┴─────┴──────┴──────┴────┴──────┘   │
│              (~95% non-ML code)            │
└────────────────────────────────────────────┘

ML Code chỉ chiếm ~5% tổng hệ thống!

技術債類型:

# 1. Data Dependency Debt
# Input data thay đổi → model hỏng
# VD: Feature từ API bên thứ 3 bị đổi format

# 2. Configuration Debt
# Hyperparams, feature flags, thresholds... không tracked
# VD: Ai đổi threshold từ 0.5 → 0.7? Khi nào?

# 3. Pipeline Debt
# Glue code nối các bước → fragile
# VD: Script bash + cron job + manual copy file

# 4. Reproducibility Debt
# Không thể reproduce kết quả cũ
# VD: "Model v2 tốt hơn v1" — nhưng không reproduce được v1

5. MLOps 工俱生態系統

┌─────────────────────────────────────────────────┐
│                 MLOps Stack                      │
├────────────┬────────────────────────────────────┤
│ Layer      │ Tools                               │
├────────────┼────────────────────────────────────┤
│ Experiment │ MLflow, W&B, Neptune, CometML      │
│ Tracking   │                                     │
├────────────┼────────────────────────────────────┤
│ Data Vers. │ DVC, LakeFS, Delta Lake            │
├────────────┼────────────────────────────────────┤
│ Feature    │ Feast, Tecton, Hopsworks           │
│ Store      │                                     │
├────────────┼────────────────────────────────────┤
│ Model Reg. │ MLflow, Vertex AI, SageMaker       │
├────────────┼────────────────────────────────────┤
│ Orchest.   │ Kubeflow, Airflow, Prefect         │
├────────────┼────────────────────────────────────┤
│ Serving    │ TorchServe, Triton, TFServing, BentoML │
├────────────┼────────────────────────────────────┤
│ Monitoring │ Evidently, Arize, WhyLabs          │
├────────────┼────────────────────────────────────┤
│ Infra      │ Docker, K8s, Terraform             │
└────────────┴────────────────────────────────────┘

根據團隊規模選擇工具:

團隊規模推薦
1-3人MLflow + DVC + Docker
3-10人+ 氣流 + 盛宴 + 顯然
10+人全平台:Kubeflow / Vertex AI / SageMaker
企業商業:Databricks、Dataiku、Domino

6. 實作:設定 MLOps 項目

"""Setup cấu trúc project MLOps chuẩn"""

# Project structure
project_structure = """
my-ml-project/
├── data/
│   ├── raw/              # Data gốc (never modify)
│   ├── processed/        # Data sau preprocessing
│   └── features/         # Feature store output
├── notebooks/            # EDA, prototyping
├── src/
│   ├── data/             # Data processing code
│   ├── features/         # Feature engineering
│   ├── models/           # Model training code
│   ├── serving/          # Inference server
│   └── monitoring/       # Monitoring code
├── configs/
│   ├── training.yaml     # Training hyperparameters
│   ├── serving.yaml      # Serving config
│   └── monitoring.yaml   # Alerting rules
├── tests/
│   ├── test_data.py      # Data validation tests
│   ├── test_model.py     # Model validation tests
│   └── test_api.py       # API tests
├── pipelines/
│   ├── training.py       # Training pipeline
│   ├── evaluation.py     # Evaluation pipeline
│   └── deployment.py     # Deployment pipeline
├── Dockerfile
├── docker-compose.yml
├── Makefile              # Common commands
├── dvc.yaml              # DVC pipeline
├── mlflow.yaml           # MLflow config
└── README.md
"""

print(project_structure)
# Makefile — Common commands
.PHONY: setup train evaluate deploy monitor

setup:
	pip install -r requirements.txt
	dvc pull

train:
	python pipelines/training.py --config configs/training.yaml

evaluate:
	python pipelines/evaluation.py --model-version latest

deploy:
	python pipelines/deployment.py --target production

monitor:
	python src/monitoring/check_drift.py

test:
	pytest tests/ -v

lint:
	ruff check src/
	mypy src/

總結

概念記住
MLOps機器學習 + DevOps + 資料工程
機器學習生命週期資料→訓練→部署→監控→重新訓練
成熟度 0-4手動 → 自動 → CI/CD → 完整 MLOps
科技債務ML 代碼只佔 5%,95% 是基礎設施
工具MLflow、DVC、盛宴、Kubeflow、顯然

練習

  1. 評估您的團隊: 您的團隊處於什麼成熟度?列出差距。
  2. 專案設定: 根據上述範本建立專案結構。初始化 git + dvc。
  3. 研究工具: 比較 2 個實驗追蹤工具:MLflow 與權重和偏差。為團隊選擇 1。

下一篇文章: 實驗追蹤 - MLflow 和權重和偏差。