Chuyển đến nội dung chính

レッスン 1: MLOps とは何ですか? — ML のライフサイクルと成熟度レベル

MLOps の基盤: ML ライフサイクル、DevOps と MLOps、成熟度レベル (0→4)、ML の技術的負債、チーム構造、ツール エコシステムの概要。

🧠 AI と ML — レッスン 0 レッスン 1: MLOps とは何ですか? — ML ライフサイクルと 成熟度レベル

MLOps と LLMOps: AI を本番環境に導入する

パート 1: MLOps の基礎

xdev.asia

はじめに

ML モデルの 87% は本番環境に導入されることはありません。問題はモデルが悪いことではなく、モデルをノートブックから現実世界に移すためのプロセスがないことです。 MLOps はこれを解決します。

🎯 MLOps = 機械学習 + DevOps + データ エンジニアリング


1. ML ライフサイクル — ML プロジェクトのライフサイクル

┌─────────────────────────────────────────────────────────┐
│                    ML LIFECYCLE                          │
│                                                         │
│  1. Problem     2. Data        3. Feature               │
│     Definition     Collection     Engineering           │
│         │              │              │                  │
│         ▼              ▼              ▼                  │
│  4. Model       5. Training    6. Evaluation            │
│     Selection      & Tuning       & Validation          │
│         │              │              │                  │
│         ▼              ▼              ▼                  │
│  7. Deployment  8. Monitoring  9. Retraining            │
│     & Serving      & Alerts       & Updates             │
│         │              │              │                  │
│         └──────────────┴──────────────┘                  │
│                   (Continuous Loop)                      │
└─────────────────────────────────────────────────────────┘

各ステップには独自の問題があります。

フェーズよくある問題
データデータの変更、スキーマのドリフト、品質の問題
トレーニング再現性のない、道に迷った実験
評価オフラインの指標 ≠ オンラインのパフォーマンス
展開「コンピュータでは実行されます」が、サーバーでは失敗します。
モニタリングモデルの崩壊、いつ再トレーニングすべきかわからない

2. DevOps と MLOps

DevOps (Software):
  Code → Build → Test → Deploy → Monitor
  ✅ Deterministic (cùng code → cùng output)

MLOps (Machine Learning):
  Data + Code + Config → Train → Evaluate → Deploy → Monitor
  ❌ Non-deterministic (cùng code, khác data → khác model)
  ❌ Data dependency (model phụ thuộc vào data quality)
  ❌ Model decay (model giảm chất lượng theo thời gian)

主な違い:

開発運用MLOps
アーティファクトバイナリ/コンテナモデル + データ + 構成
テスト単体テスト、統合+ データ検証、モデル検証
CI/CDコードの変更+ データ変更、モデルの再トレーニング
モニタリング稼働時間、遅延+ データドリフト、モデルのパフォーマンス
バージョン管理コード (Git)+ データ + モデル + パイプライン
再現性簡単非常に難しい (ランダム シード、GPU など)

3. MLOps 成熟度モデル

レベル 0: MLOps なし (手動)

Đặc điểm:
  ❌ Jupyter Notebook → Manual deploy
  ❌ Không track experiments
  ❌ Không monitoring
  ❌ Retrain = "ai đó nhớ thì làm"

Team:
  1 Data Scientist làm hết

Phù hợp: POC, hackathon

レベル 1: DevOps はあるが、まだ MLOps ではない

Đặc điểm:
  ✅ Code trên Git
  ✅ CI/CD pipeline
  ✅ Automated testing (unit tests)
  ❌ Chưa track data versions
  ❌ Chưa track experiments
  ❌ Manual retraining

Team:
  DS + ML Engineer

Phù hợp: Startup giai đoạn đầu

レベル 2: ML パイプラインの自動化

Đặc điểm:
  ✅ Automated training pipeline
  ✅ Experiment tracking (MLflow)
  ✅ Data versioning (DVC)
  ✅ Model registry
  ✅ Feature store
  ⚠️ Manual trigger retraining

Team:
  DS + ML Engineer + Data Engineer

Phù hợp: Công ty có 5-10 ML models

レベル 3: ML の CI/CD

Đặc điểm:
  ✅ Automated retraining (trigger by data/schedule)
  ✅ A/B testing, canary deployment
  ✅ Model validation pipeline
  ✅ Monitoring + alerting
  ✅ Feature store shared

Team:
  DS + ML Engineer + Data Engineer + ML Platform

Phù hợp: Công ty scale (>10 models)

レベル 4: 完全な MLOps

Đặc điểm:
  ✅ Self-healing pipelines
  ✅ Auto-retrain on drift detection
  ✅ Multi-model management
  ✅ Cost optimization
  ✅ Governance & compliance

Team:
  Full ML Platform team

Phù hợp: Big Tech, AI-first companies

4. ML における技術的負債

Google's "Hidden Technical Debt in ML Systems" (NeurIPS 2015):

┌────────────────────────────────────────────┐
│              ML System                      │
│  ┌────────────────────────────────────┐    │
│  │         ML Code (~5%)              │    │
│  └────────────────────────────────────┘    │
│  ┌────┬─────┬──────┬──────┬────┬──────┐   │
│  │Data│Data │Feat. │Config│Serv│Monit.│   │
│  │Col.│Veri.│Extr. │     │ing │oring │   │
│  └────┴─────┴──────┴──────┴────┴──────┘   │
│              (~95% non-ML code)            │
└────────────────────────────────────────────┘

ML Code chỉ chiếm ~5% tổng hệ thống!

技術的負債の種類:

# 1. Data Dependency Debt
# Input data thay đổi → model hỏng
# VD: Feature từ API bên thứ 3 bị đổi format

# 2. Configuration Debt
# Hyperparams, feature flags, thresholds... không tracked
# VD: Ai đổi threshold từ 0.5 → 0.7? Khi nào?

# 3. Pipeline Debt
# Glue code nối các bước → fragile
# VD: Script bash + cron job + manual copy file

# 4. Reproducibility Debt
# Không thể reproduce kết quả cũ
# VD: "Model v2 tốt hơn v1" — nhưng không reproduce được v1

5. MLOps ツール エコシステム

┌─────────────────────────────────────────────────┐
│                 MLOps Stack                      │
├────────────┬────────────────────────────────────┤
│ Layer      │ Tools                               │
├────────────┼────────────────────────────────────┤
│ Experiment │ MLflow, W&B, Neptune, CometML      │
│ Tracking   │                                     │
├────────────┼────────────────────────────────────┤
│ Data Vers. │ DVC, LakeFS, Delta Lake            │
├────────────┼────────────────────────────────────┤
│ Feature    │ Feast, Tecton, Hopsworks           │
│ Store      │                                     │
├────────────┼────────────────────────────────────┤
│ Model Reg. │ MLflow, Vertex AI, SageMaker       │
├────────────┼────────────────────────────────────┤
│ Orchest.   │ Kubeflow, Airflow, Prefect         │
├────────────┼────────────────────────────────────┤
│ Serving    │ TorchServe, Triton, TFServing, BentoML │
├────────────┼────────────────────────────────────┤
│ Monitoring │ Evidently, Arize, WhyLabs          │
├────────────┼────────────────────────────────────┤
│ Infra      │ Docker, K8s, Terraform             │
└────────────┴────────────────────────────────────┘

チームの規模に応じて選択されるツール:

チームの規模推薦
1 ~ 3 名様MLflow + DVC + Docker
3 ~ 10 名様+ 空気の流れ + 饗宴 + 明らかに
10 名以上フルプラットフォーム: Kubeflow / Vertex AI / SageMaker
エンタープライズコマーシャル: Databricks、Dataiku、Domino

6. ハンズオン: MLOps プロジェクトのセットアップ

"""Setup cấu trúc project MLOps chuẩn"""

# Project structure
project_structure = """
my-ml-project/
├── data/
│   ├── raw/              # Data gốc (never modify)
│   ├── processed/        # Data sau preprocessing
│   └── features/         # Feature store output
├── notebooks/            # EDA, prototyping
├── src/
│   ├── data/             # Data processing code
│   ├── features/         # Feature engineering
│   ├── models/           # Model training code
│   ├── serving/          # Inference server
│   └── monitoring/       # Monitoring code
├── configs/
│   ├── training.yaml     # Training hyperparameters
│   ├── serving.yaml      # Serving config
│   └── monitoring.yaml   # Alerting rules
├── tests/
│   ├── test_data.py      # Data validation tests
│   ├── test_model.py     # Model validation tests
│   └── test_api.py       # API tests
├── pipelines/
│   ├── training.py       # Training pipeline
│   ├── evaluation.py     # Evaluation pipeline
│   └── deployment.py     # Deployment pipeline
├── Dockerfile
├── docker-compose.yml
├── Makefile              # Common commands
├── dvc.yaml              # DVC pipeline
├── mlflow.yaml           # MLflow config
└── README.md
"""

print(project_structure)
# Makefile — Common commands
.PHONY: setup train evaluate deploy monitor

setup:
	pip install -r requirements.txt
	dvc pull

train:
	python pipelines/training.py --config configs/training.yaml

evaluate:
	python pipelines/evaluation.py --model-version latest

deploy:
	python pipelines/deployment.py --target production

monitor:
	python src/monitoring/check_drift.py

test:
	pytest tests/ -v

lint:
	ruff check src/
	mypy src/

概要

コンセプト覚えておいてください
MLOpsML + DevOps + データ エンジニアリング
ML ライフサイクルデータ → トレーニング → 導入 → 監視 → 再トレーニング
成熟度 0-4手動 → 自動 → CI/CD → フル MLOps
技術的負債ML コードはわずか ~5%、95% はインフラストラクチャです。
ツールMLflow、DVC、Feast、Kubeflow、明らかに

演習

  1. チームを評価します: あなたのチームはどのくらいの成熟度レベルにありますか?ギャップをリストします。
  2. プロジェクトのセットアップ: 上記のテンプレートに従ってプロジェクト構造を作成します。 git + dvc を初期化します。
  3. 調査ツール: 2 つの実験追跡ツールを比較します: MLflow と重みとバイアス。チームに 1 つを選択します。

次の記事: 実験の追跡 — MLflow、重み、バイアス。