Lesson 1: What is MLOps? — ML Lifecycle & Maturity Levels
MLOps foundation: ML lifecycle, DevOps vs MLOps, maturity levels (0→4), technical debt in ML, team structure, tools ecosystem overview.
MLOps & LLMOps: Bringing AI to Production
Part 1: MLOps Foundations
xdev.asia
Introduction
87% of ML models never go to production. The problem isn't that the model is bad — it's that there is no process for getting the model from the notebook to the real world. MLOps solves this.
🎯 MLOps = Machine Learning + DevOps + Data Engineering
DevOps (Software):
Code → Build → Test → Deploy → Monitor
✅ Deterministic (cùng code → cùng output)
MLOps (Machine Learning):
Data + Code + Config → Train → Evaluate → Deploy → Monitor
❌ Non-deterministic (cùng code, khác data → khác model)
❌ Data dependency (model phụ thuộc vào data quality)
❌ Model decay (model giảm chất lượng theo thời gian)
Key differences:
DevOps
MLOps
Artifact
Binary / Container
Model + Data + Config
Testing
Unit tests, integration
+ Data validation, model validation
CI/CD
Code changes
+ Data changes, model retraining
Monitoring
Uptime, latency
+ Data drift, model performance
Versioning
Code (Git)
+ Data + Model + Pipeline
Reproducibility
Easy
Very difficult (random seeds, GPU, ...)
3. MLOps Maturity Model
Level 0: No MLOps (Manual)
Đặc điểm:
❌ Jupyter Notebook → Manual deploy
❌ Không track experiments
❌ Không monitoring
❌ Retrain = "ai đó nhớ thì làm"
Team:
1 Data Scientist làm hết
Phù hợp: POC, hackathon
Level 1: DevOps but not yet MLOps
Đặc điểm:
✅ Code trên Git
✅ CI/CD pipeline
✅ Automated testing (unit tests)
❌ Chưa track data versions
❌ Chưa track experiments
❌ Manual retraining
Team:
DS + ML Engineer
Phù hợp: Startup giai đoạn đầu
Level 2: ML Pipeline Automation
Đặc điểm:
✅ Automated training pipeline
✅ Experiment tracking (MLflow)
✅ Data versioning (DVC)
✅ Model registry
✅ Feature store
⚠️ Manual trigger retraining
Team:
DS + ML Engineer + Data Engineer
Phù hợp: Công ty có 5-10 ML models
Level 3: CI/CD for ML
Đặc điểm:
✅ Automated retraining (trigger by data/schedule)
✅ A/B testing, canary deployment
✅ Model validation pipeline
✅ Monitoring + alerting
✅ Feature store shared
Team:
DS + ML Engineer + Data Engineer + ML Platform
Phù hợp: Công ty scale (>10 models)
Level 4: Full MLOps
Đặc điểm:
✅ Self-healing pipelines
✅ Auto-retrain on drift detection
✅ Multi-model management
✅ Cost optimization
✅ Governance & compliance
Team:
Full ML Platform team
Phù hợp: Big Tech, AI-first companies
4. Technical Debt in ML
Google's "Hidden Technical Debt in ML Systems" (NeurIPS 2015):
┌────────────────────────────────────────────┐
│ ML System │
│ ┌────────────────────────────────────┐ │
│ │ ML Code (~5%) │ │
│ └────────────────────────────────────┘ │
│ ┌────┬─────┬──────┬──────┬────┬──────┐ │
│ │Data│Data │Feat. │Config│Serv│Monit.│ │
│ │Col.│Veri.│Extr. │ │ing │oring │ │
│ └────┴─────┴──────┴──────┴────┴──────┘ │
│ (~95% non-ML code) │
└────────────────────────────────────────────┘
ML Code chỉ chiếm ~5% tổng hệ thống!
Types of technical debt:
# 1. Data Dependency Debt
# Input data thay đổi → model hỏng
# VD: Feature từ API bên thứ 3 bị đổi format
# 2. Configuration Debt
# Hyperparams, feature flags, thresholds... không tracked
# VD: Ai đổi threshold từ 0.5 → 0.7? Khi nào?
# 3. Pipeline Debt
# Glue code nối các bước → fragile
# VD: Script bash + cron job + manual copy file
# 4. Reproducibility Debt
# Không thể reproduce kết quả cũ
# VD: "Model v2 tốt hơn v1" — nhưng không reproduce được v1