Introduction
ML model runs on laptop → transferred to server → fails. "Works on my machine" syndrome but ML version: different from CUDA, different from PyTorch, different from numpy version...
🎯 This article: Docker containerize ML, Kubernetes orchestrate, Cloud ML platforms to scale.
1. Docker for ML
1.1 Basic ML Dockerfile
# Dockerfile — ML Training
FROM python:3.11-slim
# System dependencies
RUN apt-get update && apt-get install -y --no-install-recommends \
build-essential \
libgomp1 \
&& rm -rf /var/lib/apt/lists/*
WORKDIR /app
# Install Python deps (cache layer)
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
# Copy code
COPY src/ src/
COPY configs/ configs/
# Default command
CMD ["python", "src/models/train.py", "--config", "configs/training.yaml"]
1.2 GPU Docker
# Dockerfile.gpu — ML Training với GPU
FROM nvidia/cuda:12.1-cudnn8-runtime-ubuntu22.04
# Python
RUN apt-get update && apt-get install -y --no-install-recommends \
python3 python3-pip \
&& rm -rf /var/lib/apt/lists/*
WORKDIR /app
# PyTorch with CUDA
COPY requirements.txt .
RUN pip3 install --no-cache-dir torch torchvision --index-url \
https://download.pytorch.org/whl/cu121
RUN pip3 install --no-cache-dir -r requirements.txt
COPY src/ src/
COPY configs/ configs/
CMD ["python3", "src/models/train.py"]
1.3 Multi-stage Build (Production Serving)
# Dockerfile.serving — Optimized cho serving
# === Stage 1: Build ===
FROM python:3.11 AS builder
WORKDIR /build
COPY requirements-serving.txt .
RUN pip install --no-cache-dir --prefix=/install -r requirements-serving.txt
# === Stage 2: Runtime ===
FROM python:3.11-slim
# Copy chỉ những gì cần
COPY --from=builder /install /usr/local
WORKDIR /app
COPY src/serving/ src/serving/
COPY models/ models/
# Non-root user
RUN useradd -m appuser
USER appuser
EXPOSE 8000
HEALTHCHECK --interval=30s --timeout=3s \
CMD curl -f http://localhost:8000/health || exit 1
CMD ["uvicorn", "src.serving.api:app", \
"--host", "0.0.0.0", "--port", "8000", \
"--workers", "4"]
1.4 Docker Compose for ML Stack
# docker-compose.yml
services:
# Training
trainer:
build:
context: .
dockerfile: Dockerfile.gpu
deploy:
resources:
reservations:
devices:
- capabilities: [gpu]
volumes:
- ./data:/app/data
- ./models:/app/models
environment:
- MLFLOW_TRACKING_URI=http://mlflow:5000
- WANDB_API_KEY=${WANDB_API_KEY}
# Serving
api:
build:
context: .
dockerfile: Dockerfile.serving
ports:
- "8000:8000"
volumes:
- ./models:/app/models:ro
environment:
- MODEL_NAME=churn-predictor
- MODEL_STAGE=Production
# MLflow
mlflow:
image: ghcr.io/mlflow/mlflow:latest
ports:
- "5000:5000"
volumes:
- mlflow_data:/mlflow
command: >
mlflow server
--host 0.0.0.0
--backend-store-uri sqlite:///mlflow/mlflow.db
--default-artifact-root /mlflow/artifacts
# Monitoring
prometheus:
image: prom/prometheus:latest
ports:
- "9090:9090"
volumes:
- ./monitoring/prometheus.yml:/etc/prometheus/prometheus.yml
grafana:
image: grafana/grafana:latest
ports:
- "3000:3000"
depends_on:
- prometheus
volumes:
- grafana_data:/var/lib/grafana
volumes:
mlflow_data:
grafana_data:
2. Kubernetes for ML
2.1 ML Training Job
# k8s/training-job.yaml
apiVersion: batch/v1
kind: Job
metadata:
name: model-training-v2
labels:
app: ml-training
model: churn-predictor
spec:
backoffLimit: 2
template:
spec:
restartPolicy: Never
containers:
- name: trainer
image: my-registry/ml-trainer:latest
command: ["python", "src/models/train.py"]
args:
- "--config"
- "configs/training.yaml"
- "--experiment"
- "churn-v2"
resources:
requests:
cpu: "4"
memory: "16Gi"
nvidia.com/gpu: "1"
limits:
cpu: "8"
memory: "32Gi"
nvidia.com/gpu: "1"
env:
- name: MLFLOW_TRACKING_URI
value: "http://mlflow-service:5000"
volumeMounts:
- name: data
mountPath: /app/data
- name: models
mountPath: /app/models
volumes:
- name: data
persistentVolumeClaim:
claimName: training-data-pvc
- name: models
persistentVolumeClaim:
claimName: models-pvc
nodeSelector:
gpu-type: a100
2.2 Model Serving Deployment
# k8s/serving-deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: churn-predictor
labels:
app: churn-predictor
spec:
replicas: 3
selector:
matchLabels:
app: churn-predictor
template:
metadata:
labels:
app: churn-predictor
annotations:
prometheus.io/scrape: "true"
prometheus.io/port: "8000"
spec:
containers:
- name: api
image: my-registry/ml-serving:latest
ports:
- containerPort: 8000
resources:
requests:
cpu: "1"
memory: "2Gi"
limits:
cpu: "2"
memory: "4Gi"
readinessProbe:
httpGet:
path: /health
port: 8000
initialDelaySeconds: 10
periodSeconds: 5
livenessProbe:
httpGet:
path: /health
port: 8000
initialDelaySeconds: 30
periodSeconds: 10
env:
- name: MODEL_NAME
value: "churn-predictor"
---
apiVersion: v1
kind: Service
metadata:
name: churn-predictor-svc
spec:
selector:
app: churn-predictor
ports:
- port: 80
targetPort: 8000
type: ClusterIP
---
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: churn-predictor-hpa
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: churn-predictor
minReplicas: 2
maxReplicas: 10
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
2.3 CronJob for Scheduled Retraining
# k8s/retrain-cronjob.yaml
apiVersion: batch/v1
kind: CronJob
metadata:
name: weekly-retrain
spec:
schedule: "0 2 * * 1" # Monday 2am
jobTemplate:
spec:
template:
spec:
restartPolicy: OnFailure
containers:
- name: retrain
image: my-registry/ml-trainer:latest
command: ["python", "pipelines/retrain.py"]
resources:
requests:
nvidia.com/gpu: "1"
3. Cloud ML Platforms
3.1 AWS SageMaker
"""AWS SageMaker — Managed ML platform"""
import sagemaker
from sagemaker.estimator import Estimator
session = sagemaker.Session()
role = "arn:aws:iam::role/SageMakerRole"
# Training
estimator = Estimator(
image_uri="my-training-image:latest",
role=role,
instance_count=1,
instance_type="ml.g5.xlarge", # GPU instance
output_path="s3://my-bucket/models/",
hyperparameters={
"epochs": 20,
"learning_rate": 0.001,
"batch_size": 32,
},
)
estimator.fit({
"training": "s3://my-bucket/data/train/",
"validation": "s3://my-bucket/data/val/",
})
# Deploy
predictor = estimator.deploy(
initial_instance_count=1,
instance_type="ml.m5.xlarge",
)
# Predict
result = predictor.predict(input_data)
3.2 GCP Vertex AI
"""Google Cloud Vertex AI"""
from google.cloud import aiplatform
aiplatform.init(
project="my-project",
location="us-central1",
)
# Training
job = aiplatform.CustomTrainingJob(
display_name="churn-training",
script_path="src/models/train.py",
container_uri="us-docker.pkg.dev/vertex-ai/training/pytorch-gpu.1-13:latest",
requirements=["scikit-learn", "pandas"],
)
model = job.run(
replica_count=1,
machine_type="n1-standard-8",
accelerator_type="NVIDIA_TESLA_T4",
accelerator_count=1,
)
# Deploy
endpoint = model.deploy(
machine_type="n1-standard-4",
min_replica_count=1,
max_replica_count=5,
)
# Predict
prediction = endpoint.predict(instances=[{"features": [1, 2, 3]}])
3.3 Compare Cloud Platforms
| Features | AWS SageMaker | GCP Vertex AI | Azure ML |
|---|---|---|---|
| Ease of use | ⭐⭐⭐ | ⭐⭐⭐⭐ | ⭐⭐⭐ |
| GPU availability | ⭐⭐⭐⭐⭐ | ⭐⭐⭐⭐ | ⭐⭐⭐⭐ |
| AutoML | ✅ Autopilot | ✅ AutoML | ✅ AutoML |
| Experiment tracking | ✅ Built-in | ✅ Built-in | ✅ Built-in |
| Model registry | ✅ | ✅ | ✅ |
| Feature store | ✅ | ✅ | ✅ |
| Pricing | 💰💰💰 | 💰💰 | 💰💰💰 |
| Best for | Enterprise | Research + Prod | Microsoft shops |
4. Infrastructure as Code
# terraform/ml-infra/main.tf
# Infrastructure as Code với Terraform
provider "aws" {
region = "us-east-1"
}
# S3 bucket cho data & models
resource "aws_s3_bucket" "ml_artifacts" {
bucket = "my-ml-artifacts"
versioning {
enabled = true
}
}
# ECR cho Docker images
resource "aws_ecr_repository" "ml_training" {
name = "ml-training"
}
resource "aws_ecr_repository" "ml_serving" {
name = "ml-serving"
}
# ECS Cluster cho serving
resource "aws_ecs_cluster" "ml_serving" {
name = "ml-serving-cluster"
}
# SageMaker Endpoint
resource "aws_sagemaker_endpoint_configuration" "churn" {
name = "churn-predictor-config"
production_variants {
variant_name = "primary"
model_name = aws_sagemaker_model.churn.name
initial_instance_count = 2
instance_type = "ml.m5.xlarge"
}
}
# RDS cho MLflow
resource "aws_db_instance" "mlflow" {
engine = "postgres"
instance_class = "db.t3.medium"
allocated_storage = 20
db_name = "mlflow"
}
5. Best Practices
Docker:
✅ Multi-stage builds (build vs runtime)
✅ Pin dependency versions
✅ Non-root user
✅ Health checks
✅ .dockerignore (exclude data, notebooks)
Kubernetes:
✅ Resource requests & limits
✅ HPA for auto-scaling
✅ Readiness & liveness probes
✅ Node selectors cho GPU
✅ Separate namespaces (dev/staging/prod)
Cloud:
✅ Spot instances cho training (70% cheaper)
✅ Auto-scaling cho serving
✅ Monitoring & alerting
✅ Infrastructure as Code
✅ Cost monitoring & budgets
Summary
| Concepts | Remember |
|---|---|
| Docker | Containerize ML → reproducible environment |
| Multi-stage | Separate build vs runtime → smaller images |
| GPU Docker | nvidia/cuda base image + NVIDIA Container Toolkit |
| Kubernetes | Orchestrate: Jobs (train), Deployments (serve), CronJobs (retrain) |
| HPA | Auto-scale serving pods |
| SageMaker | AWS managed ML, end-to-end |
| Vertex AI | GCP managed ML, good for research |
| Terraform | Infrastructure as Code |
Exercises
- Docker: Create Dockerfile for ML serving (multi-stage, < 500MB image size).
- Compose: Create docker-compose with API + MLflow + Prometheus + Grafana.
- K8s: Deploy model serving to K8s cluster (minikube/kind). Add HPA.
- Cloud: Deploy a model to SageMaker or Vertex AI (free tier).
Next article: LLMOps vs MLOps — Paradigm Shift.