Chuyển đến nội dung chính

Lesson 6: Infrastructure — Docker, Kubernetes & Cloud ML

ML infrastructure: Docker for ML, multi-stage builds. Kubernetes for ML workloads, GPU scheduling. Cloud ML platforms: AWS SageMaker, GCP Vertex AI, Azure ML. Infrastructure as Code with Terraform.

🧠 AI & ML — Lesson 5 Lesson 6: Infrastructure — Docker, Kubernetes & Cloud ML

MLOps & LLMOps: Bringing AI to Production

Part 2: ML Infrastructure

xdev.asia

Introduction

ML model runs on laptop → transferred to server → fails. "Works on my machine" syndrome but ML version: different from CUDA, different from PyTorch, different from numpy version...

🎯 This article: Docker containerize ML, Kubernetes orchestrate, Cloud ML platforms to scale.


1. Docker for ML

1.1 Basic ML Dockerfile

# Dockerfile — ML Training
FROM python:3.11-slim

# System dependencies
RUN apt-get update && apt-get install -y --no-install-recommends \
    build-essential \
    libgomp1 \
    && rm -rf /var/lib/apt/lists/*

WORKDIR /app

# Install Python deps (cache layer)
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt

# Copy code
COPY src/ src/
COPY configs/ configs/

# Default command
CMD ["python", "src/models/train.py", "--config", "configs/training.yaml"]

1.2 GPU Docker

# Dockerfile.gpu — ML Training với GPU
FROM nvidia/cuda:12.1-cudnn8-runtime-ubuntu22.04

# Python
RUN apt-get update && apt-get install -y --no-install-recommends \
    python3 python3-pip \
    && rm -rf /var/lib/apt/lists/*

WORKDIR /app

# PyTorch with CUDA
COPY requirements.txt .
RUN pip3 install --no-cache-dir torch torchvision --index-url \
    https://download.pytorch.org/whl/cu121
RUN pip3 install --no-cache-dir -r requirements.txt

COPY src/ src/
COPY configs/ configs/

CMD ["python3", "src/models/train.py"]

1.3 Multi-stage Build (Production Serving)

# Dockerfile.serving — Optimized cho serving
# === Stage 1: Build ===
FROM python:3.11 AS builder

WORKDIR /build
COPY requirements-serving.txt .
RUN pip install --no-cache-dir --prefix=/install -r requirements-serving.txt

# === Stage 2: Runtime ===
FROM python:3.11-slim

# Copy chỉ những gì cần
COPY --from=builder /install /usr/local

WORKDIR /app
COPY src/serving/ src/serving/
COPY models/ models/

# Non-root user
RUN useradd -m appuser
USER appuser

EXPOSE 8000

HEALTHCHECK --interval=30s --timeout=3s \
    CMD curl -f http://localhost:8000/health || exit 1

CMD ["uvicorn", "src.serving.api:app", \
     "--host", "0.0.0.0", "--port", "8000", \
     "--workers", "4"]

1.4 Docker Compose for ML Stack

# docker-compose.yml
services:
  # Training
  trainer:
    build:
      context: .
      dockerfile: Dockerfile.gpu
    deploy:
      resources:
        reservations:
          devices:
            - capabilities: [gpu]
    volumes:
      - ./data:/app/data
      - ./models:/app/models
    environment:
      - MLFLOW_TRACKING_URI=http://mlflow:5000
      - WANDB_API_KEY=${WANDB_API_KEY}

  # Serving
  api:
    build:
      context: .
      dockerfile: Dockerfile.serving
    ports:
      - "8000:8000"
    volumes:
      - ./models:/app/models:ro
    environment:
      - MODEL_NAME=churn-predictor
      - MODEL_STAGE=Production

  # MLflow
  mlflow:
    image: ghcr.io/mlflow/mlflow:latest
    ports:
      - "5000:5000"
    volumes:
      - mlflow_data:/mlflow
    command: >
      mlflow server
      --host 0.0.0.0
      --backend-store-uri sqlite:///mlflow/mlflow.db
      --default-artifact-root /mlflow/artifacts

  # Monitoring
  prometheus:
    image: prom/prometheus:latest
    ports:
      - "9090:9090"
    volumes:
      - ./monitoring/prometheus.yml:/etc/prometheus/prometheus.yml

  grafana:
    image: grafana/grafana:latest
    ports:
      - "3000:3000"
    depends_on:
      - prometheus
    volumes:
      - grafana_data:/var/lib/grafana

volumes:
  mlflow_data:
  grafana_data:

2. Kubernetes for ML

2.1 ML Training Job

# k8s/training-job.yaml
apiVersion: batch/v1
kind: Job
metadata:
  name: model-training-v2
  labels:
    app: ml-training
    model: churn-predictor
spec:
  backoffLimit: 2
  template:
    spec:
      restartPolicy: Never
      containers:
        - name: trainer
          image: my-registry/ml-trainer:latest
          command: ["python", "src/models/train.py"]
          args:
            - "--config"
            - "configs/training.yaml"
            - "--experiment"
            - "churn-v2"
          resources:
            requests:
              cpu: "4"
              memory: "16Gi"
              nvidia.com/gpu: "1"
            limits:
              cpu: "8"
              memory: "32Gi"
              nvidia.com/gpu: "1"
          env:
            - name: MLFLOW_TRACKING_URI
              value: "http://mlflow-service:5000"
          volumeMounts:
            - name: data
              mountPath: /app/data
            - name: models
              mountPath: /app/models
      volumes:
        - name: data
          persistentVolumeClaim:
            claimName: training-data-pvc
        - name: models
          persistentVolumeClaim:
            claimName: models-pvc
      nodeSelector:
        gpu-type: a100

2.2 Model Serving Deployment

# k8s/serving-deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: churn-predictor
  labels:
    app: churn-predictor
spec:
  replicas: 3
  selector:
    matchLabels:
      app: churn-predictor
  template:
    metadata:
      labels:
        app: churn-predictor
      annotations:
        prometheus.io/scrape: "true"
        prometheus.io/port: "8000"
    spec:
      containers:
        - name: api
          image: my-registry/ml-serving:latest
          ports:
            - containerPort: 8000
          resources:
            requests:
              cpu: "1"
              memory: "2Gi"
            limits:
              cpu: "2"
              memory: "4Gi"
          readinessProbe:
            httpGet:
              path: /health
              port: 8000
            initialDelaySeconds: 10
            periodSeconds: 5
          livenessProbe:
            httpGet:
              path: /health
              port: 8000
            initialDelaySeconds: 30
            periodSeconds: 10
          env:
            - name: MODEL_NAME
              value: "churn-predictor"
---
apiVersion: v1
kind: Service
metadata:
  name: churn-predictor-svc
spec:
  selector:
    app: churn-predictor
  ports:
    - port: 80
      targetPort: 8000
  type: ClusterIP
---
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: churn-predictor-hpa
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: churn-predictor
  minReplicas: 2
  maxReplicas: 10
  metrics:
    - type: Resource
      resource:
        name: cpu
        target:
          type: Utilization
          averageUtilization: 70

2.3 CronJob for Scheduled Retraining

# k8s/retrain-cronjob.yaml
apiVersion: batch/v1
kind: CronJob
metadata:
  name: weekly-retrain
spec:
  schedule: "0 2 * * 1"  # Monday 2am
  jobTemplate:
    spec:
      template:
        spec:
          restartPolicy: OnFailure
          containers:
            - name: retrain
              image: my-registry/ml-trainer:latest
              command: ["python", "pipelines/retrain.py"]
              resources:
                requests:
                  nvidia.com/gpu: "1"

3. Cloud ML Platforms

3.1 AWS SageMaker

"""AWS SageMaker — Managed ML platform"""
import sagemaker
from sagemaker.estimator import Estimator

session = sagemaker.Session()
role = "arn:aws:iam::role/SageMakerRole"

# Training
estimator = Estimator(
    image_uri="my-training-image:latest",
    role=role,
    instance_count=1,
    instance_type="ml.g5.xlarge",  # GPU instance
    output_path="s3://my-bucket/models/",
    hyperparameters={
        "epochs": 20,
        "learning_rate": 0.001,
        "batch_size": 32,
    },
)

estimator.fit({
    "training": "s3://my-bucket/data/train/",
    "validation": "s3://my-bucket/data/val/",
})

# Deploy
predictor = estimator.deploy(
    initial_instance_count=1,
    instance_type="ml.m5.xlarge",
)

# Predict
result = predictor.predict(input_data)

3.2 GCP Vertex AI

"""Google Cloud Vertex AI"""
from google.cloud import aiplatform

aiplatform.init(
    project="my-project",
    location="us-central1",
)

# Training
job = aiplatform.CustomTrainingJob(
    display_name="churn-training",
    script_path="src/models/train.py",
    container_uri="us-docker.pkg.dev/vertex-ai/training/pytorch-gpu.1-13:latest",
    requirements=["scikit-learn", "pandas"],
)

model = job.run(
    replica_count=1,
    machine_type="n1-standard-8",
    accelerator_type="NVIDIA_TESLA_T4",
    accelerator_count=1,
)

# Deploy
endpoint = model.deploy(
    machine_type="n1-standard-4",
    min_replica_count=1,
    max_replica_count=5,
)

# Predict
prediction = endpoint.predict(instances=[{"features": [1, 2, 3]}])

3.3 Compare Cloud Platforms

FeaturesAWS SageMakerGCP Vertex AIAzure ML
Ease of use⭐⭐⭐⭐⭐⭐⭐⭐⭐⭐
GPU availability⭐⭐⭐⭐⭐⭐⭐⭐⭐⭐⭐⭐⭐
AutoML✅ Autopilot✅ AutoML✅ AutoML
Experiment tracking✅ Built-in✅ Built-in✅ Built-in
Model registry✅✅✅
Feature store✅✅✅
Pricing💰💰💰💰💰💰💰💰
Best forEnterpriseResearch + ProdMicrosoft shops

4. Infrastructure as Code

# terraform/ml-infra/main.tf
# Infrastructure as Code với Terraform

provider "aws" {
  region = "us-east-1"
}

# S3 bucket cho data & models
resource "aws_s3_bucket" "ml_artifacts" {
  bucket = "my-ml-artifacts"
  
  versioning {
    enabled = true
  }
}

# ECR cho Docker images
resource "aws_ecr_repository" "ml_training" {
  name = "ml-training"
}

resource "aws_ecr_repository" "ml_serving" {
  name = "ml-serving"
}

# ECS Cluster cho serving
resource "aws_ecs_cluster" "ml_serving" {
  name = "ml-serving-cluster"
}

# SageMaker Endpoint
resource "aws_sagemaker_endpoint_configuration" "churn" {
  name = "churn-predictor-config"

  production_variants {
    variant_name           = "primary"
    model_name             = aws_sagemaker_model.churn.name
    initial_instance_count = 2
    instance_type          = "ml.m5.xlarge"
  }
}

# RDS cho MLflow
resource "aws_db_instance" "mlflow" {
  engine         = "postgres"
  instance_class = "db.t3.medium"
  allocated_storage = 20
  db_name        = "mlflow"
}

5. Best Practices

Docker:
  ✅ Multi-stage builds (build vs runtime)
  ✅ Pin dependency versions
  ✅ Non-root user
  ✅ Health checks
  ✅ .dockerignore (exclude data, notebooks)

Kubernetes:
  ✅ Resource requests & limits
  ✅ HPA for auto-scaling
  ✅ Readiness & liveness probes
  ✅ Node selectors cho GPU
  ✅ Separate namespaces (dev/staging/prod)

Cloud:
  ✅ Spot instances cho training (70% cheaper)
  ✅ Auto-scaling cho serving
  ✅ Monitoring & alerting
  ✅ Infrastructure as Code
  ✅ Cost monitoring & budgets

Summary

ConceptsRemember
DockerContainerize ML → reproducible environment
Multi-stageSeparate build vs runtime → smaller images
GPU Dockernvidia/cuda base image + NVIDIA Container Toolkit
KubernetesOrchestrate: Jobs (train), Deployments (serve), CronJobs (retrain)
HPAAuto-scale serving pods
SageMakerAWS managed ML, end-to-end
Vertex AIGCP managed ML, good for research
TerraformInfrastructure as Code

Exercises

  1. Docker: Create Dockerfile for ML serving (multi-stage, < 500MB image size).
  2. Compose: Create docker-compose with API + MLflow + Prometheus + Grafana.
  3. K8s: Deploy model serving to K8s cluster (minikube/kind). Add HPA.
  4. Cloud: Deploy a model to SageMaker or Vertex AI (free tier).

Next article: LLMOps vs MLOps — Paradigm Shift.