Chuyển đến nội dung chính

Lesson 18: Docker & Microservices Architecture for AI

Docker for AI: multi-stage builds, GPU support, model caching. Docker Compose for local dev. Microservices patterns: API Gateway, service mesh. Message queue (Redis, RabbitMQ). Separating services: inference, embedding, retrieval, orchestration.

You've finished coding the AI ​​Agent, pushed it to GitHub — but when your colleague clones it, it "works on my machine". Model needs CUDA 12.1, LangChain requires Python 3.11, ChromaDB needs the latest SQLite... Docker solves this entire dependency hell, and Microservices Architecture helps you scale each part independently.

1. Why Docker for AI?

1.1. Dependency hell in AI/ML

AI projects have much more complex dependencies than regular web apps:

Python App thường:         AI/ML App:
├── flask==3.0             ├── torch==2.3.0+cu121
├── requests==2.31         ├── transformers==4.41
└── redis==5.0             ├── langchain==0.2.5
                           ├── chromadb==0.5.0
   ~5 packages                ├── sentence-transformers==3.0
   ~50MB                      ├── opencv-python==4.10
                           ├── numpy==1.26 (conflict với torch?)
                           ├── CUDA 12.1 runtime
                           └── cuDNN 8.9
                              ~200+ packages
                              ~5-15GB

1.2. What does Docker solve?

ProblemNo DockerThere is Docker
Python version conflictpyenv, conda puppet1 Python container per container
CUDA version mismatchInstall system-wideCUDA in image
"Works on my machine"Always happensImage is the same everywhere
Reproducibilityrequirements.txt missingDockerfile = full recipe
Scale horizontalManually install each serverdocker compose scale
Rollback deploymentFeardocker rollback 1 command

1.3. Containers vs VM for AI workloads

┌──────────────────────────────────────────────────┐
│              Virtual Machine                      │
│  ┌────────┐  ┌────────┐  ┌────────┐             │
│  │ App A  │  │ App B  │  │ App C  │             │
│  │ Bins/  │  │ Bins/  │  │ Bins/  │             │
│  │ Libs   │  │ Libs   │  │ Libs   │             │
│  │ OS     │  │ OS     │  │ OS     │  ← Mỗi VM  │
│  └────────┘  └────────┘  └────────┘    1 OS     │
│           Hypervisor                              │
│           Host OS                                 │
│           Hardware (GPU)                          │
└──────────────────────────────────────────────────┘

┌──────────────────────────────────────────────────┐
│              Container (Docker)                   │
│  ┌────────┐  ┌────────┐  ┌────────┐             │
│  │ App A  │  │ App B  │  │ App C  │             │
│  │ Bins/  │  │ Bins/  │  │ Bins/  │             │
│  │ Libs   │  │ Libs   │  │ Libs   │             │
│  └────────┘  └────────┘  └────────┘             │
│           Docker Engine                           │
│           Host OS ← Share kernel                  │
│           Hardware (GPU passthrough)              │
└──────────────────────────────────────────────────┘

Key insight: Container shares kernel with host → almost zero overhead. GPU passthrough via NVIDIA Container Toolkit → performance = bare metal.

2. Dockerfile for AI Services

2.1. Multi-stage build — reduces image size

Multi-stage build separates build dependencies (gcc, build tools) from runtime, reducing image size from 5GB to 1-2GB:

# ============================================
# Stage 1: Builder — cài dependencies
# ============================================
FROM python:3.11-slim AS builder

WORKDIR /app

# Cài build tools (chỉ cần trong build stage)
RUN apt-get update && apt-get install -y --no-install-recommends \
    build-essential \
    gcc \
    && rm -rf /var/lib/apt/lists/*

# Copy requirements trước (layer caching)
COPY requirements.txt .
RUN pip install --no-cache-dir --prefix=/install -r requirements.txt

# ============================================
# Stage 2: Runtime — chỉ giữ những gì cần
# ============================================
FROM python:3.11-slim AS runtime

WORKDIR /app

# Copy installed packages từ builder
COPY --from=builder /install /usr/local

# Copy source code
COPY src/ ./src/
COPY config/ ./config/

# Non-root user (security best practice)
RUN useradd -m -r appuser && chown -R appuser:appuser /app
USER appuser

EXPOSE 8000

CMD ["uvicorn", "src.main:app", "--host", "0.0.0.0", "--port", "8000"]

2.2. Layer caching — optimize build time

Docker caches each layer. Copy order is very important — small changes should be placed first:

# ❌ BAD: Copy tất cả → mỗi lần sửa code đều reinstall packages
COPY . .
RUN pip install -r requirements.txt

# ✅ GOOD: Copy requirements trước → chỉ reinstall khi thêm package
COPY requirements.txt .
RUN pip install -r requirements.txt
COPY src/ ./src/
Layer cache visualization:

COPY requirements.txt .     ← Cached (không đổi)
RUN pip install ...          ← Cached (requirements không đổi)
COPY src/ ./src/             ← Rebuilt  (code thay đổi)
CMD [...]                    ← Rebuilt

→ Build time: 5 phút → 10 giây (khi chỉ sửa code)

2.3. .dockerignore for AI projects

# .dockerignore
__pycache__/
*.pyc
*.pyo

# Model files (mount volume thay vì bake vào image)
models/
*.bin
*.safetensors
*.gguf

# Dev/test files
.git/
.github/
tests/
notebooks/
*.ipynb

# Environment
.env
.env.*
.venv/
venv/

# Data files
data/raw/
data/tmp/
*.csv
*.parquet

# IDE
.vscode/
.idea/

# Docker
docker-compose*.yml
Dockerfile*

3. GPU Support with Docker

3.1. NVIDIA Container Toolkit setup

# Cài NVIDIA Container Toolkit (Ubuntu/Debian)
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey \
  | sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg

curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list \
  | sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' \
  | sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list

sudo apt-get update
sudo apt-get install -y nvidia-container-toolkit

# Configure Docker runtime
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker

# Verify
docker run --rm --gpus all nvidia/cuda:12.1.0-base-ubuntu22.04 nvidia-smi

3.2. CUDA base images

Base ImageSizeUse Case
nvidia/cuda:12.1.0-base-ubuntu22.04~120MBOnly CUDA runtime
nvidia/cuda:12.1.0-runtime-ubuntu22.04~800MBRuntime + cuDNN
nvidia/cuda:12.1.0-devel-ubuntu22.04~3.5GBBuild custom CUDA ops
pytorch/pytorch:2.3.0-cuda12.1-cudnn8-runtime~5.5GBPyTorch pre-installed
# Dockerfile.gpu — AI inference service với GPU
FROM nvidia/cuda:12.1.0-runtime-ubuntu22.04 AS runtime

# Cài Python
RUN apt-get update && apt-get install -y --no-install-recommends \
    python3.11 \
    python3.11-venv \
    python3-pip \
    && rm -rf /var/lib/apt/lists/*

WORKDIR /app

COPY requirements-gpu.txt .
RUN pip install --no-cache-dir -r requirements-gpu.txt

COPY src/ ./src/

# GPU memory settings
ENV CUDA_VISIBLE_DEVICES=0
ENV PYTORCH_CUDA_ALLOC_CONF=max_split_size_mb:512

EXPOSE 8000
CMD ["uvicorn", "src.main:app", "--host", "0.0.0.0", "--port", "8000"]
# Chạy container với GPU
docker run --gpus all -p 8000:8000 my-ai-service:gpu

# Chỉ định GPU cụ thể
docker run --gpus '"device=0,1"' -p 8000:8000 my-ai-service:gpu

4. Model Caching Strategies

AI models are usually 1-10GB — should not be baked into Docker images. There are 3 strategies:

4.1. Strategy comparison

StrategyBuild TimeCold StartDisk UsageBest For
Volume mountInstantHost dependentSharedDev, on-prem
Download at startupsFast buildsSlow (first time)Per-containerCloud, auto-scale
Model registry (MLflow)Fast buildsFast (cached)CentralizedProduction, teams

4.2. Volume mount — simplest

# docker-compose.yml
services:
  inference:
    build: .
    volumes:
      - ./models:/app/models:ro    # Read-only mount
      - model-cache:/root/.cache   # HuggingFace cache
    environment:
      - TRANSFORMERS_CACHE=/root/.cache

volumes:
  model-cache:   # Docker managed volume
# src/model_loader.py
import os
from pathlib import Path

MODEL_DIR = Path(os.getenv("MODEL_DIR", "/app/models"))

def load_model(model_name: str):
    """Load model từ mounted volume."""
    model_path = MODEL_DIR / model_name
    if not model_path.exists():
        raise FileNotFoundError(
            f"Model {model_name} not found. "
            f"Mount volume chứa model vào {MODEL_DIR}"
        )
    return AutoModelForCausalLM.from_pretrained(str(model_path))

4.3. Download at startup with caching

# src/startup.py
import hashlib
from huggingface_hub import snapshot_download

CACHE_DIR = "/app/model-cache"

def ensure_model(model_id: str, revision: str = "main") -> str:
    """Download model nếu chưa có trong cache."""
    cache_key = hashlib.sha256(
        f"{model_id}:{revision}".encode()
    ).hexdigest()[:12]

    model_path = f"{CACHE_DIR}/{cache_key}"

    if os.path.exists(f"{model_path}/config.json"):
        print(f"✅ Model {model_id} found in cache")
        return model_path

    print(f"⬇️  Downloading {model_id}...")
    snapshot_download(
        repo_id=model_id,
        revision=revision,
        local_dir=model_path,
        local_dir_use_symlinks=False,
    )
    return model_path

5. Docker Compose for AI Stack

5.1. Full AI development stack

# docker-compose.yml
version: "3.9"

services:
  # ============================================
  # Core AI API
  # ============================================
  api:
    build:
      context: .
      dockerfile: Dockerfile
    ports:
      - "8000:8000"
    environment:
      - OPENAI_API_KEY=${OPENAI_API_KEY}
      - REDIS_URL=redis://redis:6379/0
      - CHROMA_HOST=chromadb
      - CHROMA_PORT=8100
      - POSTGRES_URL=postgresql://ai:secret@postgres:5432/aidb
    volumes:
      - model-cache:/app/models
    depends_on:
      redis:
        condition: service_healthy
      chromadb:
        condition: service_healthy
      postgres:
        condition: service_healthy
    healthcheck:
      test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
      interval: 30s
      timeout: 10s
      retries: 3
      start_period: 40s

  # ============================================
  # Background Worker (async tasks)
  # ============================================
  worker:
    build: .
    command: >
      celery -A src.worker worker
      --loglevel=info
      --concurrency=2
      --queues=inference,embedding,document
    environment:
      - REDIS_URL=redis://redis:6379/0
      - CHROMA_HOST=chromadb
    volumes:
      - model-cache:/app/models
    depends_on:
      - redis
      - chromadb

  # ============================================
  # Vector Database
  # ============================================
  chromadb:
    image: chromadb/chroma:0.5.0
    ports:
      - "8100:8000"
    volumes:
      - chroma-data:/chroma/chroma
    environment:
      - ANONYMIZED_TELEMETRY=false
    healthcheck:
      test: ["CMD", "curl", "-f", "http://localhost:8000/api/v1/heartbeat"]
      interval: 10s
      timeout: 5s
      retries: 5

  # ============================================
  # Cache & Message Broker
  # ============================================
  redis:
    image: redis:7-alpine
    ports:
      - "6379:6379"
    volumes:
      - redis-data:/data
    command: redis-server --maxmemory 512mb --maxmemory-policy allkeys-lru
    healthcheck:
      test: ["CMD", "redis-cli", "ping"]
      interval: 10s
      timeout: 5s
      retries: 5

  # ============================================
  # Relational Database
  # ============================================
  postgres:
    image: postgres:16-alpine
    ports:
      - "5432:5432"
    environment:
      - POSTGRES_USER=ai
      - POSTGRES_PASSWORD=secret
      - POSTGRES_DB=aidb
    volumes:
      - postgres-data:/var/lib/postgresql/data
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U ai -d aidb"]
      interval: 10s
      timeout: 5s
      retries: 5

volumes:
  model-cache:
  chroma-data:
  redis-data:
  postgres-data:

5.2. Launch and manage

# Start toàn bộ stack
docker compose up -d

# Xem logs real-time
docker compose logs -f api worker

# Scale worker lên 3 instances
docker compose up -d --scale worker=3

# Restart 1 service
docker compose restart api

# Xem resource usage
docker compose stats

# Dừng và xoá data
docker compose down           # Giữ volumes
docker compose down -v        # Xoá cả volumes

6. Microservices Architecture for AI

6.1. Monolith vs Microservices — when to split?

Monolith AI Service:              Microservices AI:
┌─────────────────────┐     ┌──────────┐ ┌──────────┐
│  /chat               │     │ Chat API │ │ Embed    │
│  /embed              │     │ Service  │ │ Service  │
│  /search             │     └────┬─────┘ └────┬─────┘
│  /process-doc        │          │             │
│                      │     ┌────┴─────┐ ┌────┴─────┐
│  Model Loading       │     │ Inference│ │ Retrieval│
│  Embedding           │     │ Service  │ │ Service  │
│  Vector Search       │     └──────────┘ └──────────┘
│  Document Processing │
└─────────────────────┘     Mỗi service scale riêng
1 process, tất cả memory     GPU chỉ cho inference
CriteriaMonolithMicroservices
Team size1-3 devs4+ devs
Scale patternUniformPer-service (Dedicated GPU)
Deploy frequencyAll at onceEach service
ComplexityLowHigh (networking, discovery)
AI-specificModel shares memoryBetter model isolation
When to chooseMVP, prototype, POCProduction, multi-model

Practical rule: Start Monolith → separate Microservices when you need to scale inference separately, or when the team is > 4 people, or when deploy frequency is different between components.

6.2. Service decomposition for AI

┌─────────────────────────────────────────────────────────────┐
│                     API Gateway (Kong/Traefik)              │
│                 Rate limiting, Auth, Routing                │
└────┬──────────┬──────────┬──────────┬──────────┬────────────┘
     │          │          │          │          │
     ▼          ▼          ▼          ▼          ▼
┌─────────┐ ┌────────┐ ┌────────┐ ┌────────┐ ┌──────────┐
│ Orchest-│ │Inferen-│ │Embed-  │ │Retriev-│ │Document  │
│ rator   │ │ce SVC  │ │ding SVC│ │al SVC  │ │Processor │
│         │ │        │ │        │ │        │ │          │
│ Agent   │ │ LLM    │ │Sentence│ │Vector  │ │ PDF/DOCX │
│ Logic   │ │ Models │ │Transf. │ │Search  │ │ Chunking │
│ Routing │ │ GPU ⚡ │ │ GPU ⚡ │ │ChromaDB│ │ Parsing  │
└────┬────┘ └────────┘ └────────┘ └────────┘ └──────────┘
     │                                │            │
     ▼                                ▼            ▼
┌──────────┐                   ┌──────────┐  ┌──────────┐
│  Redis   │                   │ ChromaDB │  │PostgreSQL│
│ Cache +  │                   │ Vectors  │  │ Metadata │
│ Queue    │                   └──────────┘  └──────────┘
└──────────┘

Each service has 1 responsibility (Single Responsibility):

ServiceResponsibilityScale FactorGPU?
OrchestratorAgent logic, routing, tool callingCPU-bound, I/O-wait❌
InferenceLLM generation (chat, completion)GPU memory-bound✅
EmbeddingText → vectorsGPU compute-bound✅
RetrievalVector search, re-rankingMemory-bound❌
Document ProcessorParse, chunk, extractCPU-bound, I/O❌

6.3. Service interfaces

# Mỗi service expose REST API rõ ràng
# inference_service/main.py
from fastapi import FastAPI

app = FastAPI(title="Inference Service")

@app.post("/v1/generate")
async def generate(request: GenerateRequest) -> GenerateResponse:
    """LLM text generation."""
    ...

@app.post("/v1/generate/stream")
async def generate_stream(request: GenerateRequest):
    """Streaming generation via SSE."""
    ...

# embedding_service/main.py
app = FastAPI(title="Embedding Service")

@app.post("/v1/embed")
async def embed(request: EmbedRequest) -> EmbedResponse:
    """Text → vector embedding."""
    ...

@app.post("/v1/embed/batch")
async def embed_batch(request: BatchEmbedRequest) -> BatchEmbedResponse:
    """Batch embedding cho hiệu năng cao."""
    ...

7. API Gateway Pattern

7.1. Why do we need API Gateway?

Clients should not call each microservice directly — need 1 entry point to handle cross-cutting concerns:

Không Gateway:                  Có Gateway:
Client ──→ Auth Service         Client ──→ API Gateway ──→ Services
Client ──→ Inference SVC                   ├── Auth
Client ──→ Embedding SVC                   ├── Rate Limit
Client ──→ Retrieval SVC                   ├── Routing
   ↑ Client biết mọi service              ├── CORS
   ↑ Auth lặp lại mọi nơi                 └── Logging

7.2. Traefik configuration for AI services

# docker-compose.gateway.yml
services:
  traefik:
    image: traefik:v3.0
    command:
      - "--api.insecure=true"
      - "--providers.docker=true"
      - "--providers.docker.exposedbydefault=false"
      - "--entryPoints.web.address=:80"
    ports:
      - "80:80"
      - "8080:8080"    # Traefik dashboard
    volumes:
      - /var/run/docker.sock:/var/run/docker.sock:ro

  orchestrator:
    build: ./services/orchestrator
    labels:
      - "traefik.enable=true"
      - "traefik.http.routers.orchestrator.rule=PathPrefix(`/api/v1/chat`)"
      - "traefik.http.routers.orchestrator.entrypoints=web"
      - "traefik.http.services.orchestrator.loadbalancer.server.port=8000"
      # Rate limiting
      - "traefik.http.middlewares.ai-ratelimit.ratelimit.average=10"
      - "traefik.http.middlewares.ai-ratelimit.ratelimit.burst=20"
      - "traefik.http.routers.orchestrator.middlewares=ai-ratelimit"

  inference:
    build: ./services/inference
    deploy:
      replicas: 2
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]
    labels:
      - "traefik.enable=true"
      - "traefik.http.routers.inference.rule=PathPrefix(`/api/v1/generate`)"
      - "traefik.http.services.inference.loadbalancer.server.port=8001"

  embedding:
    build: ./services/embedding
    deploy:
      replicas: 3
    labels:
      - "traefik.enable=true"
      - "traefik.http.routers.embedding.rule=PathPrefix(`/api/v1/embed`)"
      - "traefik.http.services.embedding.loadbalancer.server.port=8002"

8. Message Queue for Async AI Tasks

8.1. Why do we need a message queue?

Many AI tasks do not require immediate response — document processing takes 30 seconds, batch embedding takes 5 minutes. HTTP timeout will kill the request.

Synchronous (❌ cho heavy tasks):
Client ──HTTP──→ API ──process 30s──→ Response
                      ↑ Connection timeout!

Asynchronous (✅):
Client ──HTTP──→ API ──enqueue──→ Response (task_id)
                        │
                  Queue ─┤
                        │
                Worker ──┘ process in background
Client ──HTTP──→ API ──check status──→ Result

8.2. Redis Streams for AI task queue

# src/queue/producer.py
import redis
import json
import uuid

redis_client = redis.Redis(host="redis", port=6379, db=0)

def enqueue_task(task_type: str, payload: dict) -> str:
    """Đẩy task vào Redis Stream."""
    task_id = str(uuid.uuid4())
    message = {
        "task_id": task_id,
        "type": task_type,
        "payload": json.dumps(payload),
        "status": "pending",
    }
    # XADD vào stream tương ứng
    redis_client.xadd(f"tasks:{task_type}", message)
    # Lưu status riêng để query nhanh
    redis_client.hset(f"task:{task_id}", mapping=message)
    return task_id

# API endpoint
@app.post("/api/v1/process-document")
async def process_document(file: UploadFile):
    """Upload document → queue → background processing."""
    content = await file.read()
    task_id = enqueue_task("document", {
        "filename": file.filename,
        "content_b64": base64.b64encode(content).decode(),
    })
    return {"task_id": task_id, "status": "queued"}

@app.get("/api/v1/tasks/{task_id}")
async def get_task_status(task_id: str):
    """Check task completion status."""
    status = redis_client.hgetall(f"task:{task_id}")
    return status
# src/queue/consumer.py
import redis
import json

redis_client = redis.Redis(host="redis", port=6379, db=0)
CONSUMER_GROUP = "ai-workers"
CONSUMER_NAME = f"worker-{os.getpid()}"

def process_tasks(stream: str):
    """Consumer group — nhiều worker xử lý song song."""
    # Tạo consumer group (idempotent)
    try:
        redis_client.xgroup_create(stream, CONSUMER_GROUP, id="0", mkstream=True)
    except redis.ResponseError:
        pass  # Group already exists

    while True:
        # Đọc message chưa ai xử lý
        messages = redis_client.xreadgroup(
            CONSUMER_GROUP, CONSUMER_NAME,
            {stream: ">"},
            count=1, block=5000  # Block 5s nếu không có message
        )

        for stream_name, stream_messages in messages:
            for msg_id, data in stream_messages:
                task_id = data[b"task_id"].decode()
                payload = json.loads(data[b"payload"])

                try:
                    # Xử lý task
                    result = handle_task(data[b"type"].decode(), payload)

                    # Update status
                    redis_client.hset(f"task:{task_id}", mapping={
                        "status": "completed",
                        "result": json.dumps(result),
                    })

                    # Acknowledge message
                    redis_client.xack(stream_name, CONSUMER_GROUP, msg_id)

                except Exception as e:
                    redis_client.hset(f"task:{task_id}", mapping={
                        "status": "failed",
                        "error": str(e),
                    })

8.3. Task types in AI system

Task TypePrioritiesAvg TimeWorkers
inference (chat)🔴 High2-10sGPU workers
embedding🟡 Medium0.5-2sGPU workers
document_process🟢 Low10-60sCPU workers
batch_inference🟢 Low1-30 minGPU workers (off-peak)
reindex⚪ Background5-60 minCPU workers

9. REST vs gRPC for AI Services

9.1. Compare details

CriteriaREST/JSONgRPC/Protobuf
SerializationJSON (text)Protobuf (binary)
Payload sizeLarge (~2-5x)Small, compact
Speed ​​(latency)~30% slowerFaster
StreamingSSE (1-way), WebSocketBi-directional native
Browser support✅ Native⚠️ Need grpc-web proxy
Debuggingcurl, Postmangrpcurl, Postman (new)
Code generationManual / OpenAPIAuto from .proto
Best for AIClient-facing APIService-to-service

9.2. Hybrid approach — most practical

External clients (browser, mobile):
  └──→ REST/JSON qua API Gateway

Internal service-to-service:
  └──→ gRPC cho tốc độ

┌──────────┐  REST   ┌──────────┐  gRPC   ┌──────────┐
│  Client  │────────→│ API GW / │────────→│Inference │
│ (React)  │  JSON   │ Orchest. │ Protobuf│ Service  │
└──────────┘         └──────────┘         └──────────┘
// protos/inference.proto
syntax = "proto3";

service InferenceService {
  rpc Generate (GenerateRequest) returns (GenerateResponse);
  rpc GenerateStream (GenerateRequest) returns (stream Token);
}

message GenerateRequest {
  string prompt = 1;
  string model = 2;
  float temperature = 3;
  int32 max_tokens = 4;
}

message Token {
  string text = 1;
  bool is_finished = 2;
}

10. Docker Networking & Service Discovery

10.1. Docker network basics

# docker-compose.yml
services:
  api:
    networks:
      - frontend    # Exposed to client
      - backend     # Internal services

  inference:
    networks:
      - backend     # KHÔNG expose ra ngoài

  chromadb:
    networks:
      - backend

  traefik:
    networks:
      - frontend
      - backend

networks:
  frontend:
    driver: bridge
  backend:
    driver: bridge
    internal: true    # Không route ra internet

10.2. Service discovery

In Docker Compose, each service automatically has DNS name = service name:

# Trong code, dùng service name thay vì IP
import os

# ✅ Docker DNS tự resolve
INFERENCE_URL = os.getenv("INFERENCE_URL", "http://inference:8001")
EMBEDDING_URL = os.getenv("EMBEDDING_URL", "http://embedding:8002")
REDIS_URL = os.getenv("REDIS_URL", "redis://redis:6379/0")
CHROMA_URL = os.getenv("CHROMA_URL", "http://chromadb:8000")

# ❌ KHÔNG hardcode IP
# INFERENCE_URL = "http://172.18.0.5:8001"
Docker DNS resolution:
┌──────────────────────────────────────────┐
│ Docker Network: ai-network               │
│                                          │
│  api ─────────→ inference:8001  ✅       │
│  api ─────────→ redis:6379      ✅       │
│  api ─────────→ chromadb:8000   ✅       │
│                                          │
│  DNS: service_name → container_IP        │
│  Load Balance: round-robin khi replicas  │
└──────────────────────────────────────────┘

11. Complete Docker Compose for AI Agent Stack

Summary from Lesson 16 (RAG Agent) + Lesson 17 (FastAPI):

# docker-compose.prod.yml
# Complete AI Agent Stack
version: "3.9"

x-common-env: &common-env
  REDIS_URL: redis://redis:6379/0
  CHROMA_HOST: chromadb
  CHROMA_PORT: 8000
  POSTGRES_URL: postgresql://ai:${DB_PASSWORD}@postgres:5432/aidb
  LOG_LEVEL: ${LOG_LEVEL:-info}

services:
  # ── Gateway ──────────────────────────────
  traefik:
    image: traefik:v3.0
    command:
      - "--providers.docker=true"
      - "--providers.docker.exposedbydefault=false"
      - "--entryPoints.web.address=:80"
      - "--accesslog=true"
      - "--metrics.prometheus=true"
    ports:
      - "80:80"
    volumes:
      - /var/run/docker.sock:/var/run/docker.sock:ro
    networks:
      - frontend
      - backend

  # ── Orchestrator (Agent logic) ───────────
  orchestrator:
    build:
      context: ./services/orchestrator
      dockerfile: Dockerfile
    environment:
      <<: *common-env
      INFERENCE_URL: http://inference:8001
      EMBEDDING_URL: http://embedding:8002
      RETRIEVAL_URL: http://retrieval:8003
    labels:
      - "traefik.enable=true"
      - "traefik.http.routers.orch.rule=PathPrefix(`/api/v1`)"
      - "traefik.http.services.orch.loadbalancer.server.port=8000"
    depends_on:
      redis:
        condition: service_healthy
    networks:
      - backend
    deploy:
      replicas: 2
      resources:
        limits:
          memory: 1G
          cpus: "1.0"
    healthcheck:
      test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
      interval: 30s
      timeout: 10s
      retries: 3

  # ── Inference Service (LLM) ──────────────
  inference:
    build: ./services/inference
    environment:
      <<: *common-env
      MODEL_NAME: ${MODEL_NAME:-gpt-4o-mini}
      OPENAI_API_KEY: ${OPENAI_API_KEY}
    networks:
      - backend
    deploy:
      replicas: 1
      resources:
        limits:
          memory: 4G
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]
    healthcheck:
      test: ["CMD", "curl", "-f", "http://localhost:8001/health"]
      interval: 15s
      timeout: 5s
      retries: 5

  # ── Embedding Service ────────────────────
  embedding:
    build: ./services/embedding
    environment:
      <<: *common-env
      EMBED_MODEL: sentence-transformers/all-MiniLM-L6-v2
    volumes:
      - model-cache:/app/models
    networks:
      - backend
    deploy:
      replicas: 2
    healthcheck:
      test: ["CMD", "curl", "-f", "http://localhost:8002/health"]
      interval: 15s
      timeout: 5s
      retries: 5

  # ── Retrieval Service ────────────────────
  retrieval:
    build: ./services/retrieval
    environment:
      <<: *common-env
    networks:
      - backend
    deploy:
      replicas: 2
    healthcheck:
      test: ["CMD", "curl", "-f", "http://localhost:8003/health"]
      interval: 15s
      timeout: 5s
      retries: 5

  # ── Worker (async tasks) ─────────────────
  worker:
    build: ./services/orchestrator
    command: >
      celery -A src.tasks worker
      --loglevel=info
      --queues=documents,batch
      --concurrency=4
    environment:
      <<: *common-env
    networks:
      - backend
    deploy:
      replicas: 2
      resources:
        limits:
          memory: 2G

  # ── Data Stores ──────────────────────────
  redis:
    image: redis:7-alpine
    command: >
      redis-server
      --maxmemory 1gb
      --maxmemory-policy allkeys-lru
      --appendonly yes
    volumes:
      - redis-data:/data
    networks:
      - backend
    healthcheck:
      test: ["CMD", "redis-cli", "ping"]
      interval: 10s
      timeout: 3s
      retries: 5

  chromadb:
    image: chromadb/chroma:0.5.0
    environment:
      - ANONYMIZED_TELEMETRY=false
    volumes:
      - chroma-data:/chroma/chroma
    networks:
      - backend
    healthcheck:
      test: ["CMD", "curl", "-f", "http://localhost:8000/api/v1/heartbeat"]
      interval: 10s
      timeout: 5s
      retries: 5

  postgres:
    image: postgres:16-alpine
    environment:
      POSTGRES_USER: ai
      POSTGRES_PASSWORD: ${DB_PASSWORD}
      POSTGRES_DB: aidb
    volumes:
      - postgres-data:/var/lib/postgresql/data
      - ./init-db:/docker-entrypoint-initdb.d
    networks:
      - backend
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U ai -d aidb"]
      interval: 10s
      timeout: 5s
      retries: 5

networks:
  frontend:
    driver: bridge
  backend:
    driver: bridge
    internal: true

volumes:
  model-cache:
  redis-data:
  chroma-data:
  postgres-data:
# .env file (KHÔNG commit vào git!)
OPENAI_API_KEY=sk-...
DB_PASSWORD=strong-random-password-here
MODEL_NAME=gpt-4o-mini
LOG_LEVEL=info

12. Health Checks, Logging & Monitoring

12.1. Health check pattern for AI services

# src/health.py
from fastapi import APIRouter, Response
from datetime import datetime
import asyncio

router = APIRouter()

@router.get("/health")
async def health_check():
    """Liveness probe — service còn sống không?"""
    return {"status": "ok", "timestamp": datetime.utcnow().isoformat()}

@router.get("/ready")
async def readiness_check():
    """Readiness probe — service sẵn sàng nhận request không?"""
    checks = {}

    # Check model loaded
    checks["model"] = model_manager.is_loaded()

    # Check Redis connection
    try:
        await redis_client.ping()
        checks["redis"] = True
    except Exception:
        checks["redis"] = False

    # Check vector DB
    try:
        await chroma_client.heartbeat()
        checks["chromadb"] = True
    except Exception:
        checks["chromadb"] = False

    all_healthy = all(checks.values())
    status_code = 200 if all_healthy else 503

    return Response(
        content=json.dumps({
            "status": "ready" if all_healthy else "not_ready",
            "checks": checks,
        }),
        status_code=status_code,
        media_type="application/json",
    )

12.2. Structured logging

# src/logging_config.py
import structlog
import logging

def setup_logging():
    """Structured JSON logging — dễ parse bởi ELK/Grafana."""
    structlog.configure(
        processors=[
            structlog.contextvars.merge_contextvars,
            structlog.processors.add_log_level,
            structlog.processors.TimeStamper(fmt="iso"),
            structlog.processors.JSONRenderer(),
        ],
        logger_factory=structlog.PrintLoggerFactory(),
    )

# Usage trong AI service
logger = structlog.get_logger()

@app.post("/api/v1/chat")
async def chat(request: ChatRequest):
    logger.info(
        "chat_request",
        model=request.model,
        message_count=len(request.messages),
        user_id=request.user_id,
    )

    start = time.time()
    response = await inference_service.generate(request)
    duration = time.time() - start

    logger.info(
        "chat_response",
        model=request.model,
        tokens_used=response.usage.total_tokens,
        duration_seconds=round(duration, 3),
        user_id=request.user_id,
    )
    return response

12.3. Container monitoring commands

# Real-time resource usage
docker stats --format "table {{.Name}}\t{{.CPUPerc}}\t{{.MemUsage}}\t{{.NetIO}}"

# Output:
# NAME          CPU %   MEM USAGE / LIMIT   NET I/O
# api           12.5%   256MiB / 1GiB       5.2MB / 1.1MB
# inference     85.3%   3.2GiB / 4GiB       12MB / 45MB
# worker        5.2%    512MiB / 2GiB       1.1MB / 0.5MB
# redis         0.5%    128MiB / 512MiB     8.5MB / 4.2MB
# chromadb      2.1%    1.1GiB / 2GiB       3.2MB / 1.8MB

# Xem logs với timestamps
docker compose logs --timestamps --tail=100 api

# Inspect container details
docker inspect --format='{{.State.Health.Status}}' ai-api-1

# GPU monitoring (trong container)
docker exec inference nvidia-smi --query-gpu=utilization.gpu,memory.used,memory.total \
  --format=csv,noheader,nounits

12.4. Monitoring architecture overview

┌─────────────────────────────────────────────────┐
│                 Grafana Dashboard                │
│  ┌──────────┐ ┌──────────┐ ┌──────────────┐   │
│  │ Request  │ │ GPU Usage│ │ Queue Depth  │   │
│  │ Latency  │ │ Memory   │ │ Task Status  │   │
│  └──────────┘ └──────────┘ └──────────────┘   │
└────────────────────┬────────────────────────────┘
                     │
              ┌──────┴──────┐
              │ Prometheus  │  ← Scrape /metrics
              └──────┬──────┘
                     │
    ┌────────────────┼────────────────┐
    │                │                │
┌───┴───┐     ┌─────┴────┐    ┌──────┴─────┐
│Traefik│     │ AI Svcs  │    │   Redis    │
│metrics│     │ /metrics │    │  Exporter  │
└───────┘     └──────────┘    └────────────┘

Summary

This article covers all containerization + microservices for AI systems:

  • ✅ Docker fundamentals — multi-stage builds, layer caching, .dockerignore optimized for AI
  • ✅ GPU support — NVIDIA Container Toolkit, CUDA base images, device reservation
  • ✅ Model caching — volume mount, download at startup, 3 strategies depending on use case
  • ✅ Docker Compose — full AI stack: API + Worker + VectorDB + Redis + PostgreSQL
  • ✅ Microservices decomposition — Orchestrator, Inference, Embedding, Retrieval, Document Processor
  • ✅ API Gateway — Traefik routing, rate limiting, load balancing
  • ✅ Message Queue — Redis Streams for async tasks (document processing, batch inference)
  • ✅ REST vs gRPC — hybrid approach: REST for client, gRPC for service-to-service
  • ✅ Networking — Docker networks, service discovery, DNS resolution
  • ✅ Monitoring — health checks, structured logging, Prometheus + Grafana stack

Next article: Lesson 19 will cover CI/CD & MLOps — GitHub Actions, model versioning, automated testing, blue-green deployment for AI services.

Exercises

Exercise 1: Dockerfile for RAG Service (30 minutes)

Write multi-stage Dockerfile for RAG service from Lesson 16:

  • Stage 1: Install dependencies (requirements.txt has langchain, chromadb, sentence-transformers)
  • Stage 2: Copy code, non-root user, health check
  • Create a suitable .dockerignore
  • Build and test run: docker build -t rag-service . && docker run -p 8000:8000 rag-service

Exercise 2: Docker Compose AI Stack (45 minutes)

Create docker-compose.yml for system:

  • api: FastAPI service from Lesson 17 (port 8000)
  • chromadb: Vector database (port 8100)
  • redis: Cache + message broker
  • postgres: Metadata storage
  • Health checks for all services
  • Shared network, environment variables from .env file

Exercise 3: Microservices Decomposition (60 minutes)

Split the monolith AI service into 3 microservices:

  1. orchestrator — receives requests, calls other services, returns responses
  2. embedding-service — text → vector (endpoint /v1/embed)
  3. retrieval-service — vector search from ChromaDB (endpoint /v1/search)

Requirements:

  • Each service has its own Dockerfile
  • Docker Compose connects it all
  • Orchestrator calls embedding + retrieval over HTTP
  • Test with curl: send query → receive RAG results

Exercise 4: Async Task Queue (45 minutes)

Implement document processing pipeline:

  • POST /upload — upload PDF, pay task_id
  • Redis Stream as queue
  • Worker reads from queue, parses PDF, chunks, embeds, stores into ChromaDB
  • GET /tasks/{task_id} — check status (pending/processing/completed/failed)
  • Test with 3 documents uploaded at the same time