Chuyển đến nội dung chính

BÀI 48: KUBERNETES CHO AI/ML WORKLOADS

Kubernetes AI/ML 2026: GPU scheduling với DRA (Dynamic Resource Allocation) GA K8s 1.34, time-slicing, MIG partitioning. Kubernetes Inference Extension (KIE), KEDA scale to zero, ResourceFlavor, Kueue batch scheduling.

🔒 DevSecOps — Bài 48 BÀI 48: KUBERNETES CHO AI/ML WORKLOADS

KUBERNETES: TỪ CƠ BẢN ĐẾN NÂNG CAO

Module 10: Cloud & Production

xdev.asia

🎯 Mục tiêu bài học

Hiểu cách Kubernetes 2026 hỗ trợ AI/ML workloads: GPU scheduling với DRA, inference serving với KIE, batch training với Kueue và JobSet, và autoscaling theo request queue.

1. Tại sao Kubernetes cho AI/ML?

  • GPU as a Service: share GPU cluster cho nhiều teams
  • Scale to zero: inference server không nhận request → giảm về 0 pods, tiết kiệm GPU
  • Batch scheduling: training jobs được queue và schedule hiệu quả
  • Reproducibility: container images đảm bảo môi trường consistent
  • Multi-cloud portability: chạy training trên cloud bất kỳ có GPU

2. GPU Scheduling — Dynamic Resource Allocation (DRA) GA K8s 1.34

DRA thay thế Device Plugin API cũ, cho phép GPU sharing linh hoạt hơn:

# DeviceClass: define loại GPU available
apiVersion: resource.k8s.io/v1beta1
kind: DeviceClass
metadata:
  name: nvidia-gpu
spec:
  selectors:
  - cel:
      expression: device.driver == "gpu.nvidia.com" && device.attributes["memory"].quantity >= "40Gi"
# ResourceClaim: request GPU resources
apiVersion: resource.k8s.io/v1beta1
kind: ResourceClaim
metadata:
  name: gpu-claim
spec:
  devices:
    requests:
    - name: gpu
      deviceClassName: nvidia-gpu
      count: 2          # request 2 GPUs
      adminAccess: false
# Pod sử dụng ResourceClaim
apiVersion: v1
kind: Pod
metadata:
  name: training-job
spec:
  resourceClaims:
  - name: gpu-resource
    resourceClaimName: gpu-claim
  containers:
  - name: trainer
    image: nvcr.io/nvidia/pytorch:24.12-py3
    resources:
      claims:
      - name: gpu-resource
    command: ["python", "train.py"]

3. GPU Time-slicing và MIG

# Cách 1: Time-slicing (chia sẻ GPU theo thời gian)
# Phù hợp cho inference, batch jobs nhỏ

ConfigMap cho NVIDIA device plugin

cat <<EOF | kubectl apply -f - apiVersion: v1 kind: ConfigMap metadata: name: time-slicing-config namespace: kube-system data: any: |- version: v1 sharing: timeSlicing: resources: - name: nvidia.com/gpu replicas: 4 # 1 GPU vật lý chia thành 4 logical GPUs EOF

Cách 2: MIG (Multi-Instance GPU) — partition GPU thực sự

Chỉ hỗ trợ A100, H100, A30

MIG profiles: 1g.10gb, 2g.20gb, 3g.40gb, 7g.80gb

nvidia-smi mig -cgi 2g.20gb,2g.20gb,3g.40gb -C

Verify

kubectl describe node gpu-node | grep nvidia

nvidia.com/gpu: 3 (MIG instances)

nvidia.com/mig-2g.20gb: 2

nvidia.com/mig-3g.40gb: 1

4. Kubernetes Inference Extension (KIE)

KIE (2025 CNCF project): chuẩn hóa LLM inference serving trên Kubernetes:

# InferenceModel: define model spec
apiVersion: inference.networking.x-k8s.io/v1alpha2
kind: InferenceModel
metadata:
  name: llama3-8b
  namespace: ai
spec:
  modelName: "meta-llama/Llama-3-8B-Instruct"
  criticality: Standard   # Critical / Standard / Sheddable
  poolRef:
    name: llm-pool
# InferencePool: group of serving instances
apiVersion: inference.networking.x-k8s.io/v1alpha2
kind: InferencePool
metadata:
  name: llm-pool
spec:
  targetPortNumber: 8080
  selector:
    matchLabels:
      app: vllm-server
  extensionRef:
    name: kie-ext-proc   # gRPC extension for intelligent routing
# KIE features:
# - Per-request routing dựa trên model đang loaded
# - Prefix caching awareness (route requests tới server đã cache prompt)
# - Priority scheduling (critical requests trước)
# - LoRA adapter management

KIE + vLLM + K8s Gateway API

Gateway → KIE Extension Proc → vLLM pods

KIE biết mỗi vLLM đang cache gì → route hiệu quả hơn

5. Kueue — Batch Job Scheduling

Kueue (CNCF): fair queuing cho batch workloads (training jobs, data processing):

# ResourceFlavor: define GPU types
apiVersion: kueue.x-k8s.io/v1beta1
kind: ResourceFlavor
metadata:
  name: nvidia-a100
spec:
  nodeLabels:
    accelerator: nvidia-a100
---
apiVersion: kueue.x-k8s.io/v1beta1
kind: ResourceFlavor
metadata:
  name: nvidia-h100
spec:
  nodeLabels:
    accelerator: nvidia-h100
# ClusterQueue: tổng capacity
apiVersion: kueue.x-k8s.io/v1beta1
kind: ClusterQueue
metadata:
  name: gpu-cluster-queue
spec:
  namespaceSelector: {}
  resourceGroups:
  - coveredResources: ["nvidia.com/gpu", "cpu", "memory"]
    flavors:
    - name: nvidia-h100
      resources:
      - name: nvidia.com/gpu
        nominalQuota: 16     # 16 H100 GPUs total
      - name: cpu
        nominalQuota: 512
      - name: memory
        nominalQuota: 2048Gi
    - name: nvidia-a100
      resources:
      - name: nvidia.com/gpu
        nominalQuota: 32     # 32 A100 GPUs total
# LocalQueue: per-team quota
apiVersion: kueue.x-k8s.io/v1beta1
kind: LocalQueue
metadata:
  name: team-nlp-queue
  namespace: nlp-team
spec:
  clusterQueueName: gpu-cluster-queue
# Job sử dụng Kueue
apiVersion: batch/v1
kind: Job
metadata:
  name: llama-finetune
  namespace: nlp-team
  labels:
    kueue.x-k8s.io/queue-name: team-nlp-queue
spec:
  completions: 1
  parallelism: 1
  template:
    spec:
      containers:
      - name: trainer
        image: nvcr.io/nvidia/pytorch:24.12-py3
        resources:
          requests:
            nvidia.com/gpu: "4"    # request 4 GPUs
            cpu: "32"
            memory: "256Gi"
        command: ["python", "finetune.py", "--model", "llama3-8b"]

6. KEDA — Scale to Zero cho Inference

# KEDA ScaledObject: scale inference server dựa trên queue depth
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: vllm-scaler
  namespace: ai
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: vllm-llama3
  minReplicaCount: 0          # scale to zero khi không có request
  maxReplicaCount: 4
  pollingInterval: 15
  cooldownPeriod: 300         # đợi 5 phút trước khi scale down
  triggers:
  - type: prometheus
    metadata:
      serverAddress: http://prometheus.monitoring:9090
      metricName: vllm_request_queue_depth
      query: sum(vllm_num_requests_waiting)
      threshold: "5"          # scale up khi có > 5 requests chờ

7. Tổng quan AI/ML Stack trên K8s 2026

Layer             Tools
────────────────────────────────────────────────────
Orchestration     Kubernetes 1.34+
GPU Scheduling    DRA (Dynamic Resource Allocation) GA
Training Jobs     JobSet (CNCF), Kueue (fair scheduling)
Training Framework PyTorch Distributed, JAX, DeepSpeed
Inference Serving vLLM, TGI (Text Generation Inference)
Inference Routing Kubernetes Inference Extension (KIE)
Autoscaling       KEDA (scale to zero)
Model Storage     OCI Artifacts, PVC với ReadOnlyMany
MLOps             MLflow, Kubeflow Pipelines, ZenML
Monitoring        Prometheus + Grafana + DCGM Exporter (GPU)

Tóm tắt

  • DRA GA K8s 1.34: flexible GPU sharing, thay Device Plugin API
  • MIG: hardware partitioning A100/H100 cho isolation mạnh hơn
  • Kueue: fair batch scheduling cho GPU cluster, per-team quotas
  • KIE: intelligent LLM inference routing với prefix caching awareness
  • KEDA: scale inference servers về 0 khi idle → tiết kiệm GPU costs