Chuyển đến nội dung chính

LECTURE 48: KUBERNETES FOR AI/ML WORKLOADS

Kubernetes AI/ML 2026: GPU scheduling with DRA (Dynamic Resource Allocation) GA K8s 1.34, time-slicing, MIG partitioning. Kubernetes Inference Extension (KIE), KEDA scale to zero, ResourceFlavor, Kueue batch scheduling.

🔒 DevSecOps — Lesson 48 LESSON 48: KUBERNETES FOR AI/ML WORKLOADS

KUBERNETES: FROM BASIC TO ADVANCED

Module 10: Cloud & Production

xdev.asia

🎯 Lesson Objective

Understand how Kubernetes 2026 supports AI/ML workloads: GPU scheduling with DRA, inference serving with KIE, batch training with Kueue and JobSet, and autoscaling by request queue.

1. Why Kubernetes for AI/ML?

  • GPU as a Service: share GPU cluster for multiple teams
  • Scale to zero: inference server does not receive requests → reduces to 0 pods, saves GPU
  • Batch scheduling: training jobs are queued and scheduled efficiently
  • Reproducibility: container images ensure a consistent environment
  • Multi-cloud portability: run training on any cloud with GPU

2. GPU Scheduling — Dynamic Resource Allocation (DRA) GA K8s 1.34

DRA replaces the old Device Plugin API, allowing more flexible GPU sharing:

# DeviceClass: define loại GPU available
apiVersion: resource.k8s.io/v1beta1
kind: DeviceClass
metadata:
  name: nvidia-gpu
spec:
  selectors:
  - cel:
      expression: device.driver == "gpu.nvidia.com" && device.attributes["memory"].quantity >= "40Gi"
# ResourceClaim: request GPU resources
apiVersion: resource.k8s.io/v1beta1
kind: ResourceClaim
metadata:
  name: gpu-claim
spec:
  devices:
    requests:
    - name: gpu
      deviceClassName: nvidia-gpu
      count: 2          # request 2 GPUs
      adminAccess: false
# Pod sử dụng ResourceClaim
apiVersion: v1
kind: Pod
metadata:
  name: training-job
spec:
  resourceClaims:
  - name: gpu-resource
    resourceClaimName: gpu-claim
  containers:
  - name: trainer
    image: nvcr.io/nvidia/pytorch:24.12-py3
    resources:
      claims:
      - name: gpu-resource
    command: ["python", "train.py"]

3. GPU Time-slicing and MIG

# Cách 1: Time-slicing (chia sẻ GPU theo thời gian)
# Phù hợp cho inference, batch jobs nhỏ

ConfigMap cho NVIDIA device plugin

cat <<EOF | kubectl apply -f - apiVersion: v1 kind: ConfigMap metadata: name: time-slicing-config namespace: kube-system data: any: |- version: v1 sharing: timeSlicing: resources: - name: nvidia.com/gpu replicas: 4 # 1 GPU vật lý chia thành 4 logical GPUs EOF

Cách 2: MIG (Multi-Instance GPU) — partition GPU thực sự

Chỉ hỗ trợ A100, H100, A30

MIG profiles: 1g.10gb, 2g.20gb, 3g.40gb, 7g.80gb

nvidia-smi mig -cgi 2g.20gb,2g.20gb,3g.40gb -C

Verify

kubectl describe node gpu-node | grep nvidia

nvidia.com/gpu: 3 (MIG instances)

nvidia.com/mig-2g.20gb: 2

nvidia.com/mig-3g.40gb: 1

4. Kubernetes Inference Extension (KIE)

KIE (2025 CNCF project): standardize LLM inference serving on Kubernetes:

# InferenceModel: define model spec
apiVersion: inference.networking.x-k8s.io/v1alpha2
kind: InferenceModel
metadata:
  name: llama3-8b
  namespace: ai
spec:
  modelName: "meta-llama/Llama-3-8B-Instruct"
  criticality: Standard   # Critical / Standard / Sheddable
  poolRef:
    name: llm-pool
# InferencePool: group of serving instances
apiVersion: inference.networking.x-k8s.io/v1alpha2
kind: InferencePool
metadata:
  name: llm-pool
spec:
  targetPortNumber: 8080
  selector:
    matchLabels:
      app: vllm-server
  extensionRef:
    name: kie-ext-proc   # gRPC extension for intelligent routing
# KIE features:
# - Per-request routing dựa trên model đang loaded
# - Prefix caching awareness (route requests tới server đã cache prompt)
# - Priority scheduling (critical requests trước)
# - LoRA adapter management

KIE + vLLM + K8s Gateway API

Gateway → KIE Extension Proc → vLLM pods

KIE biết mỗi vLLM đang cache gì → route hiệu quả hơn

5. Kueue — Batch Job Scheduling

Kueue (CNCF): fair queuing for batch workloads (training jobs, data processing):

# ResourceFlavor: define GPU types
apiVersion: kueue.x-k8s.io/v1beta1
kind: ResourceFlavor
metadata:
  name: nvidia-a100
spec:
  nodeLabels:
    accelerator: nvidia-a100
---
apiVersion: kueue.x-k8s.io/v1beta1
kind: ResourceFlavor
metadata:
  name: nvidia-h100
spec:
  nodeLabels:
    accelerator: nvidia-h100
# ClusterQueue: tổng capacity
apiVersion: kueue.x-k8s.io/v1beta1
kind: ClusterQueue
metadata:
  name: gpu-cluster-queue
spec:
  namespaceSelector: {}
  resourceGroups:
  - coveredResources: ["nvidia.com/gpu", "cpu", "memory"]
    flavors:
    - name: nvidia-h100
      resources:
      - name: nvidia.com/gpu
        nominalQuota: 16     # 16 H100 GPUs total
      - name: cpu
        nominalQuota: 512
      - name: memory
        nominalQuota: 2048Gi
    - name: nvidia-a100
      resources:
      - name: nvidia.com/gpu
        nominalQuota: 32     # 32 A100 GPUs total
# LocalQueue: per-team quota
apiVersion: kueue.x-k8s.io/v1beta1
kind: LocalQueue
metadata:
  name: team-nlp-queue
  namespace: nlp-team
spec:
  clusterQueueName: gpu-cluster-queue
# Job sử dụng Kueue
apiVersion: batch/v1
kind: Job
metadata:
  name: llama-finetune
  namespace: nlp-team
  labels:
    kueue.x-k8s.io/queue-name: team-nlp-queue
spec:
  completions: 1
  parallelism: 1
  template:
    spec:
      containers:
      - name: trainer
        image: nvcr.io/nvidia/pytorch:24.12-py3
        resources:
          requests:
            nvidia.com/gpu: "4"    # request 4 GPUs
            cpu: "32"
            memory: "256Gi"
        command: ["python", "finetune.py", "--model", "llama3-8b"]

6. KEDA — Scale to Zero for Inference

# KEDA ScaledObject: scale inference server dựa trên queue depth
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: vllm-scaler
  namespace: ai
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: vllm-llama3
  minReplicaCount: 0          # scale to zero khi không có request
  maxReplicaCount: 4
  pollingInterval: 15
  cooldownPeriod: 300         # đợi 5 phút trước khi scale down
  triggers:
  - type: prometheus
    metadata:
      serverAddress: http://prometheus.monitoring:9090
      metricName: vllm_request_queue_depth
      query: sum(vllm_num_requests_waiting)
      threshold: "5"          # scale up khi có > 5 requests chờ

7. Overview of AI/ML Stack on K8s 2026

Layer             Tools
────────────────────────────────────────────────────
Orchestration     Kubernetes 1.34+
GPU Scheduling    DRA (Dynamic Resource Allocation) GA
Training Jobs     JobSet (CNCF), Kueue (fair scheduling)
Training Framework PyTorch Distributed, JAX, DeepSpeed
Inference Serving vLLM, TGI (Text Generation Inference)
Inference Routing Kubernetes Inference Extension (KIE)
Autoscaling       KEDA (scale to zero)
Model Storage     OCI Artifacts, PVC với ReadOnlyMany
MLOps             MLflow, Kubeflow Pipelines, ZenML
Monitoring        Prometheus + Grafana + DCGM Exporter (GPU)

Summary

  • DRA GA K8s 1.34: flexible GPU sharing, replace Device Plugin API
  • MIG: hardware partitioning A100/H100 for stronger isolation__HTMLTAG_117___
  • Kueue: fair batch scheduling for GPU cluster, per-team quotas
  • KIE: intelligent LLM inference routing with caching awareness prefix__HTMLTAG_121___
  • KEDA: scale inference servers to 0 when idle → save GPU costs