🎯 Lesson Objective
Understand how Kubernetes 2026 supports AI/ML workloads: GPU scheduling with DRA, inference serving with KIE, batch training with Kueue and JobSet, and autoscaling by request queue.
1. Why Kubernetes for AI/ML?
- GPU as a Service: share GPU cluster for multiple teams
- Scale to zero: inference server does not receive requests → reduces to 0 pods, saves GPU
- Batch scheduling: training jobs are queued and scheduled efficiently
- Reproducibility: container images ensure a consistent environment
- Multi-cloud portability: run training on any cloud with GPU
2. GPU Scheduling — Dynamic Resource Allocation (DRA) GA K8s 1.34
DRA replaces the old Device Plugin API, allowing more flexible GPU sharing:
# DeviceClass: define loại GPU available
apiVersion: resource.k8s.io/v1beta1
kind: DeviceClass
metadata:
name: nvidia-gpu
spec:
selectors:
- cel:
expression: device.driver == "gpu.nvidia.com" && device.attributes["memory"].quantity >= "40Gi"
# ResourceClaim: request GPU resources
apiVersion: resource.k8s.io/v1beta1
kind: ResourceClaim
metadata:
name: gpu-claim
spec:
devices:
requests:
- name: gpu
deviceClassName: nvidia-gpu
count: 2 # request 2 GPUs
adminAccess: false
# Pod sử dụng ResourceClaim
apiVersion: v1
kind: Pod
metadata:
name: training-job
spec:
resourceClaims:
- name: gpu-resource
resourceClaimName: gpu-claim
containers:
- name: trainer
image: nvcr.io/nvidia/pytorch:24.12-py3
resources:
claims:
- name: gpu-resource
command: ["python", "train.py"]
3. GPU Time-slicing and MIG
# Cách 1: Time-slicing (chia sẻ GPU theo thời gian) # Phù hợp cho inference, batch jobs nhỏConfigMap cho NVIDIA device plugin
cat <<EOF | kubectl apply -f - apiVersion: v1 kind: ConfigMap metadata: name: time-slicing-config namespace: kube-system data: any: |- version: v1 sharing: timeSlicing: resources: - name: nvidia.com/gpu replicas: 4 # 1 GPU vật lý chia thành 4 logical GPUs EOF
Cách 2: MIG (Multi-Instance GPU) — partition GPU thực sự
Chỉ hỗ trợ A100, H100, A30
MIG profiles: 1g.10gb, 2g.20gb, 3g.40gb, 7g.80gb
nvidia-smi mig -cgi 2g.20gb,2g.20gb,3g.40gb -C
Verify
kubectl describe node gpu-node | grep nvidia
nvidia.com/gpu: 3 (MIG instances)
nvidia.com/mig-2g.20gb: 2
nvidia.com/mig-3g.40gb: 1
4. Kubernetes Inference Extension (KIE)
KIE (2025 CNCF project): standardize LLM inference serving on Kubernetes:
# InferenceModel: define model spec
apiVersion: inference.networking.x-k8s.io/v1alpha2
kind: InferenceModel
metadata:
name: llama3-8b
namespace: ai
spec:
modelName: "meta-llama/Llama-3-8B-Instruct"
criticality: Standard # Critical / Standard / Sheddable
poolRef:
name: llm-pool
# InferencePool: group of serving instances
apiVersion: inference.networking.x-k8s.io/v1alpha2
kind: InferencePool
metadata:
name: llm-pool
spec:
targetPortNumber: 8080
selector:
matchLabels:
app: vllm-server
extensionRef:
name: kie-ext-proc # gRPC extension for intelligent routing
# KIE features: # - Per-request routing dựa trên model đang loaded # - Prefix caching awareness (route requests tới server đã cache prompt) # - Priority scheduling (critical requests trước) # - LoRA adapter managementKIE + vLLM + K8s Gateway API
Gateway → KIE Extension Proc → vLLM pods
KIE biết mỗi vLLM đang cache gì → route hiệu quả hơn
5. Kueue — Batch Job Scheduling
Kueue (CNCF): fair queuing for batch workloads (training jobs, data processing):
# ResourceFlavor: define GPU types
apiVersion: kueue.x-k8s.io/v1beta1
kind: ResourceFlavor
metadata:
name: nvidia-a100
spec:
nodeLabels:
accelerator: nvidia-a100
---
apiVersion: kueue.x-k8s.io/v1beta1
kind: ResourceFlavor
metadata:
name: nvidia-h100
spec:
nodeLabels:
accelerator: nvidia-h100
# ClusterQueue: tổng capacity
apiVersion: kueue.x-k8s.io/v1beta1
kind: ClusterQueue
metadata:
name: gpu-cluster-queue
spec:
namespaceSelector: {}
resourceGroups:
- coveredResources: ["nvidia.com/gpu", "cpu", "memory"]
flavors:
- name: nvidia-h100
resources:
- name: nvidia.com/gpu
nominalQuota: 16 # 16 H100 GPUs total
- name: cpu
nominalQuota: 512
- name: memory
nominalQuota: 2048Gi
- name: nvidia-a100
resources:
- name: nvidia.com/gpu
nominalQuota: 32 # 32 A100 GPUs total
# LocalQueue: per-team quota
apiVersion: kueue.x-k8s.io/v1beta1
kind: LocalQueue
metadata:
name: team-nlp-queue
namespace: nlp-team
spec:
clusterQueueName: gpu-cluster-queue
# Job sử dụng Kueue
apiVersion: batch/v1
kind: Job
metadata:
name: llama-finetune
namespace: nlp-team
labels:
kueue.x-k8s.io/queue-name: team-nlp-queue
spec:
completions: 1
parallelism: 1
template:
spec:
containers:
- name: trainer
image: nvcr.io/nvidia/pytorch:24.12-py3
resources:
requests:
nvidia.com/gpu: "4" # request 4 GPUs
cpu: "32"
memory: "256Gi"
command: ["python", "finetune.py", "--model", "llama3-8b"]
6. KEDA — Scale to Zero for Inference
# KEDA ScaledObject: scale inference server dựa trên queue depth
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
name: vllm-scaler
namespace: ai
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: vllm-llama3
minReplicaCount: 0 # scale to zero khi không có request
maxReplicaCount: 4
pollingInterval: 15
cooldownPeriod: 300 # đợi 5 phút trước khi scale down
triggers:
- type: prometheus
metadata:
serverAddress: http://prometheus.monitoring:9090
metricName: vllm_request_queue_depth
query: sum(vllm_num_requests_waiting)
threshold: "5" # scale up khi có > 5 requests chờ
7. Overview of AI/ML Stack on K8s 2026
Layer Tools
────────────────────────────────────────────────────
Orchestration Kubernetes 1.34+
GPU Scheduling DRA (Dynamic Resource Allocation) GA
Training Jobs JobSet (CNCF), Kueue (fair scheduling)
Training Framework PyTorch Distributed, JAX, DeepSpeed
Inference Serving vLLM, TGI (Text Generation Inference)
Inference Routing Kubernetes Inference Extension (KIE)
Autoscaling KEDA (scale to zero)
Model Storage OCI Artifacts, PVC với ReadOnlyMany
MLOps MLflow, Kubeflow Pipelines, ZenML
Monitoring Prometheus + Grafana + DCGM Exporter (GPU)
Summary
- DRA GA K8s 1.34: flexible GPU sharing, replace Device Plugin API
- MIG: hardware partitioning A100/H100 for stronger isolation__HTMLTAG_117___
- Kueue: fair batch scheduling for GPU cluster, per-team quotas
- KIE: intelligent LLM inference routing with caching awareness prefix__HTMLTAG_121___
- KEDA: scale inference servers to 0 when idle → save GPU costs