🎯 Mục tiêu bài học
Hiểu cách Kubernetes 2026 hỗ trợ AI/ML workloads: GPU scheduling với DRA, inference serving với KIE, batch training với Kueue và JobSet, và autoscaling theo request queue.
1. Tại sao Kubernetes cho AI/ML?
- GPU as a Service: share GPU cluster cho nhiều teams
- Scale to zero: inference server không nhận request → giảm về 0 pods, tiết kiệm GPU
- Batch scheduling: training jobs được queue và schedule hiệu quả
- Reproducibility: container images đảm bảo môi trường consistent
- Multi-cloud portability: chạy training trên cloud bất kỳ có GPU
2. GPU Scheduling — Dynamic Resource Allocation (DRA) GA K8s 1.34
DRA thay thế Device Plugin API cũ, cho phép GPU sharing linh hoạt hơn:
# DeviceClass: define loại GPU available
apiVersion: resource.k8s.io/v1beta1
kind: DeviceClass
metadata:
name: nvidia-gpu
spec:
selectors:
- cel:
expression: device.driver == "gpu.nvidia.com" && device.attributes["memory"].quantity >= "40Gi"
# ResourceClaim: request GPU resources
apiVersion: resource.k8s.io/v1beta1
kind: ResourceClaim
metadata:
name: gpu-claim
spec:
devices:
requests:
- name: gpu
deviceClassName: nvidia-gpu
count: 2 # request 2 GPUs
adminAccess: false
# Pod sử dụng ResourceClaim
apiVersion: v1
kind: Pod
metadata:
name: training-job
spec:
resourceClaims:
- name: gpu-resource
resourceClaimName: gpu-claim
containers:
- name: trainer
image: nvcr.io/nvidia/pytorch:24.12-py3
resources:
claims:
- name: gpu-resource
command: ["python", "train.py"]
3. GPU Time-slicing và MIG
# Cách 1: Time-slicing (chia sẻ GPU theo thời gian) # Phù hợp cho inference, batch jobs nhỏConfigMap cho NVIDIA device plugin
cat <<EOF | kubectl apply -f - apiVersion: v1 kind: ConfigMap metadata: name: time-slicing-config namespace: kube-system data: any: |- version: v1 sharing: timeSlicing: resources: - name: nvidia.com/gpu replicas: 4 # 1 GPU vật lý chia thành 4 logical GPUs EOF
Cách 2: MIG (Multi-Instance GPU) — partition GPU thực sự
Chỉ hỗ trợ A100, H100, A30
MIG profiles: 1g.10gb, 2g.20gb, 3g.40gb, 7g.80gb
nvidia-smi mig -cgi 2g.20gb,2g.20gb,3g.40gb -C
Verify
kubectl describe node gpu-node | grep nvidia
nvidia.com/gpu: 3 (MIG instances)
nvidia.com/mig-2g.20gb: 2
nvidia.com/mig-3g.40gb: 1
4. Kubernetes Inference Extension (KIE)
KIE (2025 CNCF project): chuẩn hóa LLM inference serving trên Kubernetes:
# InferenceModel: define model spec
apiVersion: inference.networking.x-k8s.io/v1alpha2
kind: InferenceModel
metadata:
name: llama3-8b
namespace: ai
spec:
modelName: "meta-llama/Llama-3-8B-Instruct"
criticality: Standard # Critical / Standard / Sheddable
poolRef:
name: llm-pool
# InferencePool: group of serving instances
apiVersion: inference.networking.x-k8s.io/v1alpha2
kind: InferencePool
metadata:
name: llm-pool
spec:
targetPortNumber: 8080
selector:
matchLabels:
app: vllm-server
extensionRef:
name: kie-ext-proc # gRPC extension for intelligent routing
# KIE features: # - Per-request routing dựa trên model đang loaded # - Prefix caching awareness (route requests tới server đã cache prompt) # - Priority scheduling (critical requests trước) # - LoRA adapter managementKIE + vLLM + K8s Gateway API
Gateway → KIE Extension Proc → vLLM pods
KIE biết mỗi vLLM đang cache gì → route hiệu quả hơn
5. Kueue — Batch Job Scheduling
Kueue (CNCF): fair queuing cho batch workloads (training jobs, data processing):
# ResourceFlavor: define GPU types
apiVersion: kueue.x-k8s.io/v1beta1
kind: ResourceFlavor
metadata:
name: nvidia-a100
spec:
nodeLabels:
accelerator: nvidia-a100
---
apiVersion: kueue.x-k8s.io/v1beta1
kind: ResourceFlavor
metadata:
name: nvidia-h100
spec:
nodeLabels:
accelerator: nvidia-h100
# ClusterQueue: tổng capacity
apiVersion: kueue.x-k8s.io/v1beta1
kind: ClusterQueue
metadata:
name: gpu-cluster-queue
spec:
namespaceSelector: {}
resourceGroups:
- coveredResources: ["nvidia.com/gpu", "cpu", "memory"]
flavors:
- name: nvidia-h100
resources:
- name: nvidia.com/gpu
nominalQuota: 16 # 16 H100 GPUs total
- name: cpu
nominalQuota: 512
- name: memory
nominalQuota: 2048Gi
- name: nvidia-a100
resources:
- name: nvidia.com/gpu
nominalQuota: 32 # 32 A100 GPUs total
# LocalQueue: per-team quota
apiVersion: kueue.x-k8s.io/v1beta1
kind: LocalQueue
metadata:
name: team-nlp-queue
namespace: nlp-team
spec:
clusterQueueName: gpu-cluster-queue
# Job sử dụng Kueue
apiVersion: batch/v1
kind: Job
metadata:
name: llama-finetune
namespace: nlp-team
labels:
kueue.x-k8s.io/queue-name: team-nlp-queue
spec:
completions: 1
parallelism: 1
template:
spec:
containers:
- name: trainer
image: nvcr.io/nvidia/pytorch:24.12-py3
resources:
requests:
nvidia.com/gpu: "4" # request 4 GPUs
cpu: "32"
memory: "256Gi"
command: ["python", "finetune.py", "--model", "llama3-8b"]
6. KEDA — Scale to Zero cho Inference
# KEDA ScaledObject: scale inference server dựa trên queue depth
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
name: vllm-scaler
namespace: ai
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: vllm-llama3
minReplicaCount: 0 # scale to zero khi không có request
maxReplicaCount: 4
pollingInterval: 15
cooldownPeriod: 300 # đợi 5 phút trước khi scale down
triggers:
- type: prometheus
metadata:
serverAddress: http://prometheus.monitoring:9090
metricName: vllm_request_queue_depth
query: sum(vllm_num_requests_waiting)
threshold: "5" # scale up khi có > 5 requests chờ
7. Tổng quan AI/ML Stack trên K8s 2026
Layer Tools
────────────────────────────────────────────────────
Orchestration Kubernetes 1.34+
GPU Scheduling DRA (Dynamic Resource Allocation) GA
Training Jobs JobSet (CNCF), Kueue (fair scheduling)
Training Framework PyTorch Distributed, JAX, DeepSpeed
Inference Serving vLLM, TGI (Text Generation Inference)
Inference Routing Kubernetes Inference Extension (KIE)
Autoscaling KEDA (scale to zero)
Model Storage OCI Artifacts, PVC với ReadOnlyMany
MLOps MLflow, Kubeflow Pipelines, ZenML
Monitoring Prometheus + Grafana + DCGM Exporter (GPU)
Tóm tắt
- DRA GA K8s 1.34: flexible GPU sharing, thay Device Plugin API
- MIG: hardware partitioning A100/H100 cho isolation mạnh hơn
- Kueue: fair batch scheduling cho GPU cluster, per-team quotas
- KIE: intelligent LLM inference routing với prefix caching awareness
- KEDA: scale inference servers về 0 khi idle → tiết kiệm GPU costs