🎯 Lesson Objective_
Understand what Dynamic Resource Allocation (DRA) GA in K8s 1.34 is, why it is better than the old extended resources, how to use ResourceClaim and DeviceClass to allocate GPU, FPGA for AI/ML workloads.
1. Problem with Extended Resources Old
Before DRA, Kubernetes used extended resources to manage GPU:
resources:
limits:
nvidia.com/gpu: 1 # request 1 GPU
Disadvantages:
- All-or-nothing: cannot share GPU between multiple Pods
- No structured parameters: cannot specify GPU type, memory, MIG partition
- No topology awareness: don't know which GPU is on which PCIe switch
- No deallocation hooks: no cleanup when Pod ends
2. Dynamic Resource Allocation (DRA) — GA K8s 1.34
DRA provides a flexible API to allocate hardware resources with structured parameters.
2.1 DRA Architecture
- DeviceClass: defines the device type (GPU, FPGA, NIC) — created by infrastructure team__HTMLTAG_109___
- ResourceClaim: device-specific request — created by app team
- ResourceClaimTemplate: create ResourceClaim for each Pod in Deployment/StatefulSet
- ResourceSlice: information about available devices on each Node (published by the driver)
2.2 DeviceClass
apiVersion: resource.k8s.io/v1alpha3
kind: DeviceClass
metadata:
name: gpu.nvidia.com
spec:
selectors:
# Chỉ chọn NVIDIA GPUs
- cel:
expression: device.driver == "gpu.nvidia.com"
config:
- opaque:
driver: gpu.nvidia.com
parameters:
apiVersion: gpu.resource.nvidia.com/v1alpha1
kind: GpuConfig
sharing:
strategy: TimeSlicing
timeSlicingConfig:
interval: Default
2.3 ResourceClaim — GPU specific requirements__HTMLTAG_126___
apiVersion: resource.k8s.io/v1alpha3
kind: ResourceClaim
metadata:
name: my-gpu-claim
namespace: ml-team
spec:
devices:
requests:
- name: gpu
deviceClassName: gpu.nvidia.com
selectors:
# Chỉ request GPU có ít nhất 40GB VRAM
- cel:
expression: device.attributes["gpu.nvidia.com"].memory.isGreaterThan(quantity("40Gi"))
allocationMode: ExactCount
count: 1
apiVersion: resource.k8s.io/v1alpha3
kind: ResourceClaim
metadata:
name: my-gpu-claim
namespace: ml-team
spec:
devices:
requests:
- name: gpu
deviceClassName: gpu.nvidia.com
selectors:
# Chỉ request GPU có ít nhất 40GB VRAM
- cel:
expression: device.attributes["gpu.nvidia.com"].memory.isGreaterThan(quantity("40Gi"))
allocationMode: ExactCount
count: 1
2.4 Pod uses ResourceClaim
apiVersion: v1
kind: Pod
metadata:
name: ml-training
namespace: ml-team
spec:
resourceClaims:
- name: gpu # tên trong pod spec
resourceClaimName: my-gpu-claim # ResourceClaim tạo sẵn
containers:
- name: trainer
image: nvcr.io/nvidia/pytorch:24.12-py3
resources:
claims:
- name: gpu # reference tên ở trên
command: ["python", "train.py"]
2.5 ResourceClaimTemplate — For Deployments
apiVersion: resource.k8s.io/v1alpha3
kind: ResourceClaimTemplate
metadata:
name: gpu-template
namespace: ml-team
spec:
spec:
devices:
requests:
- name: gpu
deviceClassName: gpu.nvidia.com
count: 1
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: inference-server
spec:
replicas: 3 # mỗi pod nhận 1 GPU riêng
template:
spec:
resourceClaims:
- name: gpu
resourceClaimTemplateName: gpu-template
containers:
- name: inference
image: inference:v1
resources:
claims:
- name: gpu
3. GPU Time-Slicing with DRA
Share 1 GPU for multiple Pods with time-slicing:
apiVersion: resource.k8s.io/v1alpha3
kind: DeviceClass
metadata:
name: gpu-shared.nvidia.com
spec:
selectors:
- cel:
expression: device.driver == "gpu.nvidia.com"
config:
- opaque:
driver: gpu.nvidia.com
parameters:
apiVersion: gpu.resource.nvidia.com/v1alpha1
kind: GpuConfig
sharing:
strategy: TimeSlicing # chia thời gian GPU
timeSlicingConfig:
interval: Default # ~50ms mỗi lần
4. Multi-Instance GPU (MIG) Partitioning
MIG (Multi-Instance GPU) divides the A100/H100 GPU into independent partitions:
apiVersion: resource.k8s.io/v1alpha3
kind: ResourceClaim
metadata:
name: mig-1g-10gb
spec:
devices:
requests:
- name: gpu-partition
deviceClassName: gpu.nvidia.com
selectors:
# Request MIG 1g.10gb partition (1/7 của A100 80GB)
- cel:
expression: |
device.attributes["gpu.nvidia.com"].migProfile == "1g.10gb"
5. NVIDIA GPU Operator with DRA
# Cài GPU Operator với DRA enabled helm repo add nvidia https://helm.ngc.nvidia.com/nvidia helm repo updatehelm install gpu-operator nvidia/gpu-operator
--namespace gpu-operator
--create-namespace
--set driver.enabled=true
--set toolkit.enabled=true
--set devicePlugin.enabled=false \ # Tắt device plugin cũ --set driverManager.enabled=true
--set mig.strategy=mixed
--set nfd.enabled=trueVerify DRA
kubectl get resourceslice # Xem GPUs available kubectl get deviceclass
6. Comparison: Extended Resources vs DRA
Feature Extended Resources DRA (K8s 1.34 GA)
─────────────────────────────────────────────────────────────────
GPU sharing ❌ No ✅ Time-slicing, MIG
Structured parameters ❌ No ✅ CEL expressions
Multi-device request ❌ Limited ✅ Multiple claims
Topology awareness ❌ No ✅ Yes
Cleanup hooks ❌ No ✅ Yes
Kubernetes version 1.8+ 1.34 GA
7. Other Use Cases
- FPGA: accelerate inference, crypto, network processing
- RDMA NIC: high-speed networking for distributed training
- Custom ASICs: TPU, custom AI accelerators
- SR-IOV Network Interfaces: virtual functions for high-performance networking
Summary
- DRA GA K8s 1.34: replace extended resources with flexible allocation
- DeviceClass: defines the hardware type__HTMLTAG_169___
- ResourceClaim: specific request with CEL selectors
- Supports GPU sharing: Time-slicing and MIG partitioning
- NVIDIA GPU Operator: DRA support from latest version