Chuyển đến nội dung chính

LESSON 21: DYNAMIC RESOURCE ALLOCATION (DRA) — GA K8S 1.34

Dynamic Resource Allocation (DRA) GA from K8s 1.34 — replaces old extended resources. ResourceClaim, DeviceClass, GPU sharing and FPGA allocation. NVIDIA GPU Operator with DRA for AI/ML workloads.

🔒 DevSecOps — Lesson 21 LESSON 21: DYNAMIC RESOURCE ALLOCATION (DRA) — GA K8S 1.34

KUBERNETES: FROM BASIC TO ADVANCED

Module 5: Workload Management__HTMLTAG_62___

xdev.asia

🎯 Lesson Objective_

Understand what Dynamic Resource Allocation (DRA) GA in K8s 1.34 is, why it is better than the old extended resources, how to use ResourceClaim and DeviceClass to allocate GPU, FPGA for AI/ML workloads.

1. Problem with Extended Resources Old

Before DRA, Kubernetes used extended resources to manage GPU:

resources:
  limits:
    nvidia.com/gpu: 1   # request 1 GPU

Disadvantages:

  • All-or-nothing: cannot share GPU between multiple Pods
  • No structured parameters: cannot specify GPU type, memory, MIG partition
  • No topology awareness: don't know which GPU is on which PCIe switch
  • No deallocation hooks: no cleanup when Pod ends

2. Dynamic Resource Allocation (DRA) — GA K8s 1.34

DRA provides a flexible API to allocate hardware resources with structured parameters.

2.1 DRA Architecture

  • DeviceClass: defines the device type (GPU, FPGA, NIC) — created by infrastructure team__HTMLTAG_109___
  • ResourceClaim: device-specific request — created by app team
  • ResourceClaimTemplate: create ResourceClaim for each Pod in Deployment/StatefulSet
  • ResourceSlice: information about available devices on each Node (published by the driver)

2.2 DeviceClass

apiVersion: resource.k8s.io/v1alpha3
kind: DeviceClass
metadata:
  name: gpu.nvidia.com
spec:
  selectors:
  # Chỉ chọn NVIDIA GPUs
  - cel:
      expression: device.driver == "gpu.nvidia.com"
  config:
  - opaque:
      driver: gpu.nvidia.com
      parameters:
        apiVersion: gpu.resource.nvidia.com/v1alpha1
        kind: GpuConfig
        sharing:
          strategy: TimeSlicing
          timeSlicingConfig:
            interval: Default

2.3 ResourceClaim — GPU specific requirements__HTMLTAG_126___
apiVersion: resource.k8s.io/v1alpha3
kind: ResourceClaim
metadata:
  name: my-gpu-claim
  namespace: ml-team
spec:
  devices:
    requests:
    - name: gpu
      deviceClassName: gpu.nvidia.com
      selectors:
      # Chỉ request GPU có ít nhất 40GB VRAM
      - cel:
          expression: device.attributes["gpu.nvidia.com"].memory.isGreaterThan(quantity("40Gi"))
      allocationMode: ExactCount
      count: 1

2.4 Pod uses ResourceClaim

apiVersion: v1
kind: Pod
metadata:
  name: ml-training
  namespace: ml-team
spec:
  resourceClaims:
  - name: gpu          # tên trong pod spec
    resourceClaimName: my-gpu-claim   # ResourceClaim tạo sẵn
  containers:
  - name: trainer
    image: nvcr.io/nvidia/pytorch:24.12-py3
    resources:
      claims:
      - name: gpu      # reference tên ở trên
    command: ["python", "train.py"]

2.5 ResourceClaimTemplate — For Deployments

apiVersion: resource.k8s.io/v1alpha3
kind: ResourceClaimTemplate
metadata:
  name: gpu-template
  namespace: ml-team
spec:
  spec:
    devices:
      requests:
      - name: gpu
        deviceClassName: gpu.nvidia.com
        count: 1
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: inference-server
spec:
  replicas: 3  # mỗi pod nhận 1 GPU riêng
  template:
    spec:
      resourceClaims:
      - name: gpu
        resourceClaimTemplateName: gpu-template
      containers:
      - name: inference
        image: inference:v1
        resources:
          claims:
          - name: gpu

3. GPU Time-Slicing with DRA

Share 1 GPU for multiple Pods with time-slicing:

apiVersion: resource.k8s.io/v1alpha3
kind: DeviceClass
metadata:
  name: gpu-shared.nvidia.com
spec:
  selectors:
  - cel:
      expression: device.driver == "gpu.nvidia.com"
  config:
  - opaque:
      driver: gpu.nvidia.com
      parameters:
        apiVersion: gpu.resource.nvidia.com/v1alpha1
        kind: GpuConfig
        sharing:
          strategy: TimeSlicing     # chia thời gian GPU
          timeSlicingConfig:
            interval: Default       # ~50ms mỗi lần

4. Multi-Instance GPU (MIG) Partitioning

MIG (Multi-Instance GPU) divides the A100/H100 GPU into independent partitions:

apiVersion: resource.k8s.io/v1alpha3
kind: ResourceClaim
metadata:
  name: mig-1g-10gb
spec:
  devices:
    requests:
    - name: gpu-partition
      deviceClassName: gpu.nvidia.com
      selectors:
      # Request MIG 1g.10gb partition (1/7 của A100 80GB)
      - cel:
          expression: |
            device.attributes["gpu.nvidia.com"].migProfile == "1g.10gb"

5. NVIDIA GPU Operator with DRA

# Cài GPU Operator với DRA enabled
helm repo add nvidia https://helm.ngc.nvidia.com/nvidia
helm repo update

helm install gpu-operator nvidia/gpu-operator
--namespace gpu-operator
--create-namespace
--set driver.enabled=true
--set toolkit.enabled=true
--set devicePlugin.enabled=false \ # Tắt device plugin cũ --set driverManager.enabled=true
--set mig.strategy=mixed
--set nfd.enabled=true

Verify DRA

kubectl get resourceslice # Xem GPUs available kubectl get deviceclass

6. Comparison: Extended Resources vs DRA

Feature                  Extended Resources    DRA (K8s 1.34 GA)
─────────────────────────────────────────────────────────────────
GPU sharing              ❌ No                ✅ Time-slicing, MIG
Structured parameters    ❌ No                ✅ CEL expressions
Multi-device request     ❌ Limited           ✅ Multiple claims
Topology awareness       ❌ No                ✅ Yes
Cleanup hooks            ❌ No                ✅ Yes
Kubernetes version       1.8+                 1.34 GA

7. Other Use Cases

  • FPGA: accelerate inference, crypto, network processing
  • RDMA NIC: high-speed networking for distributed training
  • Custom ASICs: TPU, custom AI accelerators
  • SR-IOV Network Interfaces: virtual functions for high-performance networking

Summary

  • DRA GA K8s 1.34: replace extended resources with flexible allocation
  • DeviceClass: defines the hardware type__HTMLTAG_169___
  • ResourceClaim: specific request with CEL selectors
  • Supports GPU sharing: Time-slicing and MIG partitioning
  • NVIDIA GPU Operator: DRA support from latest version