Chuyển đến nội dung chính

BÀI 21: DYNAMIC RESOURCE ALLOCATION (DRA) — GA K8S 1.34

Dynamic Resource Allocation (DRA) GA từ K8s 1.34 — thay thế extended resources cũ. ResourceClaim, DeviceClass, GPU sharing và FPGA allocation. NVIDIA GPU Operator với DRA cho AI/ML workloads.

🔒 DevSecOps — Bài 21 BÀI 21: DYNAMIC RESOURCE ALLOCATION (DRA) — GA K8S 1.34

KUBERNETES: TỪ CƠ BẢN ĐẾN NÂNG CAO

Module 5: Workload Management

xdev.asia

🎯 Mục tiêu bài học

Hiểu Dynamic Resource Allocation (DRA) GA trong K8s 1.34 là gì, tại sao tốt hơn extended resources cũ, cách dùng ResourceClaim và DeviceClass để allocate GPU, FPGA cho AI/ML workloads.

1. Vấn đề với Extended Resources Cũ

Trước DRA, Kubernetes dùng extended resources để quản lý GPU:

resources:
  limits:
    nvidia.com/gpu: 1   # request 1 GPU

Nhược điểm:

  • All-or-nothing: không thể share GPU giữa nhiều Pods
  • No structured parameters: không thể chỉ định loại GPU, memory, MIG partition
  • No topology awareness: không biết GPU nào nằm trên PCIe switch nào
  • No deallocation hooks: không cleanup khi Pod kết thúc

2. Dynamic Resource Allocation (DRA) — GA K8s 1.34

DRA cung cấp API linh hoạt để allocate hardware resources với structured parameters.

2.1 DRA Architecture

  • DeviceClass: định nghĩa loại device (GPU, FPGA, NIC) — do infrastructure team tạo
  • ResourceClaim: yêu cầu cụ thể cho device — do app team tạo
  • ResourceClaimTemplate: tạo ResourceClaim cho mỗi Pod trong Deployment/StatefulSet
  • ResourceSlice: thông tin về device available trên mỗi Node (do driver publish)

2.2 DeviceClass

apiVersion: resource.k8s.io/v1alpha3
kind: DeviceClass
metadata:
  name: gpu.nvidia.com
spec:
  selectors:
  # Chỉ chọn NVIDIA GPUs
  - cel:
      expression: device.driver == "gpu.nvidia.com"
  config:
  - opaque:
      driver: gpu.nvidia.com
      parameters:
        apiVersion: gpu.resource.nvidia.com/v1alpha1
        kind: GpuConfig
        sharing:
          strategy: TimeSlicing
          timeSlicingConfig:
            interval: Default

2.3 ResourceClaim — Yêu cầu GPU cụ thể

apiVersion: resource.k8s.io/v1alpha3
kind: ResourceClaim
metadata:
  name: my-gpu-claim
  namespace: ml-team
spec:
  devices:
    requests:
    - name: gpu
      deviceClassName: gpu.nvidia.com
      selectors:
      # Chỉ request GPU có ít nhất 40GB VRAM
      - cel:
          expression: device.attributes["gpu.nvidia.com"].memory.isGreaterThan(quantity("40Gi"))
      allocationMode: ExactCount
      count: 1

2.4 Pod sử dụng ResourceClaim

apiVersion: v1
kind: Pod
metadata:
  name: ml-training
  namespace: ml-team
spec:
  resourceClaims:
  - name: gpu          # tên trong pod spec
    resourceClaimName: my-gpu-claim   # ResourceClaim tạo sẵn
  containers:
  - name: trainer
    image: nvcr.io/nvidia/pytorch:24.12-py3
    resources:
      claims:
      - name: gpu      # reference tên ở trên
    command: ["python", "train.py"]

2.5 ResourceClaimTemplate — Cho Deployments

apiVersion: resource.k8s.io/v1alpha3
kind: ResourceClaimTemplate
metadata:
  name: gpu-template
  namespace: ml-team
spec:
  spec:
    devices:
      requests:
      - name: gpu
        deviceClassName: gpu.nvidia.com
        count: 1
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: inference-server
spec:
  replicas: 3  # mỗi pod nhận 1 GPU riêng
  template:
    spec:
      resourceClaims:
      - name: gpu
        resourceClaimTemplateName: gpu-template
      containers:
      - name: inference
        image: inference:v1
        resources:
          claims:
          - name: gpu

3. GPU Time-Slicing với DRA

Chia sẻ 1 GPU cho nhiều Pods với time-slicing:

apiVersion: resource.k8s.io/v1alpha3
kind: DeviceClass
metadata:
  name: gpu-shared.nvidia.com
spec:
  selectors:
  - cel:
      expression: device.driver == "gpu.nvidia.com"
  config:
  - opaque:
      driver: gpu.nvidia.com
      parameters:
        apiVersion: gpu.resource.nvidia.com/v1alpha1
        kind: GpuConfig
        sharing:
          strategy: TimeSlicing     # chia thời gian GPU
          timeSlicingConfig:
            interval: Default       # ~50ms mỗi lần

4. Multi-Instance GPU (MIG) Partitioning

MIG (Multi-Instance GPU) chia GPU A100/H100 thành các partitions độc lập:

apiVersion: resource.k8s.io/v1alpha3
kind: ResourceClaim
metadata:
  name: mig-1g-10gb
spec:
  devices:
    requests:
    - name: gpu-partition
      deviceClassName: gpu.nvidia.com
      selectors:
      # Request MIG 1g.10gb partition (1/7 của A100 80GB)
      - cel:
          expression: |
            device.attributes["gpu.nvidia.com"].migProfile == "1g.10gb"

5. NVIDIA GPU Operator với DRA

# Cài GPU Operator với DRA enabled
helm repo add nvidia https://helm.ngc.nvidia.com/nvidia
helm repo update

helm install gpu-operator nvidia/gpu-operator
--namespace gpu-operator
--create-namespace
--set driver.enabled=true
--set toolkit.enabled=true
--set devicePlugin.enabled=false \ # Tắt device plugin cũ --set driverManager.enabled=true
--set mig.strategy=mixed
--set nfd.enabled=true

Verify DRA

kubectl get resourceslice # Xem GPUs available kubectl get deviceclass

6. So sánh: Extended Resources vs DRA

Feature                  Extended Resources    DRA (K8s 1.34 GA)
─────────────────────────────────────────────────────────────────
GPU sharing              ❌ No                ✅ Time-slicing, MIG
Structured parameters    ❌ No                ✅ CEL expressions
Multi-device request     ❌ Limited           ✅ Multiple claims
Topology awareness       ❌ No                ✅ Yes
Cleanup hooks            ❌ No                ✅ Yes
Kubernetes version       1.8+                 1.34 GA

7. Use Cases khác

  • FPGA: accelerate inference, crypto, network processing
  • RDMA NIC: high-speed networking cho distributed training
  • Custom ASICs: TPU, custom AI accelerators
  • SR-IOV Network Interfaces: virtual functions cho high-performance networking

Tóm tắt

  • DRA GA K8s 1.34: thay extended resources với flexible allocation
  • DeviceClass: định nghĩa loại hardware
  • ResourceClaim: yêu cầu cụ thể với CEL selectors
  • Hỗ trợ GPU sharing: Time-slicing và MIG partitioning
  • NVIDIA GPU Operator: hỗ trợ DRA từ phiên bản mới nhất