Chuyển đến nội dung chính

LESSON 42: RESOURCE MANAGEMENT & SCHEDULING

Kubernetes resource requests/limits, QoS classes, LimitRange, ResourceQuota, node affinity/anti-affinity, taints & tolerations, topology spread, and scheduling best practices.

🔒 DevSecOps — Lesson 42 LESSON 42: RESOURCE MANAGEMENT & SCHEDULING

Deploy Microservices On-Premises with Kubernetes HA

Part 10: Deployment Patterns & Auto-Scaling

xdev.asia

🎯 LESSON OBJECTIVE__HTMLTAG_66___
  • ✅ Resource requests vs limits deep dive
  • ✅ QoS classes (Guaranteed, Burstable, BestEffort)
  • ✅ LimitRange and ResourceQuota
  • ✅ Node affinity, anti-affinity, taints/tolerations
  • ✅ Topology spread constraints
  • ✅ Pod Priority and Preemption

PART 1: RESOURCE REQUESTS & LIMITS


Resource Model:

Node Capacity: 8 CPU, 32GB RAM
│
├── kube-reserved:    0.5 CPU, 1GB  (kubelet, containerd)
├── system-reserved:  0.5 CPU, 1GB  (OS processes)
├── eviction-threshold: 100Mi       (memory.available < 100Mi → evict)
│
└── Allocatable:      7 CPU, 30GB   (available for pods)

Pod scheduling:
  requests ≤ allocatable  → pod can be scheduled
  actual usage > limits   → container OOMKilled (memory) or throttled (CPU)

                requests         limits
CPU:        guaranteed cycles   max throttle point
Memory:     guaranteed memory   OOM kill threshold
QoS ClassCondition_Eviction Priority
Guaranteedrequests == limits (all containers)Last to evict_
Burstablerequests < limitsMiddle
BestEffortNo requests or limits_First to evict
# Guaranteed QoS (production databases):
containers:
  - name: postgresql
    resources:
      requests:
        cpu: "2"
        memory: 4Gi
      limits:
        cpu: "2"
        memory: 4Gi

# Burstable QoS (most apps):
containers:
  - name: order-service
    resources:
      requests:
        cpu: 100m
        memory: 128Mi
      limits:
        cpu: 500m
        memory: 512Mi

PART 2: LIMITRANGE & RESOURCEQUOTA

# LimitRange (per container defaults):
apiVersion: v1
kind: LimitRange
metadata:
  name: default-limits
  namespace: default
spec:
  limits:
    - type: Container
      default:
        cpu: 200m
        memory: 256Mi
      defaultRequest:
        cpu: 100m
        memory: 128Mi
      max:
        cpu: "4"
        memory: 8Gi
      min:
        cpu: 50m
        memory: 64Mi

---
# ResourceQuota (per namespace total):
apiVersion: v1
kind: ResourceQuota
metadata:
  name: namespace-quota
  namespace: default
spec:
  hard:
    requests.cpu: "16"
    requests.memory: 32Gi
    limits.cpu: "32"
    limits.memory: 64Gi
    pods: "100"
    persistentvolumeclaims: "20"
    services.loadbalancers: "5"

PART 3: NODE AFFINITY & ANTI-AFFINITY

# Node affinity: schedule on specific hardware:
apiVersion: apps/v1
kind: Deployment
metadata:
  name: gpu-ml-service
spec:
  template:
    spec:
      affinity:
        # Node affinity: prefer SSD nodes:
        nodeAffinity:
          requiredDuringSchedulingIgnoredDuringExecution:
            nodeSelectorTerms:
              - matchExpressions:
                  - key: node-type
                    operator: In
                    values: ["compute"]
          preferredDuringSchedulingIgnoredDuringExecution:
            - weight: 100
              preference:
                matchExpressions:
                  - key: disk-type
                    operator: In
                    values: ["ssd"]

        # Pod anti-affinity: spread across nodes:
        podAntiAffinity:
          requiredDuringSchedulingIgnoredDuringExecution:
            - labelSelector:
                matchExpressions:
                  - key: app
                    operator: In
                    values: ["gpu-ml-service"]
              topologyKey: kubernetes.io/hostname
# Topology spread: distribute evenly across zones/nodes:
apiVersion: apps/v1
kind: Deployment
metadata:
  name: order-service
spec:
  replicas: 6
  template:
    spec:
      topologySpreadConstraints:
        - maxSkew: 1
          topologyKey: kubernetes.io/hostname
          whenUnsatisfiable: DoNotSchedule
          labelSelector:
            matchLabels:
              app: order-service
        - maxSkew: 1
          topologyKey: topology.kubernetes.io/zone
          whenUnsatisfiable: ScheduleAnyway
          labelSelector:
            matchLabels:
              app: order-service

PART 4: TAINTS & TOLERATIONS

# Taint nodes for dedicated workloads:
kubectl taint nodes worker-db-01 dedicated=database:NoSchedule
kubectl taint nodes worker-db-02 dedicated=database:NoSchedule
# Toleration: only database pods on DB nodes:
apiVersion: apps/v1
kind: StatefulSet
metadata:
  name: postgresql
spec:
  template:
    spec:
      tolerations:
        - key: dedicated
          operator: Equal
          value: database
          effect: NoSchedule
      nodeSelector:
        dedicated: database

PART 5: POD PRIORITY & PREEMPTION

# Priority classes:
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
  name: critical-production
value: 1000000
globalDefault: false
description: "Critical production workloads"
preemptionPolicy: PreemptLowerPriority

---
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
  name: default-production
value: 100000
globalDefault: true
description: "Default production priority"

---
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
  name: batch-jobs
value: 10000
description: "Batch jobs, preemptible"
preemptionPolicy: Never

---
# Use in Deployment:
apiVersion: apps/v1
kind: Deployment
metadata:
  name: payment-service
spec:
  template:
    spec:
      priorityClassName: critical-production

💡 KEY TAKEAWAYS

  1. Requests: Scheduling guarantee; Limits: Hard ceiling
  2. QoS: Guaranteed for databases, Burstable for apps
  3. LimitRange: Default/max per container; ResourceQuota: Total per namespace
  4. Topology spread: Even distribution across nodes/zones
  5. Taints: Dedicate nodes for specific workloads
  6. Priority: Critical services preempt batch jobs

🎯 EXERCISES__HTMLTAG_171___

Exercise 1: Resource Governance__HTMLTAG_173___
  • Create LimitRange and ResourceQuota for namespace__HTMLTAG_176___
  • Deploy pod without resources → verify defaults applied
  • Exceed quota → verify pod rejected

Exercise 2: Scheduling

  • Taint 2 nodes for database workloads
  • Configure topology spread for 6-replica deployment__HTMLTAG_188___
  • Create priority classes, test preemption__HTMLTAG_190___

📚 NEXT POST

In Lesson 43: Multi-Tenancy & Namespace Isolation, we will implement multi-tenant architecture on shared cluster.