🎯 LESSON OBJECTIVE__HTMLTAG_66___
- ✅ Resource requests vs limits deep dive
- ✅ QoS classes (Guaranteed, Burstable, BestEffort)
- ✅ LimitRange and ResourceQuota
- ✅ Node affinity, anti-affinity, taints/tolerations
- ✅ Topology spread constraints
- ✅ Pod Priority and Preemption
PART 1: RESOURCE REQUESTS & LIMITS
Resource Model:
Node Capacity: 8 CPU, 32GB RAM
│
├── kube-reserved: 0.5 CPU, 1GB (kubelet, containerd)
├── system-reserved: 0.5 CPU, 1GB (OS processes)
├── eviction-threshold: 100Mi (memory.available < 100Mi → evict)
│
└── Allocatable: 7 CPU, 30GB (available for pods)
Pod scheduling:
requests ≤ allocatable → pod can be scheduled
actual usage > limits → container OOMKilled (memory) or throttled (CPU)
requests limits
CPU: guaranteed cycles max throttle point
Memory: guaranteed memory OOM kill threshold
| QoS Class | Condition_ | Eviction Priority |
|---|---|---|
| Guaranteed | requests == limits (all containers) | Last to evict_ |
| Burstable | requests < limits | Middle |
| BestEffort | No requests or limits_ | First to evict |
# Guaranteed QoS (production databases):
containers:
- name: postgresql
resources:
requests:
cpu: "2"
memory: 4Gi
limits:
cpu: "2"
memory: 4Gi
# Burstable QoS (most apps):
containers:
- name: order-service
resources:
requests:
cpu: 100m
memory: 128Mi
limits:
cpu: 500m
memory: 512Mi
PART 2: LIMITRANGE & RESOURCEQUOTA
# LimitRange (per container defaults):
apiVersion: v1
kind: LimitRange
metadata:
name: default-limits
namespace: default
spec:
limits:
- type: Container
default:
cpu: 200m
memory: 256Mi
defaultRequest:
cpu: 100m
memory: 128Mi
max:
cpu: "4"
memory: 8Gi
min:
cpu: 50m
memory: 64Mi
---
# ResourceQuota (per namespace total):
apiVersion: v1
kind: ResourceQuota
metadata:
name: namespace-quota
namespace: default
spec:
hard:
requests.cpu: "16"
requests.memory: 32Gi
limits.cpu: "32"
limits.memory: 64Gi
pods: "100"
persistentvolumeclaims: "20"
services.loadbalancers: "5"
PART 3: NODE AFFINITY & ANTI-AFFINITY
# Node affinity: schedule on specific hardware:
apiVersion: apps/v1
kind: Deployment
metadata:
name: gpu-ml-service
spec:
template:
spec:
affinity:
# Node affinity: prefer SSD nodes:
nodeAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
nodeSelectorTerms:
- matchExpressions:
- key: node-type
operator: In
values: ["compute"]
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 100
preference:
matchExpressions:
- key: disk-type
operator: In
values: ["ssd"]
# Pod anti-affinity: spread across nodes:
podAntiAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
- labelSelector:
matchExpressions:
- key: app
operator: In
values: ["gpu-ml-service"]
topologyKey: kubernetes.io/hostname
# Topology spread: distribute evenly across zones/nodes:
apiVersion: apps/v1
kind: Deployment
metadata:
name: order-service
spec:
replicas: 6
template:
spec:
topologySpreadConstraints:
- maxSkew: 1
topologyKey: kubernetes.io/hostname
whenUnsatisfiable: DoNotSchedule
labelSelector:
matchLabels:
app: order-service
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: ScheduleAnyway
labelSelector:
matchLabels:
app: order-service
PART 4: TAINTS & TOLERATIONS
# Taint nodes for dedicated workloads:
kubectl taint nodes worker-db-01 dedicated=database:NoSchedule
kubectl taint nodes worker-db-02 dedicated=database:NoSchedule
# Toleration: only database pods on DB nodes:
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: postgresql
spec:
template:
spec:
tolerations:
- key: dedicated
operator: Equal
value: database
effect: NoSchedule
nodeSelector:
dedicated: database
PART 5: POD PRIORITY & PREEMPTION
# Priority classes:
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
name: critical-production
value: 1000000
globalDefault: false
description: "Critical production workloads"
preemptionPolicy: PreemptLowerPriority
---
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
name: default-production
value: 100000
globalDefault: true
description: "Default production priority"
---
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
name: batch-jobs
value: 10000
description: "Batch jobs, preemptible"
preemptionPolicy: Never
---
# Use in Deployment:
apiVersion: apps/v1
kind: Deployment
metadata:
name: payment-service
spec:
template:
spec:
priorityClassName: critical-production
💡 KEY TAKEAWAYS
- Requests: Scheduling guarantee; Limits: Hard ceiling
- QoS: Guaranteed for databases, Burstable for apps
- LimitRange: Default/max per container; ResourceQuota: Total per namespace
- Topology spread: Even distribution across nodes/zones
- Taints: Dedicate nodes for specific workloads
- Priority: Critical services preempt batch jobs
🎯 EXERCISES__HTMLTAG_171___
Exercise 1: Resource Governance__HTMLTAG_173___
- Create LimitRange and ResourceQuota for namespace__HTMLTAG_176___
- Deploy pod without resources → verify defaults applied
- Exceed quota → verify pod rejected
Exercise 2: Scheduling
- Taint 2 nodes for database workloads
- Configure topology spread for 6-replica deployment__HTMLTAG_188___
- Create priority classes, test preemption__HTMLTAG_190___
📚 NEXT POST
In Lesson 43: Multi-Tenancy & Namespace Isolation, we will implement multi-tenant architecture on shared cluster.