1. Pod Scheduling
The kube-scheduler selects a suitable node for each Pod. Scheduling is based on resources, constraints, and policies.
Node Selection Methods
| Method | Used for | Example |
|---|---|---|
| nodeSelector | Basic: select node by label | nodeSelector: {disk: ssd} |
| Node Affinity | Advanced: preferred/required rules | "prefer nodes in zone-a, required: arch=amd64" |
| Pod Affinity | Co-locate Pods | "schedule near Pods with label app=cache" |
| Pod Anti-Affinity | Spread Pods apart | "don't run 2 replicas on the same node" |
2. Taints & Tolerations
Taints & Tolerations:
Node: "I have a taint — only Pods that tolerate it can run here"
Pod: "I tolerate that taint — schedule me there"
kubectl taint nodes node1 gpu=true:NoSchedule
→ Only Pods with matching toleration are scheduled on node1
Taint Effects:
NoSchedule → New Pods won't be scheduled (existing stay)
PreferNoSchedule → Try to avoid scheduling (soft)
NoExecute → Evict existing Pods + prevent new (hard)
Exam tip: Taint is on the Node, Toleration is on the Pod. The Control Plane nodes have a default taint so that user Pods don't run there.
3. Resource Requests & Limits
resources:
requests: # Minimum required for scheduling
cpu: "250m" # 0.25 core
memory: "128Mi" # 128 MiB
limits: # Maximum the container can use
cpu: "500m" # 0.5 core
memory: "256Mi" # 256 MiB
| Field | Role | What happens if exceeded |
|---|---|---|
| requests | Minimum resource for scheduling | N/A — it's a minimum guarantee |
| limits | Maximum resource allowed | CPU: throttled. Memory: Pod OOMKilled |
Exam tip: requests are used by the scheduler to decide node placement. limits are enforced by the kubelet at runtime. No requests set = Pod can be scheduled on any node but may get evicted under resource pressure.
4. Autoscaling
| Type | What it scales | Based on | Needs |
|---|---|---|---|
| HPA (Horizontal Pod Autoscaler) | Number of Pod replicas | CPU/Memory/Custom metrics | metrics-server |
| VPA (Vertical Pod Autoscaler) | CPU/Memory requests per Pod | Historical usage | VPA controller |
| Cluster Autoscaler | Number of nodes | Unschedulable Pods due to insufficient resources | Cloud provider integration |
HPA Flow:
metrics-server → HPA controller → check current CPU %
If CPU > 70% target → scale up replicas
If CPU < 70% target → scale down replicas
kubectl autoscale deployment nginx --cpu-percent=70 --min=2 --max=10
5. Namespace & Multi-tenancy
| Resource | Used for |
|---|---|
| Namespace | Logical division (dev, staging, prod or per team) |
| ResourceQuota | Limit total resources a namespace can consume |
| LimitRange | Set default/min/max resources per container in a namespace |
| NetworkPolicy | Control traffic between namespaces |
| RBAC | Control who can access what in a namespace |
Multi-tenancy Pattern:
Cluster
├── namespace: team-alpha
│ ├── ResourceQuota: max 10 CPU, 32Gi memory
│ ├── LimitRange: default 200m CPU, 256Mi per container
│ ├── NetworkPolicy: deny all ingress from other namespaces
│ └── RBAC: only team-alpha members can access
├── namespace: team-beta
│ └── ...similar isolation...
└── namespace: kube-system (platform)
6. Cheat Sheet
| Exam question | Answer |
|---|---|
| What happens when CPU limit is exceeded? | Container is throttled (not killed) |
| What happens when memory limit is exceeded? | Container is OOMKilled |
| Scale based on CPU usage? | HPA |
| Add more nodes when Pods are unschedulable? | Cluster Autoscaler |
| Force Pods onto specific nodes? | nodeSelector or Node Affinity |
| Prevent Pods from running on certain nodes? | Taints |
| Limit total resources per namespace? | ResourceQuota |
7. Practice Questions
Q1: A container has a memory limit of 256Mi and attempts to allocate 300Mi. What happens?
- A) The container is CPU-throttled
- B) The container is terminated with OOMKilled ✓
- C) The memory limit is automatically increased
- D) The request is queued until memory becomes available
Explanation: When a container exceeds its memory limit, the kernel's OOM Killer terminates it. CPU exceeding limits causes throttling (degraded performance), but memory violations cause immediate termination.
Q2: A cluster admin wants to ensure only GPU workloads run on GPU-equipped nodes. What should they use?
- A) Node Affinity
- B) Taints on GPU nodes + Tolerations on GPU Pods ✓
- C) ResourceQuota
- D) PriorityClass
Explanation: Tainting GPU nodes (e.g., gpu=true:NoSchedule) prevents non-GPU Pods from being scheduled there. GPU workloads add the matching toleration, allowing them to use the GPU nodes. nodeSelector/affinity would attract Pods to GPU nodes but wouldn't prevent non-GPU Pods from landing there.
Q3: Which autoscaling mechanism should be used when Pods are in Pending state due to insufficient cluster resources?
- A) HPA
- B) VPA
- C) Cluster Autoscaler ✓
- D) LimitRange
Explanation: Cluster Autoscaler watches for unschedulable Pods (Pending due to insufficient CPU/memory on all nodes) and provisions new nodes from the cloud provider. HPA adds Pod replicas; VPA adjusts Pod resources — neither adds new nodes.