🎯 LESSON OBJECTIVE__HTMLTAG_68___
- ✅ HPA v2 with CPU, memory, custom metrics__HTMLTAG_71___
- ✅ VPA (Vertical Pod Autoscaler) recommendations
- ✅ KEDA event-driven autoscaling
- ✅ On-premises capacity planning (no cloud autoscaler)
- ✅ Scaling best practices and anti-patterns
PART 1: HPA V2 (HORIZONTAL POD AUTOSCALER)
HPA Flow:
Metrics Server / Prometheus
│
▼
┌──────────────┐ scale up/down ┌──────────────┐
│ HPA │───────────────────►│ Deployment │
│ │ │ replicas: │
│ target: 70% │ │ 2 → 5 → 3 │
│ min: 2 │ │ │
│ max: 20 │ │ │
└──────────────┘ └──────────────┘
# HPA with CPU + custom metrics:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: order-service-hpa
namespace: default
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: order-service
minReplicas: 2
maxReplicas: 20
metrics:
# CPU-based scaling:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
# Memory-based scaling:
- type: Resource
resource:
name: memory
target:
type: Utilization
averageUtilization: 80
# Custom metric (requests per second):
- type: Pods
pods:
metric:
name: http_requests_per_second
target:
type: AverageValue
averageValue: "100"
behavior:
scaleUp:
stabilizationWindowSeconds: 60
policies:
- type: Percent
value: 50
periodSeconds: 60
- type: Pods
value: 4
periodSeconds: 60
selectPolicy: Max
scaleDown:
stabilizationWindowSeconds: 300
policies:
- type: Percent
value: 10
periodSeconds: 120
selectPolicy: Min
# Install Metrics Server (for CPU/memory):
helm install metrics-server metrics-server/metrics-server \
--namespace kube-system \
--set args[0]=--kubelet-insecure-tls
# Install Prometheus Adapter (for custom metrics):
helm install prometheus-adapter prometheus-community/prometheus-adapter \
--namespace monitoring \
-f adapter-values.yaml
# adapter-values.yaml:
prometheus:
url: http://prometheus-kube-prometheus-prometheus.monitoring
port: 9090
rules:
default: false
custom:
- seriesQuery: 'http_server_request_duration_seconds_count{namespace!="",pod!=""}'
resources:
overrides:
namespace: { resource: "namespace" }
pod: { resource: "pod" }
name:
matches: "^(.*)_total$"
as: "http_requests_per_second"
metricsQuery: 'sum(rate(<<.Series>>{<<.LabelMatchers>>}[2m])) by (<<.GroupBy>>)'
PART 2: VERTICAL POD AUTOSCALER
# Install VPA:
git clone https://github.com/kubernetes/autoscaler.git
cd autoscaler/vertical-pod-autoscaler
./hack/vpa-up.sh
# VPA recommendation mode (safe for production):
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: order-service-vpa
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: order-service
updatePolicy:
updateMode: "Off" # Recommendation only, no auto-update
resourcePolicy:
containerPolicies:
- containerName: app
minAllowed:
cpu: 50m
memory: 64Mi
maxAllowed:
cpu: 2
memory: 2Gi
controlledResources: ["cpu", "memory"]
# Get VPA recommendations:
kubectl get vpa order-service-vpa -o yaml
# Output:
# recommendation:
# containerRecommendations:
# - containerName: app
# lowerBound: {cpu: 100m, memory: 128Mi}
# target: {cpu: 250m, memory: 256Mi}
# upperBound: {cpu: 500m, memory: 512Mi}
# uncappedTarget: {cpu: 250m, memory: 256Mi}
PART 3: KEDA — EVENT-DRIVEN AUTOSCALING
# Install KEDA:
helm repo add kedacore https://kedacore.github.io/charts
helm install keda kedacore/keda \
--namespace keda \
--create-namespace
# Scale based on RabbitMQ queue length:
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
name: order-worker
namespace: default
spec:
scaleTargetRef:
name: order-worker
pollingInterval: 15
cooldownPeriod: 300
minReplicaCount: 1
maxReplicaCount: 30
triggers:
- type: rabbitmq
metadata:
host: amqp://user:[email protected]:5672
queueName: order-processing
queueLength: "10" # 1 pod per 10 messages
protocol: amqp
---
# Scale based on Kafka consumer lag:
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
name: event-processor
spec:
scaleTargetRef:
name: event-processor
triggers:
- type: kafka
metadata:
bootstrapServers: kafka-bootstrap.default:9092
consumerGroup: event-processor-group
topic: events
lagThreshold: "100"
---
# Scale based on Prometheus metric:
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
name: api-gateway
spec:
scaleTargetRef:
name: api-gateway
triggers:
- type: prometheus
metadata:
serverAddress: http://prometheus-kube-prometheus-prometheus.monitoring:9090
metricName: http_requests_per_second
query: sum(rate(http_server_request_duration_seconds_count{service="api-gateway"}[2m]))
threshold: "500"
PART 4: ON-PREMISES CAPACITY PLANNING
On-Premises vs Cloud Scaling:
Cloud: Auto-scale nodes (add VMs)
On-Prem: Fixed hardware → must plan ahead
Capacity Zones:
┌─────────────────────────────────────────┐
│ Physical Capacity: 100 CPU, 400GB RAM │
│ ████████████████████████████████████████│
│ │
│ Allocated: 70 CPU, 280GB RAM (70%) │
│ ████████████████████████████░░░░░░░░░░░│
│ │
│ Reserved: 15 CPU, 60GB RAM (15%) │
│ ░░░░░░░░░░░░░░░░████░░░░░░░░░░░░░░░░░│
│ │
│ Buffer: 15 CPU, 60GB RAM (15%) │
│ ░░░░░░░░░░░░░░░░░░░░░░░░░░████░░░░░░░│
│ Alert at 80% ▲ │
└─────────────────────────────────────────┘
# Alert when cluster capacity is running low:
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: capacity-alerts
spec:
groups:
- name: capacity
rules:
- alert: ClusterCPUCapacityLow
expr: |
sum(kube_pod_container_resource_requests{resource="cpu"})
/
sum(kube_node_status_allocatable{resource="cpu"})
> 0.8
for: 15m
labels:
severity: warning
annotations:
summary: "Cluster CPU allocation > 80%"
- alert: ClusterMemoryCapacityLow
expr: |
sum(kube_pod_container_resource_requests{resource="memory"})
/
sum(kube_node_status_allocatable{resource="memory"})
> 0.8
for: 15m
labels:
severity: warning
💡 KEY TAKEAWAYS
- HPA v2: Scale on CPU, memory, or custom Prometheus metrics
- Behavior: Configure scale-up/down speed and stabilization
- VPA: Use "Off" mode for recommendations, avoid with HPA on same metric
- KEDA: Event-driven scaling (queue length, Kafka lag)
- On-prem: Fixed capacity → plan buffer, alert at 80%
- Anti-pattern: Don't use HPA + VPA on same metric
🎯 EXERCISE
Exercise 1: HPA + Custom Metrics__HTMLTAG_126___
- Deploy Prometheus Adapter
- Create HPA with custom requests/sec metric
- Load test and observe scaling behavior__HTMLTAG_133___
Exercise 2: KEDA Queue-Based Scaling
- Install KEDA, create ScaledObject for RabbitMQ
- Push 1000 messages to queue
- Verify worker pods scale up, then scale down
📚 NEXT POST
In Lesson 42: Resource Management & Scheduling, we will optimize resource allocation and pod scheduling.