Chuyển đến nội dung chính

BÀI 41: HORIZONTAL & VERTICAL POD AUTOSCALING

HPA với CPU/memory và custom metrics, VPA recommendations, KEDA event-driven autoscaling, Cluster Autoscaler (on-prem alternatives), và scaling best practices.

🔒 DevSecOps — Bài 41 BÀI 41: HORIZONTAL & VERTICAL POD AUTOSCALING

Deploy Microservices On-Premises với Kubernetes HA

Phần 10: Deployment Patterns & Auto-Scaling

xdev.asia

🎯 MỤC TIÊU BÀI HỌC

  • ✅ HPA v2 với CPU, memory, custom metrics
  • ✅ VPA (Vertical Pod Autoscaler) recommendations
  • ✅ KEDA event-driven autoscaling
  • ✅ On-premises capacity planning (no cloud autoscaler)
  • ✅ Scaling best practices và anti-patterns

PHẦN 1: HPA V2 (HORIZONTAL POD AUTOSCALER)


HPA Flow:

Metrics Server / Prometheus
        │
        ▼
┌──────────────┐    scale up/down    ┌──────────────┐
│     HPA      │───────────────────►│  Deployment  │
│              │                     │  replicas:   │
│ target: 70%  │                     │  2 → 5 → 3   │
│ min: 2       │                     │              │
│ max: 20      │                     │              │
└──────────────┘                     └──────────────┘
# HPA with CPU + custom metrics:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: order-service-hpa
  namespace: default
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: order-service
  minReplicas: 2
  maxReplicas: 20
  metrics:
    # CPU-based scaling:
    - type: Resource
      resource:
        name: cpu
        target:
          type: Utilization
          averageUtilization: 70

    # Memory-based scaling:
    - type: Resource
      resource:
        name: memory
        target:
          type: Utilization
          averageUtilization: 80

    # Custom metric (requests per second):
    - type: Pods
      pods:
        metric:
          name: http_requests_per_second
        target:
          type: AverageValue
          averageValue: "100"

  behavior:
    scaleUp:
      stabilizationWindowSeconds: 60
      policies:
        - type: Percent
          value: 50
          periodSeconds: 60
        - type: Pods
          value: 4
          periodSeconds: 60
      selectPolicy: Max
    scaleDown:
      stabilizationWindowSeconds: 300
      policies:
        - type: Percent
          value: 10
          periodSeconds: 120
      selectPolicy: Min
# Install Metrics Server (for CPU/memory):
helm install metrics-server metrics-server/metrics-server \
  --namespace kube-system \
  --set args[0]=--kubelet-insecure-tls

# Install Prometheus Adapter (for custom metrics):
helm install prometheus-adapter prometheus-community/prometheus-adapter \
  --namespace monitoring \
  -f adapter-values.yaml
# adapter-values.yaml:
prometheus:
  url: http://prometheus-kube-prometheus-prometheus.monitoring
  port: 9090

rules:
  default: false
  custom:
    - seriesQuery: 'http_server_request_duration_seconds_count{namespace!="",pod!=""}'
      resources:
        overrides:
          namespace: { resource: "namespace" }
          pod: { resource: "pod" }
      name:
        matches: "^(.*)_total$"
        as: "http_requests_per_second"
      metricsQuery: 'sum(rate(<<.Series>>{<<.LabelMatchers>>}[2m])) by (<<.GroupBy>>)'

PHẦN 2: VERTICAL POD AUTOSCALER

# Install VPA:
git clone https://github.com/kubernetes/autoscaler.git
cd autoscaler/vertical-pod-autoscaler
./hack/vpa-up.sh
# VPA recommendation mode (safe for production):
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
  name: order-service-vpa
spec:
  targetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: order-service
  updatePolicy:
    updateMode: "Off"  # Recommendation only, no auto-update
  resourcePolicy:
    containerPolicies:
      - containerName: app
        minAllowed:
          cpu: 50m
          memory: 64Mi
        maxAllowed:
          cpu: 2
          memory: 2Gi
        controlledResources: ["cpu", "memory"]
# Get VPA recommendations:
kubectl get vpa order-service-vpa -o yaml

# Output:
# recommendation:
#   containerRecommendations:
#     - containerName: app
#       lowerBound:   {cpu: 100m, memory: 128Mi}
#       target:       {cpu: 250m, memory: 256Mi}
#       upperBound:   {cpu: 500m, memory: 512Mi}
#       uncappedTarget: {cpu: 250m, memory: 256Mi}

PHẦN 3: KEDA — EVENT-DRIVEN AUTOSCALING

# Install KEDA:
helm repo add kedacore https://kedacore.github.io/charts
helm install keda kedacore/keda \
  --namespace keda \
  --create-namespace
# Scale based on RabbitMQ queue length:
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: order-worker
  namespace: default
spec:
  scaleTargetRef:
    name: order-worker
  pollingInterval: 15
  cooldownPeriod: 300
  minReplicaCount: 1
  maxReplicaCount: 30
  triggers:
    - type: rabbitmq
      metadata:
        host: amqp://user:[email protected]:5672
        queueName: order-processing
        queueLength: "10"  # 1 pod per 10 messages
        protocol: amqp

---
# Scale based on Kafka consumer lag:
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: event-processor
spec:
  scaleTargetRef:
    name: event-processor
  triggers:
    - type: kafka
      metadata:
        bootstrapServers: kafka-bootstrap.default:9092
        consumerGroup: event-processor-group
        topic: events
        lagThreshold: "100"

---
# Scale based on Prometheus metric:
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: api-gateway
spec:
  scaleTargetRef:
    name: api-gateway
  triggers:
    - type: prometheus
      metadata:
        serverAddress: http://prometheus-kube-prometheus-prometheus.monitoring:9090
        metricName: http_requests_per_second
        query: sum(rate(http_server_request_duration_seconds_count{service="api-gateway"}[2m]))
        threshold: "500"

PHẦN 4: ON-PREMISES CAPACITY PLANNING


On-Premises vs Cloud Scaling:

Cloud: Auto-scale nodes (add VMs)
On-Prem: Fixed hardware → must plan ahead

Capacity Zones:
┌─────────────────────────────────────────┐
│ Physical Capacity: 100 CPU, 400GB RAM   │
│ ████████████████████████████████████████│
│                                         │
│ Allocated:     70 CPU, 280GB RAM (70%)  │
│ ████████████████████████████░░░░░░░░░░░│
│                                         │
│ Reserved:      15 CPU, 60GB RAM (15%)   │
│ ░░░░░░░░░░░░░░░░████░░░░░░░░░░░░░░░░░│
│                                         │
│ Buffer:        15 CPU, 60GB RAM (15%)   │
│ ░░░░░░░░░░░░░░░░░░░░░░░░░░████░░░░░░░│
│                     Alert at 80% ▲      │
└─────────────────────────────────────────┘
# Alert when cluster capacity is running low:
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: capacity-alerts
spec:
  groups:
    - name: capacity
      rules:
        - alert: ClusterCPUCapacityLow
          expr: |
            sum(kube_pod_container_resource_requests{resource="cpu"})
            /
            sum(kube_node_status_allocatable{resource="cpu"})
            > 0.8
          for: 15m
          labels:
            severity: warning
          annotations:
            summary: "Cluster CPU allocation > 80%"

        - alert: ClusterMemoryCapacityLow
          expr: |
            sum(kube_pod_container_resource_requests{resource="memory"})
            /
            sum(kube_node_status_allocatable{resource="memory"})
            > 0.8
          for: 15m
          labels:
            severity: warning

💡 KEY TAKEAWAYS

  1. HPA v2: Scale on CPU, memory, or custom Prometheus metrics
  2. Behavior: Configure scale-up/down speed and stabilization
  3. VPA: Use "Off" mode for recommendations, avoid with HPA on same metric
  4. KEDA: Event-driven scaling (queue length, Kafka lag)
  5. On-prem: Fixed capacity → plan buffer, alert at 80%
  6. Anti-pattern: Don't use HPA + VPA on same metric

🎯 BÀI TẬP

Bài tập 1: HPA + Custom Metrics

  • Deploy Prometheus Adapter
  • Create HPA with custom requests/sec metric
  • Load test and observe scaling behavior

Bài tập 2: KEDA Queue-Based Scaling

  • Install KEDA, create ScaledObject for RabbitMQ
  • Push 1000 messages to queue
  • Verify worker pods scale up, then scale down

📚 BÀI TIẾP THEO

Trong Bài 42: Resource Management & Scheduling, chúng ta sẽ tối ưu resource allocation và pod scheduling.