Chuyển đến nội dung chính

BÀI 45: CHAOS ENGINEERING VỚI LITMUS

Chaos Engineering principles, Litmus Chaos trên K8s, pod/node/network chaos experiments, steady-state hypothesis, GameDay planning, và resilience scoring.

🔒 DevSecOps — Bài 45 BÀI 45: CHAOS ENGINEERING VỚI LITMUS

Deploy Microservices On-Premises với Kubernetes HA

Phần 11: Disaster Recovery & Chaos Engineering

xdev.asia

🎯 MỤC TIÊU BÀI HỌC

  • ✅ Chaos Engineering principles (Netflix model)
  • ✅ Deploy Litmus Chaos trên K8s
  • ✅ Pod chaos: kill, CPU stress, memory stress
  • ✅ Node chaos: drain, network partition
  • ✅ Steady-state hypothesis và probes
  • ✅ GameDay planning và resilience scoring

PHẦN 1: CHAOS ENGINEERING PRINCIPLES


Chaos Engineering Process:

1. Define Steady State
   "Order service handles 1000 req/s with P99 < 200ms"

2. Hypothesize
   "If we kill 1 pod, the system auto-recovers in < 30s"

3. Inject Failure
   Kill pod / network partition / CPU stress

4. Observe
   Metrics, logs, user impact

5. Learn & Improve
   Fix weaknesses, update runbooks

┌─────────────┐     ┌──────────────┐     ┌───────────┐
│Define Steady │────►│   Inject     │────►│  Observe  │
│   State      │     │   Chaos      │     │  & Learn  │
└─────────────┘     └──────────────┘     └─────┬─────┘
       ▲                                        │
       └────────────────────────────────────────┘
                    (improve & repeat)

PHẦN 2: DEPLOY LITMUS CHAOS

# Install Litmus:
helm repo add litmuschaos https://litmuschaos.github.io/litmus-helm/
helm install litmus litmuschaos/litmus \
  --namespace litmus \
  --create-namespace \
  --set portal.frontend.service.type=ClusterIP
# ChaosEngine: pod-kill experiment
apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
  name: order-service-chaos
  namespace: default
spec:
  engineState: active
  appinfo:
    appns: default
    applabel: app=order-service
    appkind: deployment
  chaosServiceAccount: litmus-admin
  experiments:
    - name: pod-delete
      spec:
        components:
          env:
            - name: TOTAL_CHAOS_DURATION
              value: "60"
            - name: CHAOS_INTERVAL
              value: "10"
            - name: FORCE
              value: "true"
            - name: PODS_AFFECTED_PERC
              value: "50"
        probe:
          - name: check-availability
            type: httpProbe
            mode: Continuous
            httpProbe/inputs:
              url: http://order-service:8080/health
              method:
                get:
                  criteria: ==
                  responseCode: "200"
            runProperties:
              probeTimeout: 5s
              interval: 5s
              retry: 3
              probePollingInterval: 2s

PHẦN 3: CHAOS EXPERIMENTS

# CPU stress:
apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
  name: cpu-stress-test
spec:
  engineState: active
  appinfo:
    appns: default
    applabel: app=order-service
    appkind: deployment
  experiments:
    - name: pod-cpu-hog
      spec:
        components:
          env:
            - name: CPU_CORES
              value: "2"
            - name: TOTAL_CHAOS_DURATION
              value: "120"
            - name: CPU_LOAD
              value: "80"

---
# Memory stress:
apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
  name: memory-stress-test
spec:
  experiments:
    - name: pod-memory-hog
      spec:
        components:
          env:
            - name: MEMORY_CONSUMPTION
              value: "500"
            - name: TOTAL_CHAOS_DURATION
              value: "120"

---
# Network chaos (latency injection):
apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
  name: network-chaos-test
spec:
  experiments:
    - name: pod-network-latency
      spec:
        components:
          env:
            - name: NETWORK_LATENCY
              value: "200"
            - name: TOTAL_CHAOS_DURATION
              value: "120"
            - name: DESTINATION_IPS
              value: "10.96.0.0/12"
            - name: NETWORK_INTERFACE
              value: "eth0"

---
# Node drain:
apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
  name: node-drain-test
spec:
  experiments:
    - name: node-drain
      spec:
        components:
          env:
            - name: TOTAL_CHAOS_DURATION
              value: "120"
            - name: TARGET_NODE
              value: "worker-01"

PHẦN 4: STEADY-STATE PROBES

# HTTP probe: service must respond 200:
probe:
  - name: http-health-check
    type: httpProbe
    mode: Continuous
    httpProbe/inputs:
      url: http://order-service:8080/health
      method:
        get:
          criteria: ==
          responseCode: "200"
    runProperties:
      probeTimeout: 5s
      interval: 5s
      retry: 3

# Prometheus probe: error rate must stay below SLO:
probe:
  - name: error-rate-check
    type: promProbe
    mode: Edge
    promProbe/inputs:
      endpoint: http://prometheus.monitoring:9090
      query: >
        sum(rate(http_server_request_duration_seconds_count{service="order-service",http_status_code=~"5.."}[2m]))
        /
        sum(rate(http_server_request_duration_seconds_count{service="order-service"}[2m]))
      comparator:
        type: float
        criteria: "<="
        value: "0.01"
    runProperties:
      probeTimeout: 10s
      interval: 30s

# Command probe: check database connectivity:
probe:
  - name: db-connectivity
    type: cmdProbe
    mode: Edge
    cmdProbe/inputs:
      command: "pg_isready -h postgresql-rw.database -p 5432"
      comparator:
        type: string
        criteria: contains
        value: "accepting connections"
    runProperties:
      probeTimeout: 10s

PHẦN 5: GAMEDAY PLANNING

PhaseActivityDuration
PreparationDefine experiments, notify teams, ensure monitoring1 week before
BriefingReview experiments, assign observers, confirm rollback30 min
ExecutionRun chaos experiments one-by-one2-4 hours
ObservationMonitor dashboards, note anomaliesDuring execution
DebriefReview findings, create action items1 hour
Follow-upImplement fixes, schedule next GameDay2 weeks
# GameDay experiment sequence:

# Round 1: Single pod kill
# Hypothesis: "Service recovers in < 30s, zero user errors"

# Round 2: Kill 50% pods
# Hypothesis: "Remaining pods handle load, P99 < 500ms"

# Round 3: Node failure
# Hypothesis: "Pods reschedule in < 2 min"

# Round 4: Network partition
# Hypothesis: "Circuit breaker activates, graceful degradation"

# Round 5: Database failover
# Hypothesis: "< 5s downtime, no data loss"

💡 KEY TAKEAWAYS

  1. Chaos Engineering: Proactively find weaknesses before production incidents
  2. Litmus: K8s-native chaos framework, CRD-based experiments
  3. Probes: Validate steady-state during chaos (HTTP, Prometheus, command)
  4. Start small: Pod kill → node drain → network chaos
  5. GameDay: Structured team exercise, not random destruction
  6. Always have rollback: Know how to stop chaos immediately

🎯 BÀI TẬP

Bài tập 1: Litmus Setup

  • Install Litmus, run pod-delete experiment
  • Add HTTP probe to validate service availability
  • Review ChaosResult for pass/fail

Bài tập 2: GameDay

  • Plan 5-round chaos experiment sequence
  • Execute with monitoring dashboards open
  • Document findings and improvement actions

📚 BÀI TIẾP THEO

Trong Bài 46: Production Readiness Checklist, chúng ta sẽ bắt đầu Section 12 — Production Operations & Capstone Project.