Chuyển đến nội dung chính

Lesson 9: Chaos Engineering — Combining Load Testing and Fault Injection

Steady-state hypothesis, Litmus Chaos, Chaos Mesh, Gremlin. Game days and resilience scoring.

🔒 DevSecOps — Lesson 9 Lesson 9: Chaos Engineering — Combining Load Testing and Fault Injection__HTMLTAG_55___

Performance Testing & Pentest: Enterprise Standard Process 2026

Part 2: Advanced Performance Testing

xdev.asia

1. What is Chaos Engineering?

Chaos Engineering = "Break things on purpose to learn"

Quy trình:
  1. Định nghĩa STEADY STATE (hệ thống hoạt động bình thường)
     → p95 latency < 300ms, error rate < 0.1%, 99.9% availability

  2. Đặt HYPOTHESIS (giả thuyết)
     → "Nếu 1 pod bị kill, hệ thống vẫn đáp ứng SLO"

  3. Inject FAULT (gây lỗi có kiểm soát)
     → kubectl delete pod api-server-abc123

  4. OBSERVE impact
     → Latency tăng 20% trong 30s, sau đó recovery

  5. LEARN & improve
     → Thêm PodDisruptionBudget, tune readiness probe

2. Litmus Chaos on Kubernetes

# Cài đặt LitmusChaos
helm repo add litmuschaos https://litmuschaos.github.io/litmus-helm
helm install litmus litmuschaos/litmus \
  --namespace litmus --create-namespace
# Pod delete experiment
apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
  name: api-pod-delete
spec:
  appinfo:
    appns: production
    applabel: "app=api-server"
    appkind: deployment
  engineState: active
  chaosServiceAccount: litmus-admin
  experiments:
    - name: pod-delete
      spec:
        components:
          env:
            - name: TOTAL_CHAOS_DURATION
              value: "60"           # 60 giây
            - name: CHAOS_INTERVAL
              value: "10"           # Kill pod mỗi 10s
            - name: FORCE
              value: "true"
        probe:
          - name: "api-health-check"
            type: httpProbe
            httpProbe/inputs:
              url: "http://api-server.production:3000/health"
              method:
                get:
                  criteria: ==
                  responseCode: "200"
            mode: Continuous
            runProperties:
              probeTimeout: 5s
              interval: 2s
              retry: 3

3. Chaos Mesh — Network Chaos

# Network delay injection
apiVersion: chaos-mesh.org/v1alpha1
kind: NetworkChaos
metadata:
  name: network-delay
spec:
  action: delay
  mode: all
  selector:
    namespaces:
      - production
    labelSelectors:
      app: api-server
  delay:
    latency: "200ms"
    jitter: "50ms"
    correlation: "50"
  duration: "5m"
  
---
# Network partition between services
apiVersion: chaos-mesh.org/v1alpha1
kind: NetworkChaos
metadata:
  name: partition-db
spec:
  action: partition
  mode: all
  selector:
    namespaces:
      - production
    labelSelectors:
      app: api-server
  direction: both
  target:
    selector:
      namespaces:
        - production
      labelSelectors:
        app: postgresql
  duration: "2m"
# CPU stress
apiVersion: chaos-mesh.org/v1alpha1
kind: StressChaos
metadata:
  name: cpu-stress
spec:
  mode: one
  selector:
    labelSelectors:
      app: api-server
  stressors:
    cpu:
      workers: 2
      load: 80            # 80% CPU
  duration: "3m"

4. Game Day — Load Test + Chaos Combined

Game Day Runbook:

Phase 1: Baseline (15 min)
  └── k6 load test: 500 RPS, tất cả metrics normal

Phase 2: Pod failure (15 min)
  ├── Tiếp tục k6 load test
  └── Chaos: Kill 1/3 pods mỗi 30 giây
  → Expect: Latency tăng nhẹ, no 5xx errors

Phase 3: Network degradation (15 min)
  ├── Tiếp tục k6 load test
  └── Chaos: 200ms latency injection giữa API → DB
  → Expect: p99 tăng ~200ms, no timeouts

Phase 4: Database failover (15 min)
  ├── Tiếp tục k6 load test
  └── Chaos: Kill primary DB, promote replica
  → Expect: < 30s downtime, auto-recovery

Phase 5: Combined stress (15 min)
  ├── k6: Spike to 2000 RPS
  ├── Chaos: 50% CPU stress + network delay
  └── → Measure degradation gracefully
// k6 script chạy song song với Chaos experiments
export const options = {
  scenarios: {
    game_day: {
      executor: 'ramping-arrival-rate',
      startRate: 100,
      timeUnit: '1s',
      stages: [
        { duration: '15m', target: 500 },   // Phase 1: Baseline
        { duration: '15m', target: 500 },   // Phase 2: Pod failure
        { duration: '15m', target: 500 },   // Phase 3: Network
        { duration: '15m', target: 500 },   // Phase 4: DB failover
        { duration: '15m', target: 2000 },  // Phase 5: Spike + Chaos
      ],
      preAllocatedVUs: 500,
      maxVUs: 2000,
    },
  },
  thresholds: {
    http_req_duration: ['p(95)<500'],  // Relaxed thresholds for game day
    http_req_failed: ['rate<0.05'],    // Accept up to 5% failure
  },
};

5. Resilience Scoring

Resilience Score = Weighted average of experiment results

┌──────────────────────┬────────┬────────┬─────────┐
│ Experiment           │ Weight │ Result │ Score   │
├──────────────────────┼────────┼────────┼─────────┤
│ Pod failure          │ 25%    │ Pass   │ 25/25   │
│ Network latency      │ 20%    │ Pass   │ 20/20   │
│ Network partition    │ 15%    │ Fail   │ 0/15    │
│ CPU stress           │ 15%    │ Pass   │ 15/15   │
│ DB failover          │ 15%    │ Partial│ 10/15   │
│ Spike + Chaos combo  │ 10%    │ Pass   │ 10/10   │
├──────────────────────┼────────┼────────┼─────────┤
│ Total                │ 100%   │        │ 80/100  │
└──────────────────────┴────────┴────────┴─────────┘

Score: 80/100 = "Good" (target: > 85 = "Excellent")
Action items:
  - Fix network partition handling (add circuit breaker)
  - Improve DB failover time (< 10s target)

6. Summary

  • Chaos Engineering: Proactively find weaknesses before users do
  • Steady-state hypothesis: Define "normal" before breaking things
  • Tools: Litmus Chaos, Chaos Mesh, Gremlin, AWS FIS
  • Game Days: Combine load testing + chaos for realistic scenarios
  • Resilience Score: Quantify system robustness over time

The next article will explore Observability-driven Testing and AI-assisted Performance.