1. What is Chaos Engineering?
Chaos Engineering = "Break things on purpose to learn"
Quy trình:
1. Định nghĩa STEADY STATE (hệ thống hoạt động bình thường)
→ p95 latency < 300ms, error rate < 0.1%, 99.9% availability
2. Đặt HYPOTHESIS (giả thuyết)
→ "Nếu 1 pod bị kill, hệ thống vẫn đáp ứng SLO"
3. Inject FAULT (gây lỗi có kiểm soát)
→ kubectl delete pod api-server-abc123
4. OBSERVE impact
→ Latency tăng 20% trong 30s, sau đó recovery
5. LEARN & improve
→ Thêm PodDisruptionBudget, tune readiness probe
2. Litmus Chaos on Kubernetes
# Cài đặt LitmusChaos
helm repo add litmuschaos https://litmuschaos.github.io/litmus-helm
helm install litmus litmuschaos/litmus \
--namespace litmus --create-namespace
# Pod delete experiment
apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
name: api-pod-delete
spec:
appinfo:
appns: production
applabel: "app=api-server"
appkind: deployment
engineState: active
chaosServiceAccount: litmus-admin
experiments:
- name: pod-delete
spec:
components:
env:
- name: TOTAL_CHAOS_DURATION
value: "60" # 60 giây
- name: CHAOS_INTERVAL
value: "10" # Kill pod mỗi 10s
- name: FORCE
value: "true"
probe:
- name: "api-health-check"
type: httpProbe
httpProbe/inputs:
url: "http://api-server.production:3000/health"
method:
get:
criteria: ==
responseCode: "200"
mode: Continuous
runProperties:
probeTimeout: 5s
interval: 2s
retry: 3
3. Chaos Mesh — Network Chaos
# Network delay injection
apiVersion: chaos-mesh.org/v1alpha1
kind: NetworkChaos
metadata:
name: network-delay
spec:
action: delay
mode: all
selector:
namespaces:
- production
labelSelectors:
app: api-server
delay:
latency: "200ms"
jitter: "50ms"
correlation: "50"
duration: "5m"
---
# Network partition between services
apiVersion: chaos-mesh.org/v1alpha1
kind: NetworkChaos
metadata:
name: partition-db
spec:
action: partition
mode: all
selector:
namespaces:
- production
labelSelectors:
app: api-server
direction: both
target:
selector:
namespaces:
- production
labelSelectors:
app: postgresql
duration: "2m"
# CPU stress
apiVersion: chaos-mesh.org/v1alpha1
kind: StressChaos
metadata:
name: cpu-stress
spec:
mode: one
selector:
labelSelectors:
app: api-server
stressors:
cpu:
workers: 2
load: 80 # 80% CPU
duration: "3m"
4. Game Day — Load Test + Chaos Combined
Game Day Runbook:
Phase 1: Baseline (15 min)
└── k6 load test: 500 RPS, tất cả metrics normal
Phase 2: Pod failure (15 min)
├── Tiếp tục k6 load test
└── Chaos: Kill 1/3 pods mỗi 30 giây
→ Expect: Latency tăng nhẹ, no 5xx errors
Phase 3: Network degradation (15 min)
├── Tiếp tục k6 load test
└── Chaos: 200ms latency injection giữa API → DB
→ Expect: p99 tăng ~200ms, no timeouts
Phase 4: Database failover (15 min)
├── Tiếp tục k6 load test
└── Chaos: Kill primary DB, promote replica
→ Expect: < 30s downtime, auto-recovery
Phase 5: Combined stress (15 min)
├── k6: Spike to 2000 RPS
├── Chaos: 50% CPU stress + network delay
└── → Measure degradation gracefully
// k6 script chạy song song với Chaos experiments
export const options = {
scenarios: {
game_day: {
executor: 'ramping-arrival-rate',
startRate: 100,
timeUnit: '1s',
stages: [
{ duration: '15m', target: 500 }, // Phase 1: Baseline
{ duration: '15m', target: 500 }, // Phase 2: Pod failure
{ duration: '15m', target: 500 }, // Phase 3: Network
{ duration: '15m', target: 500 }, // Phase 4: DB failover
{ duration: '15m', target: 2000 }, // Phase 5: Spike + Chaos
],
preAllocatedVUs: 500,
maxVUs: 2000,
},
},
thresholds: {
http_req_duration: ['p(95)<500'], // Relaxed thresholds for game day
http_req_failed: ['rate<0.05'], // Accept up to 5% failure
},
};
5. Resilience Scoring
Resilience Score = Weighted average of experiment results
┌──────────────────────┬────────┬────────┬─────────┐
│ Experiment │ Weight │ Result │ Score │
├──────────────────────┼────────┼────────┼─────────┤
│ Pod failure │ 25% │ Pass │ 25/25 │
│ Network latency │ 20% │ Pass │ 20/20 │
│ Network partition │ 15% │ Fail │ 0/15 │
│ CPU stress │ 15% │ Pass │ 15/15 │
│ DB failover │ 15% │ Partial│ 10/15 │
│ Spike + Chaos combo │ 10% │ Pass │ 10/10 │
├──────────────────────┼────────┼────────┼─────────┤
│ Total │ 100% │ │ 80/100 │
└──────────────────────┴────────┴────────┴─────────┘
Score: 80/100 = "Good" (target: > 85 = "Excellent")
Action items:
- Fix network partition handling (add circuit breaker)
- Improve DB failover time (< 10s target)
6. Summary
- Chaos Engineering: Proactively find weaknesses before users do
- Steady-state hypothesis: Define "normal" before breaking things
- Tools: Litmus Chaos, Chaos Mesh, Gremlin, AWS FIS
- Game Days: Combine load testing + chaos for realistic scenarios
- Resilience Score: Quantify system robustness over time
The next article will explore Observability-driven Testing and AI-assisted Performance.