🎯 LESSON OBJECTIVE__HTMLTAG_66___
- ✅ Chaos Engineering principles (Netflix model)
- ✅ Deploy Litmus Chaos on K8s
- ✅ Pod chaos: kill, CPU stress, memory stress
- ✅ Node chaos: drain, network partition__HTMLTAG_75___
- ✅ Steady-state hypothesis and probes
- ✅ GameDay planning and resilience scoring
PART 1: CHAOS ENGINEERING PRINCIPLES
Chaos Engineering Process:
1. Define Steady State
"Order service handles 1000 req/s with P99 < 200ms"
2. Hypothesize
"If we kill 1 pod, the system auto-recovers in < 30s"
3. Inject Failure
Kill pod / network partition / CPU stress
4. Observe
Metrics, logs, user impact
5. Learn & Improve
Fix weaknesses, update runbooks
┌─────────────┐ ┌──────────────┐ ┌───────────┐
│Define Steady │────►│ Inject │────►│ Observe │
│ State │ │ Chaos │ │ & Learn │
└─────────────┘ └──────────────┘ └─────┬─────┘
▲ │
└────────────────────────────────────────┘
(improve & repeat)
PART 2: DEPLOY LITMUS CHAOS__HTMLTAG_86___
# Install Litmus:
helm repo add litmuschaos https://litmuschaos.github.io/litmus-helm/
helm install litmus litmuschaos/litmus \
--namespace litmus \
--create-namespace \
--set portal.frontend.service.type=ClusterIP
# ChaosEngine: pod-kill experiment
apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
name: order-service-chaos
namespace: default
spec:
engineState: active
appinfo:
appns: default
applabel: app=order-service
appkind: deployment
chaosServiceAccount: litmus-admin
experiments:
- name: pod-delete
spec:
components:
env:
- name: TOTAL_CHAOS_DURATION
value: "60"
- name: CHAOS_INTERVAL
value: "10"
- name: FORCE
value: "true"
- name: PODS_AFFECTED_PERC
value: "50"
probe:
- name: check-availability
type: httpProbe
mode: Continuous
httpProbe/inputs:
url: http://order-service:8080/health
method:
get:
criteria: ==
responseCode: "200"
runProperties:
probeTimeout: 5s
interval: 5s
retry: 3
probePollingInterval: 2s
# Install Litmus:
helm repo add litmuschaos https://litmuschaos.github.io/litmus-helm/
helm install litmus litmuschaos/litmus \
--namespace litmus \
--create-namespace \
--set portal.frontend.service.type=ClusterIP
# ChaosEngine: pod-kill experiment
apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
name: order-service-chaos
namespace: default
spec:
engineState: active
appinfo:
appns: default
applabel: app=order-service
appkind: deployment
chaosServiceAccount: litmus-admin
experiments:
- name: pod-delete
spec:
components:
env:
- name: TOTAL_CHAOS_DURATION
value: "60"
- name: CHAOS_INTERVAL
value: "10"
- name: FORCE
value: "true"
- name: PODS_AFFECTED_PERC
value: "50"
probe:
- name: check-availability
type: httpProbe
mode: Continuous
httpProbe/inputs:
url: http://order-service:8080/health
method:
get:
criteria: ==
responseCode: "200"
runProperties:
probeTimeout: 5s
interval: 5s
retry: 3
probePollingInterval: 2s
PART 3: CHAOS EXPERIMENTS
# CPU stress:
apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
name: cpu-stress-test
spec:
engineState: active
appinfo:
appns: default
applabel: app=order-service
appkind: deployment
experiments:
- name: pod-cpu-hog
spec:
components:
env:
- name: CPU_CORES
value: "2"
- name: TOTAL_CHAOS_DURATION
value: "120"
- name: CPU_LOAD
value: "80"
---
# Memory stress:
apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
name: memory-stress-test
spec:
experiments:
- name: pod-memory-hog
spec:
components:
env:
- name: MEMORY_CONSUMPTION
value: "500"
- name: TOTAL_CHAOS_DURATION
value: "120"
---
# Network chaos (latency injection):
apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
name: network-chaos-test
spec:
experiments:
- name: pod-network-latency
spec:
components:
env:
- name: NETWORK_LATENCY
value: "200"
- name: TOTAL_CHAOS_DURATION
value: "120"
- name: DESTINATION_IPS
value: "10.96.0.0/12"
- name: NETWORK_INTERFACE
value: "eth0"
---
# Node drain:
apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
name: node-drain-test
spec:
experiments:
- name: node-drain
spec:
components:
env:
- name: TOTAL_CHAOS_DURATION
value: "120"
- name: TARGET_NODE
value: "worker-01"
PART 4: STEADY-STATE PROBES__HTMLTAG_92___
# HTTP probe: service must respond 200:
probe:
- name: http-health-check
type: httpProbe
mode: Continuous
httpProbe/inputs:
url: http://order-service:8080/health
method:
get:
criteria: ==
responseCode: "200"
runProperties:
probeTimeout: 5s
interval: 5s
retry: 3
# Prometheus probe: error rate must stay below SLO:
probe:
- name: error-rate-check
type: promProbe
mode: Edge
promProbe/inputs:
endpoint: http://prometheus.monitoring:9090
query: >
sum(rate(http_server_request_duration_seconds_count{service="order-service",http_status_code=~"5.."}[2m]))
/
sum(rate(http_server_request_duration_seconds_count{service="order-service"}[2m]))
comparator:
type: float
criteria: "<="
value: "0.01"
runProperties:
probeTimeout: 10s
interval: 30s
# Command probe: check database connectivity:
probe:
- name: db-connectivity
type: cmdProbe
mode: Edge
cmdProbe/inputs:
command: "pg_isready -h postgresql-rw.database -p 5432"
comparator:
type: string
criteria: contains
value: "accepting connections"
runProperties:
probeTimeout: 10s
# HTTP probe: service must respond 200:
probe:
- name: http-health-check
type: httpProbe
mode: Continuous
httpProbe/inputs:
url: http://order-service:8080/health
method:
get:
criteria: ==
responseCode: "200"
runProperties:
probeTimeout: 5s
interval: 5s
retry: 3
# Prometheus probe: error rate must stay below SLO:
probe:
- name: error-rate-check
type: promProbe
mode: Edge
promProbe/inputs:
endpoint: http://prometheus.monitoring:9090
query: >
sum(rate(http_server_request_duration_seconds_count{service="order-service",http_status_code=~"5.."}[2m]))
/
sum(rate(http_server_request_duration_seconds_count{service="order-service"}[2m]))
comparator:
type: float
criteria: "<="
value: "0.01"
runProperties:
probeTimeout: 10s
interval: 30s
# Command probe: check database connectivity:
probe:
- name: db-connectivity
type: cmdProbe
mode: Edge
cmdProbe/inputs:
command: "pg_isready -h postgresql-rw.database -p 5432"
comparator:
type: string
criteria: contains
value: "accepting connections"
runProperties:
probeTimeout: 10s
PART 5: GAMEDAY PLANNING
| Phase | Activity | Duration |
|---|---|---|
| Preparation | Define experiments, notify teams, ensure monitoring | 1 week ago_ |
| Briefing | Review experiments, assign observers, confirm rollback | 30 min |
| Execution_ | Run chaos experiments one-by-one | 2-4 hours_ |
| Observation | Monitor dashboards, notes anomalies | During execution |
| Debrief | Review findings, create action items_ | 1 hour |
| Follow-up | Implement fixes, schedule next GameDay_ | 2 weeks |
# GameDay experiment sequence:
# Round 1: Single pod kill
# Hypothesis: "Service recovers in < 30s, zero user errors"
# Round 2: Kill 50% pods
# Hypothesis: "Remaining pods handle load, P99 < 500ms"
# Round 3: Node failure
# Hypothesis: "Pods reschedule in < 2 min"
# Round 4: Network partition
# Hypothesis: "Circuit breaker activates, graceful degradation"
# Round 5: Database failover
# Hypothesis: "< 5s downtime, no data loss"
💡 KEY TAKEAWAYS
- Chaos Engineering: Proactively find weaknesses before production incidents
- Litmus: K8s-native chaos framework, CRD-based experiments
- Probes: Validate steady-state during chaos (HTTP, Prometheus, command)
- Start small: Pod kill → node drain → network chaos
- GameDay: Structured team exercise, not random destruction
- Always have rollback: Know how to stop chaos immediately
🎯 EXERCISES
Exercise 1: Litmus Setup__HTMLTAG_193___
- Install Litmus, run pod-delete experiment__HTMLTAG_196___
- Add HTTP probe to validate service availability
- Review ChaosResult for pass/fail
Exercise 2: GameDay
- Plan 5-round chaos experiment sequence
- Execute with monitoring dashboards open
- Document findings and improvement actions
📚 NEXT POST
In Lesson 46: Production Readiness Checklist, we will start Section 12 — Production Operations & Capstone Project.