
簡介
您已经实施了断路器、重试、隔板……但是当生产出现问题时,您如何证明它们确实有效?
混沌工程是一門透過「主動」在生產或登台環境中引發問題來驗證系統容錯能力的學科——在問題意外發生之前。
“如果你不练习失败,你实际上并不知道你的系统在失败期间会如何表现。” — Netflix
1. 混沌工程原理
1.1 定義
混沌工程并不是为了“为了好玩而破坏系统”。這是一個受控的科學過程:
1. Xác định Steady State (baseline bình thường)
2. Đặt ra Hypothesis ("nếu X xảy ra, hệ thống vẫn ổn vì...")
3. Thiết kế experiment với blast radius nhỏ
4. Chạy experiment
5. Quan sát kết quả
6. So sánh với hypothesis
7. Nếu sai → fix → verify
8. Nếu đúng → mở rộng scope
1.2 穩態假設
# Ví dụ Steady State cho e-commerce:
steady_state:
metrics:
- name: "Order success rate"
query: "rate(http_requests_total{service='order',status='201'}[5m])"
threshold: "> 0.99" # > 99% thành công
- name: "p99 latency"
query: "histogram_quantile(0.99, ...)"
threshold: "< 500ms"
- name: "Active orders being processed"
query: "order_processing_active"
threshold: "> 0" # Vẫn đang xử lý orders
只有在实验期间和实验后保持稳定状态时,混沌实验才能通过。
1.3 混沌的類型
Infrastructure Chaos:
- Pod kill / crash loop
- Node failure
- Network partition
- Resource starvation (CPU, memory)
Application Chaos:
- Latency injection (làm service chậm giả)
- Error injection (trả về 500 ngẫu nhiên)
- Dependency failure mocking
Data Chaos:
- Database failover
- Cache eviction
- Message queue backlog
Security Chaos:
- Certificate expiry simulation
- Auth service unavailability
2. 混沌工程工具
2.1 LitmusChaos(Kubernetes 原生)
LitmusChaos 是 CNCF 项目,是 Kubernetes 的云原生混沌工程平台:
# Cài đặt LitmusChaos
kubectl apply -f https://litmuschaos.github.io/litmus/litmus-operator-v3.x.yaml
# Cài ChaosCenterCLI
curl -O https://litmusctl-bucket.s3-website.us-east-2.amazonaws.com/litmusctl-linux-amd64-latest.tar.gz
2.2 Pod 殺死實驗
apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
name: pod-kill-experiment
namespace: services-prod
spec:
engineState: 'active'
appinfo:
appns: 'services-prod'
applabel: 'app=payment-service'
appkind: 'deployment'
chaosServiceAccount: litmus-admin
experiments:
- name: pod-delete
spec:
components:
env:
- name: TOTAL_CHAOS_DURATION
value: '60' # Thử nghiệm 60 giây
- name: CHAOS_INTERVAL
value: '10' # Kill pod mỗi 10 giây
- name: FORCE
value: 'false' # Graceful delete
- name: PODS_AFFECTED_PERC
value: '50' # Kill 50% pods
2.3 網路延遲注入
# Inject 500ms latency vào 50% requests từ/đến payment-service
apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
name: network-latency-experiment
spec:
experiments:
- name: pod-network-latency
spec:
components:
env:
- name: NETWORK_INTERFACE
value: 'eth0'
- name: NETWORK_LATENCY
value: '500' # 500ms latency
- name: JITTER
value: '100' # ± 100ms jitter
- name: TOTAL_CHAOS_DURATION
value: '120'
- name: PODS_AFFECTED_PERC
value: '50' # Chỉ ảnh hưởng 50% pods
2.4 CPU 壓力
- name: pod-cpu-hog
spec:
components:
env:
- name: CPU_CORES
value: '1' # Stress 1 CPU core toàn bộ
- name: TOTAL_CHAOS_DURATION
value: '60'
- name: CPU_LOAD
value: '100' # 100% CPU usage
3.設計混沌實驗
3.1 爆炸半徑控制
始終從最小的爆炸半徑開始:
Tăng dần theo giai đoạn:
Stage 1: Dev environment, 1 pod
→ Xác nhận mechanism hoạt động
Stage 2: Staging, 25% pods
→ Xác nhận hành vi dưới partial failure
Stage 3: Production, giờ thấp điểm, 10% pods
→ Xác nhận dưới real traffic
Stage 4: Production, 50% pods
→ Xác nhận major failure handling
Stage 5: Production, node failure
→ Multi-availability zone resilience
3.2 實驗模板
# chaos-experiment-template.yaml
experiment:
name: "Payment Service Pod Kill"
version: "1.0"
hypothesis: >
Khi 50% pods của payment-service bị kill,
order success rate vẫn > 95% nhờ:
(1) Kubernetes tự restart pods mới
(2) Circuit Breaker chuyển sang Half-Open khi pods quay lại
(3) Retry mechanism handle transient failures
blast_radius:
service: payment-service
percentage: 50
duration: 60s
environment: staging
steady_state_hypothesis:
before:
- metric: order_success_rate > 99%
- metric: p99_latency < 300ms
during:
- metric: order_success_rate > 95% # Cho phép degraded
- metric: p99_latency < 800ms # Latency có thể tăng
after:
- metric: order_success_rate > 99% # Phải phục hồi
- metric: p99_latency < 300ms # Về baseline
rollback:
automatic: true
trigger: order_success_rate < 90%
3.3 中止條件
Tự động dừng experiment khi:
- Error rate > 10% (threshold)
- p99 latency > 2s
- Có manual intervention từ on-call engineer
- Business-critical metric vượt ngưỡng (orders failed > X)
4. 練習:端到端的混沌實驗
4.1 場景:支付服務降級
第 1 步:確認基線
# Kiểm tra steady state trước khi bắt đầu
kubectl get pods -n services-prod -l app=payment-service
# NAME READY STATUS RESTARTS
# payment-service-abc123-1 1/1 Running 0
# payment-service-abc123-2 1/1 Running 0
# payment-service-abc123-3 1/1 Running 0
# Check metrics
curl -s prometheus:9090/api/v1/query?query=order_success_rate
第2步:注入混沌
kubectl apply -f pod-delete-experiment.yaml
# Monitor trong real-time
watch kubectl get pods -n services-prod -l app=payment-service
第 3 步:觀察
T+0s: Experiment bắt đầu, 1 pod bị kill
→ Pod restart (Kubernetes)
→ Circuit Breaker nhận một số errors
→ Retry mechanism kick in
T+10s: Pod mới lên, nhưng pod thứ 2 bị kill
→ Requests đến pod đang restart fail
→ Error rate: 2.3% (dưới threshold 5%)
T+30s: 2 pods healthy, 1 restarting
→ System degraded nhưng functional
T+60s: Experiment kết thúc, tất cả pods healthy
→ Error rate → 0%
→ Steady state phục hồi trong 15s
第 4 步:分析結果
Expected vs Actual:
Expected Actual Pass?
Error rate (peak) < 5% 2.3% ✅
Recovery time < 30s 15s ✅
Orders lost 0 0 ✅
Alert triggered Yes Yes ✅
Kết luận: Payment service resilient trước pod kill scenario
5. 比賽日
比赛日是一场有组织的演习——工程团队聚集在一起测试灾难恢复:
Game Day Schedule:
9:00 - Briefing: scenario được chọn, team assignments
9:30 - Inject chaos (ví dụ: Database failover)
9:30-11:00 - Team response, triage, mitigation
11:00 - Khôi phục hệ thống
11:30 - Post-mortem: what happened, what worked, what didn't
12:00 - Action items: fix findings, improve runbooks
熱門遊戲日場景
Scenario 1: "Primary DB region gone"
→ Test: Database failover tự động
→ Verify: Read replicas promote, application reconnect
Scenario 2: "Kafka cluster unavailable"
→ Test: Event-driven services handle backpressure
→ Verify: Outbox pattern, dead letter queue, alerts
Scenario 3: "Auth service down"
→ Test: JWT caching cho validation
→ Verify: Graceful degradation, what percentage of traffic breaks
Scenario 4: "DDoS-like traffic spike"
→ Test: Rate limiting, auto-scaling
→ Verify: System scales, bad traffic rejected, good traffic served
6. 建立韌性文化
6.1 從被動到主動
Reactive culture:
"Sự cố xảy ra → Panic → Fix → Forget"
Proactive culture:
"Chủ động inject failure → Learn → Improve → Verify"
6.2 混亂前的清單
□ Có monitoring/alerting đầy đủ
□ Team on-call biết về experiment
□ Runbook sẵn sàng cho rollback
□ Blast radius được giới hạn rõ ràng
□ Abort conditions được định nghĩa
□ Stakeholders đã được thông báo (nếu production)
□ Experiment có thể dừng ngay bằng một lệnh
6.3 混沌工程成熟度模型
Level 1 — Manual (Beginner)
Chạy chaos manually, thủ công trên dev/staging
Không có automation
Level 2 — Automated experiments trong CI
Chaos tests tự động chạy mỗi tuần
Kết quả phải pass trước khi release
Level 3 — Continuous Chaos (Advanced)
Chaos experiments chạy liên tục 24/7 ở production
Netflix Chaos Monkey — kill random instances
Luôn test khả năng tự phục hồi
Level 4 — Full Automation + GameDays
Tất cả scenarios được automated
Quarterly game days cho catastrophic scenarios
SRE team chuyên về resilience
總結
| 概念 | 目的 |
|---|---|
| 穩態假說 | 基線與預期行為的定義 |
| 爆炸半徑 | 限制實驗的影響範圍 |
| 石蕊混沌 | Kubernetes-原生混沌工具 |
| Pod 殺死 | 測試自動重新啟動和斷路器 |
| 网络延迟注入 | 测试超时和重试行为 |
| 比賽日 | 以團隊形式演練災難應變 |
| 持续的混乱 | 始终验证弹性 24/7(Netflix 级别) |
下一篇文章:微服務的 CI/CD 管道