Chuyển đến nội dung chính

レッスン 20: 高可用性とフォールト トレランス

可用性メトリック (9)。冗長パターン: アクティブ-アクティブ、アクティブ-パッシブ。フェイルオーバー戦略。ヘルスチェックと心拍数。カオスエンジニアリングの原則。優雅な劣化。失敗の考え方をデザインする。

🏗️ アーキテクチャ — レッスン 20 レッスン 20: 高可用性と障害 許容範囲

システムアーキテクチャ: ゼロからヒーローへ

パート 6: 信頼性、セキュリティ、可観測性

xdev.asia

はじめに

「いつもすべてが失敗する。」 — Amazon CTO、Werner Vogels 氏。高可用性 (HA) は障害を防ぐことではなく、障害が発生した場合でもシステムが 動作を継続することです。


1. 可用性メトリクス

1.1 ナインズ

Availability   Downtime/year   Downtime/month   Downtime/week
99%            3.65 days       7.31 hours       1.68 hours
99.9%          8.77 hours      43.8 minutes     10.1 minutes
99.95%         4.38 hours      21.9 minutes     5.04 minutes
99.99%         52.6 minutes    4.38 minutes     1.01 minutes
99.999%        5.26 minutes    26.3 seconds     6.05 seconds

Availability = Uptime / (Uptime + Downtime)
MTBF = Mean Time Between Failures
MTTR = Mean Time To Recover
Availability = MTBF / (MTBF + MTTR)

Tăng MTBF → Ít failures hơn (khó)
Giảm MTTR → Recovery nhanh hơn (dễ hơn!)

1.2 複雑なシステムの可用性

Sequential (cả 2 phải up):
  A(99.9%) ──► B(99.9%)
  System = 99.9% × 99.9% = 99.8%

  Thêm components → Availability GIẢM!

Parallel (1 trong 2 up là đủ):
  A(99.9%) ──┐
             ├──► System
  B(99.9%) ──┘
  System = 1 - (0.1% × 0.1%) = 99.9999%

  Thêm redundancy → Availability TĂNG!

2. 冗長パターン

2.1 アクティブ/パッシブ (フェイルオーバー)

Normal:
  Traffic ──► Active Server (processing) 
              Passive Server (standby, syncing data)

Failover:
  Traffic ──► Active Server ✗ (down!)
              Passive Server → Promoted to Active
  Traffic ──► New Active Server (was passive)

Types:
  Hot Standby:  Passive chạy sẵn, failover nhanh (<30s)
  Warm Standby: Passive chạy nhưng không sync real-time
  Cold Standby: Passive tắt, bật lên khi cần (phút-giờ)

2.2 アクティブ-アクティブ

Traffic ──► Load Balancer
            ├──► Server A (processing)
            └──► Server B (processing)

Cả 2 servers đều nhận traffic
Nếu A down → B nhận 100% traffic
Không cần failover (tự động)
Tốt hơn Active-Passive nhưng phức tạp hơn
  - Session management
  - Data consistency
  - Split-brain problem

2.3 マルチレベルの冗長性

┌──────────────────────────────────────────┐
│ Region: Vietnam                          │
│                                          │
│ AZ-1               AZ-2                 │
│ ┌────────────┐     ┌────────────┐       │
│ │ LB (active)│     │ LB (active)│       │
│ ├────────────┤     ├────────────┤       │
│ │ App × 3    │     │ App × 3    │       │
│ ├────────────┤     ├────────────┤       │
│ │ DB Primary │←───►│ DB Replica │       │
│ └────────────┘     └────────────┘       │
└──────────────────────────────────────────┘

Redundancy levels:
  Process: Multiple app instances
  Server:  Multiple AZs (Availability Zones)
  Region:  Multi-region (cho global services)

3. ヘルスチェック

Types:
  1. Liveness:  "App còn sống không?"
     GET /healthz → 200 OK
     Fail → Restart container

  2. Readiness: "App sẵn sàng nhận traffic?"
     GET /readyz → 200 OK (DB connected, cache warm)
     Fail → Remove from load balancer

  3. Deep health check:
     GET /health/detailed
     {
       "status": "healthy",
       "checks": {
         "database": { "status": "up", "latency": "5ms" },
         "redis":    { "status": "up", "latency": "1ms" },
         "disk":     { "status": "up", "free": "50GB" },
         "memory":   { "status": "warning", "used": "85%" }
       }
     }

Health Check Cascade:
  Tránh: Service A health check gọi Service B
  → Service B slow → Service A "unhealthy" → Cascading!
  Correct: Health check chỉ check LOCAL resources

4. 優雅な劣化

Khi một component fail, hệ thống vẫn hoạt động
với reduced functionality

Ví dụ: E-commerce
  Normal:
    Product page: title + description + reviews + recommendations
  
  Review Service down:
    Product page: title + description + "Reviews temporarily unavailable"
  
  Recommendation Service down:
    Product page: title + description + reviews + "Popular products" (static)
  
  Search Service down:
    Homepage: Categories navigation + "Search coming back soon"

Patterns:
  1. Feature flags: Disable features instantly
  2. Fallback values: Default/cached responses
  3. Read-only mode: Disable writes, serve reads
  4. Static content: Serve cached HTML khi backend down

5. カオスエンジニアリング

5.1 原則

"Inject failures INTENTIONALLY to discover weaknesses
 BEFORE they surprise you in production"

Steps:
  1. Define "steady state" (normal behavior)
  2. Hypothesize: "System survives X failure"
  3. Inject failure (kill server, add latency, corrupt data)
  4. Observe: Did system behave as expected?
  5. Fix: Address discovered weaknesses

Tools:
  - Chaos Monkey (Netflix): Kill random instances
  - Litmus Chaos: K8s chaos engineering
  - Gremlin: Enterprise chaos platform
  - Toxiproxy: Simulate network conditions

5.2 カオス実験

Experiment 1: Kill an instance
  Action: Terminate 1 of 3 app servers
  Expected: LB routes to remaining 2, no user impact
  Verify: Error rate unchanged, latency < 2x

Experiment 2: Network partition
  Action: Block traffic between App and Database
  Expected: Circuit breaker opens, cached responses served
  Verify: Graceful error message, no crash

Experiment 3: Dependency slow
  Action: Add 5s latency to Payment Service
  Expected: Timeout after 3s, show "Try again later"
  Verify: Other features unaffected (bulkhead)

Experiment 4: Disk full
  Action: Fill disk to 100%
  Expected: Alert triggered, log rotation kicks in
  Verify: App doesn't crash, monitoring shows disk alert

6. 障害に備えた設計チェックリスト

□ Mọi external call có timeout
□ Circuit breakers cho downstream services
□ Retry với exponential backoff + jitter
□ Health check endpoints (liveness + readiness)
□ Graceful shutdown (drain connections)
□ Graceful degradation (fallbacks)
□ Data replication (ít nhất 2 copies)
□ Multi-AZ deployment
□ Auto-scaling configured
□ Runbooks cho common failures
□ Chaos experiments scheduled
□ Monitoring + alerting configured
□ Backup + restore tested regularly

概要

戦略MTBF への影響MTTR の影響複雑さ
冗長性ニュートラル⬇️ 高速フェイルオーバー中
健康診断⬆️ 早期発見⬇️ 自動回復低い
優雅な劣化ニュートラル⬇️部分サービス中
カオスエンジニアリング⬆️ 弱点を見つける⬇️ より良いランブック高
自動スケーリング⬆️ ハンドルスパイクニュートラル中

演習

  1. 可用性の計算: システムには次が含まれます: LB (99.99%) → 3 つのアプリケーション サーバー (それぞれ 99.9%、アクティブ/アクティブ) → DB プライマリ (99.95%) + DB レプリカ (99.95%、ホット スタンバイ)。全体的な可用性を計算します。

  2. フェイルオーバー設計: PostgreSQL クラスター: 1 つのプライマリ + 2 つのレプリカ。一次クラッシュ。フェイルオーバー手順の詳細を記述します: 検出、昇格、再接続、検証。

  3. カオス プラン: 電子商取引システム (API、データベース、Redis、S3、ペイメント ゲートウェイ) の 5 つのカオス実験を設計します。各実験: アクション、仮説、予想される動作、ロールバック計画。