簡介
當整個資料中心出現問題(停電、自然災害、網路中斷)時,一個區域的高可用性是不夠的。 災難復原 (DR) 確保系統可以從災難性故障中復原。
1.RPO 和 RTO
RPO (Recovery Point Objective):
"Mất tối đa bao nhiêu data có thể chấp nhận?"
RPO = 1 giờ → Backup mỗi giờ → Mất tối đa 1 giờ data
RPO = 0 → Synchronous replication → Không mất data
RTO (Recovery Time Objective):
"Hệ thống phải recovery trong bao lâu?"
RTO = 4 giờ → Có 4 giờ để khôi phục
RTO = 0 → Instant failover (Active-Active)
Timeline:
◄──── RPO ────► ◄──── RTO ────►
Last backup Disaster Start Recovered
│ │ Recovery │
────┼──────────────┼───────┼─────────────────┼────
▲ ▲ ▲ ▲
Data preserved Data lost Downtime Back online
2. 災難復原策略
2.1 備份與還原
┌──────────────┐ ┌──────────────┐
│ Primary │ backup │ S3/GCS │
│ Region │────────►│ (cold store) │
│ (running) │ └──────┬───────┘
└──────────────┘ │ restore
┌─────▼────────┐
│ DR Region │
│ (provisioned │
│ on demand) │
└──────────────┘
RPO: Hours (last backup)
RTO: Hours (provision + restore)
Cost: $ (chỉ trả storage)
Use case: Non-critical systems, dev/staging
2.2 指示燈
┌──────────────┐ ┌──────────────┐
│ Primary │ replica │ DR Region │
│ Region │────────►│ │
│ App servers │ │ DB Replica │ ← chạy sẵn
│ DB Primary │ │ (no app) │
│ Cache │ │ │
└──────────────┘ └──────────────┘
Disaster → Scale up DR:
1. Promote DB replica → Primary
2. Launch app servers (AMI/container)
3. Update DNS → DR region
RPO: Minutes (async replication)
RTO: 10-30 minutes
Cost: $$ (DB replica running)
2.3 熱備
┌──────────────┐ ┌──────────────┐
│ Primary │ replica │ DR Region │
│ Region │────────►│ │
│ App × 10 │ │ App × 2 │ ← scaled down
│ DB Primary │ │ DB Replica │
│ Cache │ │ Cache │
└──────────────┘ └──────────────┘
Disaster → Scale up DR:
1. Scale app 2 → 10
2. Promote DB
3. Switch traffic
RPO: Seconds-minutes
RTO: Minutes
Cost: $$$ (minimal infra running)
2.4 多站點主動-主動
┌──────────────┐ ┌──────────────┐
│ Region A │◄───────►│ Region B │
│ App × 10 │ sync │ App × 10 │
│ DB Primary │ │ DB Primary │
│ Full traffic │ │ Full traffic │
└──────────────┘ └──────────────┘
▲ ▲
└────── Global LB ──────┘
(GeoDNS/Anycast)
RPO: 0 (synchronous) hoặc seconds (async)
RTO: 0 (automatic failover)
Cost: $$$$ (2x infrastructure)
Use case: Mission-critical, global services
2.5 比較
| 戰略 | 恢復點目標 | RTO | 成本 | 複雜度 |
|---|---|---|---|---|
| 備份與恢復 | 營業時間 | 營業時間 | $ | 低 |
| 指示燈 | 分鐘 | 10-30 分鐘 | $$ | 中 |
| 溫暖待機 | 秒 | 分鐘 | $$$ | 高 |
| 主動-主動 | 〜0 | 〜0 | $$$$ | 非常高 |
3. 多區域資料挑戰
3.1 資料複製
Synchronous:
Region A write → Wait for Region B confirm → Return
✅ Strong consistency (RPO=0)
❌ Latency tăng (cross-region: 50-200ms)
❌ Region B down → Region A blocked
Asynchronous:
Region A write → Return immediately
Region A → replicate to Region B (background)
✅ Nhanh, Region B down không ảnh hưởng
❌ Replication lag → Data inconsistency
❌ RPO > 0 (có thể mất data)
Conflict Resolution (Active-Active):
Region A: UPDATE user SET name='Alice'
Region B: UPDATE user SET name='Bob' (cùng lúc)
→ Conflict! Giải quyết bằng:
- Last-Writer-Wins (LWW)
- Application-level merge
- CRDTs
3.2 全域負載平衡
GeoDNS:
User ở Vietnam → DNS trả IP Region Asia
User ở US → DNS trả IP Region US
user.example.com
├── Vietnam user → 10.0.1.1 (Asia region)
├── US user → 10.0.2.1 (US region)
└── EU user → 10.0.3.1 (EU region)
Anycast:
Cùng 1 IP, nhiều locations
BGP routing → nearest location
Dùng cho CDN, DNS servers
Latency-based:
AWS Route 53: Route to region có latency thấp nhất
Health check: Nếu region down → route sang region khác
4. 災難復原測試
1. Tabletop Exercise (hàng quý):
Team ngồi lại, giả lập scenario trên giấy
"Database chính bị corrupt, phải làm gì?"
Lên plan, identify gaps
2. Failover Test (hàng quý-năm):
Thực sự failover sang DR region
Verify data integrity
Measure actual RTO
3. Chaos Day (hàng tháng):
Inject failures vào production
Kill processes, add latency
Netflix "Chaos Monkey" style
4. Backup Restore Test (hàng tháng):
Restore backup to new environment
Verify data completeness
Measure restore time
5.災難復原運作手冊模板
# Runbook: Database Failover
## Trigger Conditions
- Primary DB unreachable > 5 minutes
- Data corruption detected
- Region-level outage declared
## Pre-Checks
□ Verify primary is actually down (not network issue)
□ Check replica lag (pg_stat_replication)
□ Notify on-call team lead
## Failover Steps
1. Stop writes to primary (if accessible)
2. Verify replica is caught up
3. Promote replica: SELECT pg_promote()
4. Update connection strings (DNS/config)
5. Verify app connects to new primary
6. Monitor error rates for 15 minutes
## Post-Failover
□ Notify stakeholders
□ Update status page
□ Create incident ticket
□ Plan for original primary recovery
□ Post-mortem within 48 hours
## Rollback Plan
If failover fails:
1. Restore from latest backup
2. Point apps to restored DB
3. Accept data loss from RPO gap
總結
| 決定 | 因素 |
|---|---|
| RPO選擇 | 資料關鍵性、合規性、成本 |
| RTO選擇 | 每分鐘停機時間對業務的影響 |
| 災難復原策略 | 預算、RPO/RTO 需求 |
| 地區 | 用戶位置、合規性、延遲 |
| 測試 | 頻率與重要性相符 |
練習
-
災難復原策略選擇: 金融科技應用程式:交易、使用者餘額、合規報告。停機時間 > 30 分鐘 = 違反法規。 RPO = 0,RTO < 5 分鐘。預算:DR 每月 5 萬美元。選擇哪種災難復原策略?
-
多區域設計: 東南亞社交媒體應用程式。越南 (60%)、印尼 (20%)、菲律賓 (10%) 和其他 (10%) 的使用者。設計多區域架構。哪些資料被複製,哪些資料會依區域分片?
-
DR Runbook: 針對場景撰寫詳細的 DR Runbook:AWS ap-southeast-1(新加坡)完全不可用。系統包括:EKS叢集、RDS PostgreSQL、ElastiCache Redis、S3。 DR 站點:ap-northeast-1(東京)。