Chuyển đến nội dung chính

LESSON 44: DISASTER RECOVERY & BACKUP STRATEGIES

DR planning for on-premises K8s, Velero backup/restore, etcd disaster recovery, cross-site replication, RPO/RTO targets, and DR runbook automation.

🔒 DevSecOps — Lesson 44 LESSON 44: DISASTER RECOVERY & BACKUP STARTEGIES

Deploy Microservices On-Premises with Kubernetes HA

Part 11: Disaster Recovery & Chaos Engineering

xdev.asia

🎯 LESSON OBJECTIVE__HTMLTAG_68___
  • ✅ DR planning: RPO, RTO, failure scenarios
  • ✅ Velero backup & restore for K8s resources + PVs
  • ✅ etcd snapshot backup & restore
  • ✅ Cross-site DR strategies
  • ✅ DR testing and runbook automation

PART 1: DR PLANNING


Failure Scenarios & DR Strategy:

Level 1: Pod/Container failure
  → K8s auto-restart (ReplicaSet, liveness probe)
  → RPO: 0  |  RTO: seconds

Level 2: Node failure
  → K8s reschedule pods to other nodes
  → RPO: 0  |  RTO: minutes

Level 3: Storage failure
  → Ceph replication (replica 3)
  → RPO: 0  |  RTO: seconds

Level 4: Control plane failure
  → HA control plane (3 masters)
  → RPO: 0  |  RTO: seconds

Level 5: Entire cluster failure
  → Velero restore + etcd snapshot
  → RPO: last backup  |  RTO: 1-4 hours

Level 6: Data center / site failure
  → Cross-site DR (active-passive)
  → RPO: minutes  |  RTO: hours

┌──────────────────┬───────────┬───────────┐
│ Component        │   RPO     │   RTO     │
├──────────────────┼───────────┼───────────┤
│ K8s resources    │ 1 hour    │ 30 min    │
│ PostgreSQL       │ 5 min     │ 15 min    │
│ etcd             │ 1 hour    │ 30 min    │
│ Ceph data        │ 0 (3x)   │ seconds   │
│ Configurations   │ 0 (Git)   │ minutes   │
└──────────────────┴───────────┴───────────┘

PART 2: VELERO BACKUP & RESTORE__HTMLTAG_86___
# Install Velero:
helm repo add vmware-tanzu https://vmware-tanzu.github.io/helm-charts
helm install velero vmware-tanzu/velero \
  --namespace velero \
  --create-namespace \
  -f velero-values.yaml
# velero-values.yaml:
configuration:
  backupStorageLocation:
    - name: default
      provider: aws
      bucket: velero-backups
      config:
        region: us-east-1
        s3ForcePathStyle: "true"
        s3Url: http://ceph-rgw.storage:8080

  volumeSnapshotLocation:
    - name: default
      provider: csi

credentials:
  secretContents:
    cloud: |
      [default]
      aws_access_key_id=velero
      aws_secret_access_key=velero-secret

initContainers:
  - name: velero-plugin-for-aws
    image: velero/velero-plugin-for-aws:v1.9.0
    volumeMounts:
      - mountPath: /target
        name: plugins
  - name: velero-plugin-for-csi
    image: velero/velero-plugin-for-csi:v0.7.0
    volumeMounts:
      - mountPath: /target
        name: plugins

snapshotsEnabled: true

schedules:
  daily-backup:
    disabled: false
    schedule: "0 2 * * *"
    template:
      ttl: 720h
      includedNamespaces:
        - default
        - production
        - database
      snapshotVolumes: true
      storageLocation: default
      volumeSnapshotLocations:
        - default
# Manual backup:
velero backup create full-backup \
  --include-namespaces default,production,database \
  --snapshot-volumes

# List backups:
velero backup get

# Restore from backup:
velero restore create --from-backup full-backup \
  --include-namespaces production

# Restore specific resources:
velero restore create --from-backup full-backup \
  --include-resources deployments,services,configmaps \
  --include-namespaces production

# Backup status:
velero backup describe full-backup --details

PART 3: ETCD DISASTER RECOVERY

# Automated etcd backup (CronJob):
# (Already covered in Bài 10, recap here for DR context)

# Backup etcd snapshot:
ETCDCTL_API=3 etcdctl snapshot save /backup/etcd-$(date +%Y%m%d-%H%M).db \
  --endpoints=https://127.0.0.1:2379 \
  --cacert=/etc/kubernetes/pki/etcd/ca.crt \
  --cert=/etc/kubernetes/pki/etcd/server.crt \
  --key=/etc/kubernetes/pki/etcd/server.key

# Restore etcd from snapshot (DR scenario):
# 1. Stop kube-apiserver on all masters
# 2. Restore on each etcd member:
ETCDCTL_API=3 etcdctl snapshot restore /backup/etcd-latest.db \
  --name master-1 \
  --initial-cluster master-1=https://10.0.1.11:2380,master-2=https://10.0.1.12:2380,master-3=https://10.0.1.13:2380 \
  --initial-cluster-token etcd-cluster-1 \
  --initial-advertise-peer-urls https://10.0.1.11:2380 \
  --data-dir=/var/lib/etcd-restore

# 3. Replace /var/lib/etcd with restored data
# 4. Restart etcd + kube-apiserver

PART 4: CROSS-SITE DR


Cross-Site DR Architecture:

Site A (Primary)                Site B (DR)
┌──────────────────┐            ┌──────────────────┐
│  K8s Cluster     │            │  K8s Cluster     │
│  (Active)        │            │  (Standby)       │
│                  │            │                  │
│  ┌────────────┐  │  replicate │  ┌────────────┐  │
│  │ PostgreSQL │──┼───────────►│  │ PostgreSQL │  │
│  │ (Primary)  │  │  streaming │  │ (Replica)  │  │
│  └────────────┘  │            │  └────────────┘  │
│                  │            │                  │
│  ┌────────────┐  │  S3 sync  │  ┌────────────┐  │
│  │ Ceph/S3    │──┼───────────►│  │ Ceph/S3    │  │
│  │ (backups)  │  │            │  │ (backups)  │  │
│  └────────────┘  │            │  └────────────┘  │
│                  │            │                  │
│  ┌────────────┐  │  GitOps   │  ┌────────────┐  │
│  │ ArgoCD     │  │  (shared) │  │ ArgoCD     │  │
│  └────────────┘  │            │  └────────────┘  │
└──────────────────┘            └──────────────────┘

Failover: DNS switch + promote PostgreSQL replica

PART 5: DR RUNBOOK

# DR Runbook — Full Cluster Recovery:

# Step 1: Restore infrastructure (Ansible/Terraform)
ansible-playbook -i inventory/dr site.yml

# Step 2: Bootstrap K8s cluster
kubeadm init --config kubeadm-dr-config.yaml

# Step 3: Restore etcd
etcdctl snapshot restore /backup/etcd-latest.db ...

# Step 4: Install core components (Cilium, MetalLB, Rook-Ceph)
helmfile -f helmfile-core.yaml apply

# Step 5: Restore Velero backups
velero restore create --from-backup latest-daily

# Step 6: Restore databases (PostgreSQL PITR)
kubectl apply -f postgresql-restore-cluster.yaml

# Step 7: Verify services
./scripts/verify-all-services.sh

# Step 8: Switch DNS
# Update DNS records to point to DR site
# DR testing schedule:
# Monthly:  Velero backup/restore test
# Quarterly: Full cluster DR drill
# Annually:  Cross-site failover test

# Monitoring DR readiness:
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: dr-alerts
spec:
  groups:
    - name: dr-readiness
      rules:
        - alert: VeleroBackupFailed
          expr: |
            velero_backup_failure_total > velero_backup_success_total
          for: 5m
          labels:
            severity: critical

        - alert: VeleroBackupStale
          expr: |
            time() - velero_backup_last_successful_timestamp > 86400
          for: 1h
          labels:
            severity: warning
          annotations:
            summary: "No successful Velero backup in 24h"

💡 KEY TAKEAWAYS

  1. RPO/RTO: Define targets per component before disaster
  2. Velero: K8s resource + PV backup to S3/Ceph
  3. etcd: Regular snapshots, tested restore procedure
  4. GitOps: Infrastructure-as-Code = instant re-deploy
  5. Cross-site DR: PostgreSQL streaming replication + S3 sync
  6. Test regularly: Untested backups are not backups

🎯 EXERCISES__HTMLTAG_127___

Exercise 1: Velero DR

  • Setup Velero with S3 backend
  • Create scheduled backup for all namespaces__HTMLTAG_134___
  • Simulate namespace deletion → restore from backup

Exercise 2: DR Drill

  • Document DR runbook for your cluster
  • Perform etcd restore drill__HTMLTAG_144___
  • Measure actual RTO vs target

📚 NEXT POST

In Lesson 45: Chaos Engineering with Litmus, we will proactively test system resilience.