🎯 LESSON OBJECTIVE__HTMLTAG_68___
- ✅ Configure backup destination (S3/Ceph Object Store)
- ✅ Setup ScheduledBackup CRD for automated backups
- ✅ Practice Point-in-Time Recovery (PITR)
- ✅ Restore cluster from backup
- ✅ Disaster recovery procedures
PART 1: BACKUP ARCHITECTURE
CloudNativePG Backup Flow:
┌──────────────────────────────────────────────────────┐
│ Kubernetes Cluster │
│ │
│ ┌──────────────┐ ┌──────────────────────────┐ │
│ │ Primary │────►│ Barman Cloud Plugin │ │
│ │ pg-1 │ │ - Base backup │ │
│ │ │ │ - WAL archiving │ │
│ └──────────────┘ └───────────┬──────────────┘ │
│ │ │
└───────────────────────────────────┼───────────────────┘
│
▼
┌──────────────────────┐
│ Object Storage │
│ - Ceph RGW (S3) │
│ - MinIO │
│ - AWS S3 │
│ │
│ /base/ │
│ YYYYMMDDTHHMMSS/ │
│ /wals/ │
│ 00000001/ │
└──────────────────────┘
Recovery:
- Full restore: Base backup + ALL WALs
- PITR: Base backup + WALs up to target timestamp
PART 2: SETUP BACKUP DESTINATION__HTMLTAG_86___
2.1. Use Ceph RGW (S3-compatible)
# Tạo CephObjectStore (nếu chưa có):
apiVersion: ceph.rook.io/v1
kind: CephObjectStore
metadata:
name: pg-backup-store
namespace: rook-ceph
spec:
metadataPool:
replicated:
size: 3
dataPool:
replicated:
size: 3
gateway:
port: 80
instances: 2
# Tạo S3 user cho backup:
kubectl -n rook-ceph exec deploy/rook-ceph-tools -- \
radosgw-admin user create --uid=pg-backup --display-name="PG Backup"
# Tạo bucket:
# aws s3 --endpoint-url http://rook-ceph-rgw-pg-backup-store.rook-ceph mb s3://pg-backups
# Tạo secret cho backup credentials:
kubectl -n database create secret generic pg-backup-s3-creds \
--from-literal=ACCESS_KEY_ID="" \
--from-literal=ACCESS_SECRET_KEY=""
2.2. Update Cluster with Backup Config
# Thêm vào pg-cluster.yaml spec:
backup:
barmanObjectStore:
destinationPath: s3://pg-backups/production-pg
endpointURL: http://rook-ceph-rgw-pg-backup-store.rook-ceph:80
s3Credentials:
accessKeyId:
name: pg-backup-s3-creds
key: ACCESS_KEY_ID
secretAccessKey:
name: pg-backup-s3-creds
key: ACCESS_SECRET_KEY
wal:
compression: gzip
maxParallel: 4
data:
compression: gzip
immediateCheckpoint: true
retentionPolicy: "30d" # Giữ backup 30 ngày
PART 3: SCHEDULED BACKUP
# scheduled-backup.yaml:
apiVersion: postgresql.cnpg.io/v1
kind: ScheduledBackup
metadata:
name: production-pg-daily-backup
namespace: database
spec:
schedule: "0 2 * * *" # Mỗi ngày lúc 2 AM
backupOwnerReference: self
cluster:
name: production-pg
immediate: true # Tạo backup ngay lập tức (lần đầu)
target: prefer-standby # Backup từ standby (giảm load primary)
kubectl apply -f scheduled-backup.yaml
# Verify backup:
kubectl -n database get backup
# NAME AGE CLUSTER METHOD PHASE STARTED COMPLETED
# production-pg-daily-backup-20250402 5m production-pg barman completed 2025-04-02T02:00:00Z 2025-04-02T02:05:00Z
# Manual backup:
kubectl -n database create -f - <<'EOF'
apiVersion: postgresql.cnpg.io/v1
kind: Backup
metadata:
name: manual-backup-$(date +%Y%m%d)
namespace: database
spec:
cluster:
name: production-pg
target: prefer-standby
EOF
PART 4: POINT-IN-TIME RECOVERY (PITR)
4.1. Simulate Data Loss
# Insert important data: kubectl -n database exec production-pg-1 -- psql -U appuser -d appdb -c \ "INSERT INTO test (data) VALUES ('important data 1'); INSERT INTO test (data) VALUES ('important data 2'); INSERT INTO test (data) VALUES ('important data 3');"Ghi nhận timestamp TRƯỚC KHI XÓA:
RECOVERY_TARGET="2025-04-02 08:00:00"
⚠️ Simulate accidental delete:
kubectl -n database exec production-pg-1 -- psql -U appuser -d appdb -c
"DELETE FROM test;"DELETE 3 ← Dữ liệu bị xóa!
4.2. Restore with PITR
# pitr-restore.yaml — Tạo cluster MỚI từ backup: apiVersion: postgresql.cnpg.io/v1 kind: Cluster metadata: name: production-pg-restored namespace: database spec: instances: 3 imageName: ghcr.io/cloudnative-pg/postgresql:16.4storage: storageClass: ceph-block size: 50Gi
bootstrap: recovery: source: production-pg-backup recoveryTarget: targetTime: "2025-04-02T08:00:00Z" # Restore tới thời điểm TRƯỚC delete
externalClusters: - name: production-pg-backup barmanObjectStore: destinationPath: s3://pg-backups/production-pg endpointURL: http://rook-ceph-rgw-pg-backup-store.rook-ceph:80 s3Credentials: accessKeyId: name: pg-backup-s3-creds key: ACCESS_KEY_ID secretAccessKey: name: pg-backup-s3-creds key: ACCESS_SECRET_KEY wal: maxParallel: 4
kubectl apply -f pitr-restore.yaml
# Monitor restore:
kubectl -n database get cluster production-pg-restored -w
# Đợi cluster healthy
# Verify data recovered:
kubectl -n database exec production-pg-restored-1 -- psql -U appuser -d appdb -c \
"SELECT * FROM test;"
# id | data | created_at
# ---+--------------------+----------------------------
# 1 | important data 1 | 2025-04-02 07:55:00
# 2 | important data 2 | 2025-04-02 07:55:01
# 3 | important data 3 | 2025-04-02 07:55:02
# ✅ Data recovered to point-in-time TRƯỚC delete!
# Switchover: đổi app connection sang cluster restored
# (hoặc rename cluster)
PART 5: DISASTER RECOVERY PROCEDURES
5.1. Full Cluster Restore (no PITR)
# full-restore.yaml:
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
name: production-pg-dr
namespace: database
spec:
instances: 3
imageName: ghcr.io/cloudnative-pg/postgresql:16.4
storage:
storageClass: ceph-block
size: 50Gi
bootstrap:
recovery:
source: production-pg-backup
# Không có recoveryTarget → restore latest
externalClusters:
- name: production-pg-backup
barmanObjectStore:
destinationPath: s3://pg-backups/production-pg
endpointURL: http://rook-ceph-rgw-pg-backup-store.rook-ceph:80
s3Credentials:
accessKeyId:
name: pg-backup-s3-creds
key: ACCESS_KEY_ID
secretAccessKey:
name: pg-backup-s3-creds
key: ACCESS_SECRET_KEY
5.2. Verify Backup Integrity
# List backups: kubectl -n database get backup --sort-by=.status.startedAt # NAME PHASE STARTED COMPLETED # production-pg-daily-backup-20250401 completed 2025-04-01T02:00:00Z 2025-04-01T02:03:00Z # production-pg-daily-backup-20250402 completed 2025-04-02T02:00:00Z 2025-04-02T02:04:00ZWAL archiving status:
kubectl -n database get cluster production-pg -o jsonpath='{.status.firstRecoverabilityPoint}'
2025-03-03T02:00:00Z ← Có thể restore tới 30 ngày trước
💡 KEY TAKEAWAYS
- Barman + S3: Base backup + continuous WAL archiving
- ScheduledBackup: Automated daily backup, prefer-standby
- PITR: Restore to any point in retention period
- Recovery = Create new cluster from backup, then switchover
- retentionPolicy: 30d: Keep backup 30 days
- Test restore periodically: Verify monthly backup integrity
🎯 EXERCISES
Exercise 1: Backup Lab
- Configure backup destination (S3/MinIO)
- Setup ScheduledBackup, verify backup created successfully
Exercise 2: PITR Lab
- Insert data, note timestamp
- Delete data (simulate accident)
- Restore to timestamp before delete
- Verify data recovered__HTMLTAG_158___
📚 NEXT POST
In Lesson 19: PostgreSQL Failover Testing and Switchover, we will test failover scenarios and planned switchover.