🎯 MỤC TIÊU BÀI HỌC
- ✅ Test automatic failover khi primary pod bị kill
- ✅ Planned switchover (maintenance window)
- ✅ Monitor replication lag
- ✅ Application connection handling during failover
- ✅ Fencing và split-brain prevention
PHẦN 1: AUTOMATIC FAILOVER
1.1. Test: Kill Primary Pod
# Current state: kubectl cnpg status production-pg -n database # Primary: production-pg-1Kill primary pod:
kubectl -n database delete pod production-pg-1
Watch failover (real-time):
kubectl -n database get pods -w
production-pg-1 1/1 Terminating 0 1h
production-pg-2 1/1 Running 0 1h ← Promoted to Primary!
production-pg-3 1/1 Running 0 1h
production-pg-1 0/1 Pending 0 0s ← Recreating as Standby
Verify new primary:
kubectl cnpg status production-pg -n database
Primary instance: production-pg-2 ← New primary!
Check services updated:
kubectl -n database get endpoints production-pg-rw
ENDPOINTS
10.244.2.5:5432 ← production-pg-2 IP
⏱️ Failover time: ~5-15 seconds
1.2. Failover Timeline
Timeline: T+0s: Primary pod deleted T+1s: K8s detects pod termination T+3s: Operator detects primary failure T+5s: Operator selects standby with highest LSN T+8s: Standby promoted to primary (pg_promote) T+10s: Service rw endpoint updated T+12s: Other standbys repoint to new primary T+15s: Old primary restarts as standby
Total downtime: ~10-15 seconds (write unavailability) Read queries: ~0s downtime (ro service still serves)
PHẦN 2: PLANNED SWITCHOVER
2.1. Planned Switchover (zero data loss)
# Switchover = planned promotion (graceful) # Use case: maintenance, node drain, rebalancingInstall cnpg kubectl plugin:
curl -sSfL https://github.com/cloudnative-pg/cloudnative-pg/releases/latest/download/kubectl-cnpg_linux_x86_64.tar.gz |
tar xz -C /usr/local/binPlanned switchover tới pg-3:
kubectl cnpg promote production-pg production-pg-3 -n database
Monitor:
kubectl cnpg status production-pg -n database
Primary instance: production-pg-3 ← Switched!
Switchover process:
1. Operator stops writes on current primary
2. Wait for all standbys to catch up (WAL replay)
3. Promote target standby → Primary
4. Demote old primary → Standby
5. Update services
→ Zero data loss! ✅
PHẦN 3: MONITORING REPLICATION LAG
# Check replication lag on Primary:
kubectl -n database exec production-pg-3 -- psql -U postgres -c \
"SELECT
client_addr,
application_name,
state,
sync_state,
sent_lsn,
write_lsn,
flush_lsn,
replay_lsn,
pg_wal_lsn_diff(sent_lsn, replay_lsn) AS replay_lag_bytes,
write_lag,
flush_lag,
replay_lag
FROM pg_stat_replication;"
# Output:
# client_addr | application_name | state | replay_lag_bytes | replay_lag
# 10.244.1.5 | production-pg-1 | streaming | 0 | 00:00:00.001
# 10.244.2.5 | production-pg-2 | streaming | 0 | 00:00:00.001
# ✅ replay_lag_bytes = 0 → fully caught up
# ⚠️ replay_lag_bytes > 0 → replication delay
3.1. Alerting on Replication Lag
# pg-alerts.yaml: apiVersion: monitoring.coreos.com/v1 kind: PrometheusRule metadata: name: postgresql-alerts namespace: database spec: groups: - name: postgresql rules: - alert: PostgreSQLReplicationLag expr: cnpg_pg_replication_streaming_replics < 2 for: 5m labels: severity: warning annotations: summary: "PostgreSQL replication lag detected"- alert: PostgreSQLDown expr: cnpg_collector_up == 0 for: 1m labels: severity: critical annotations: summary: "PostgreSQL instance down"
PHẦN 4: APPLICATION CONNECTION HANDLING
4.1. Xử lý failover trong ứng dụng
# Application deployment sử dụng service names:
apiVersion: apps/v1
kind: Deployment
metadata:
name: my-app
namespace: default
spec:
template:
spec:
containers:
- name: app
env:
# Write connection (via PgBouncer):
- name: DATABASE_URL
value: "postgresql://appuser:$(PG_PASSWORD)@production-pg-pooler-rw.database:5432/appdb?sslmode=require"
# Read connection:
- name: DATABASE_READ_URL
value: "postgresql://appuser:$(PG_PASSWORD)@production-pg-pooler-ro.database:5432/appdb?sslmode=require"
- name: PG_PASSWORD
valueFrom:
secretKeyRef:
name: pg-app-secret
key: password
# Python app connection handling:
import psycopg2
from psycopg2 import OperationalError
import time
def get_connection(retries=3, delay=2):
"""Retry connection on failover"""
for attempt in range(retries):
try:
conn = psycopg2.connect(
host="production-pg-pooler-rw.database",
port=5432,
dbname="appdb",
user="appuser",
password=os.environ["PG_PASSWORD"],
connect_timeout=5,
# Auto-reconnect params:
target_session_attrs="read-write"
)
return conn
except OperationalError as e:
if attempt < retries - 1:
time.sleep(delay)
else:
raise
PHẦN 5: FENCING VÀ SPLIT-BRAIN PREVENTION
Split-brain scenario:
┌──────────┐ Network ┌──────────┐
│ pg-1 │◄──── SPLIT ────►│ pg-2 │
│ PRIMARY │ partition │ PRIMARY? │ ⚠️ TWO PRIMARIES!
└──────────┘ └──────────┘
CloudNativePG prevention:
1. K8s Lease-based fencing:
- Primary phải maintain K8s Lease
- Nếu không renew → tự demote
2. Pod deletion fencing:
- Old primary pod bị delete khi new primary promoted
- K8s ensures only 1 primary at a time
3. pg_rewind:
- Khi old primary restart → pg_rewind tự fix timeline
- Join lại cluster as standby
💡 KEY TAKEAWAYS
- Automatic failover: ~10-15 seconds, operator detects + promotes
- Planned switchover: zero data loss,
kubectl cnpg promote - Replication lag monitoring: pg_stat_replication, Prometheus metrics
- Application: Dùng service names, retry logic, target_session_attrs
- Fencing: K8s Lease prevents split-brain
- PgBouncer hides failover from applications (transparent reconnect)
🎯 BÀI TẬP
Bài tập 1: Failover Lab
- Kill primary pod, measure failover time
- Verify data consistency after failover
- Planned switchover to specific standby
Bài tập 2: App Connection Test
- Deploy simple app connecting to PG via PgBouncer
- Trigger failover while app is writing
- Verify app recovers automatically
📚 BÀI TIẾP THEO
Trong Bài 20: PostgreSQL Monitoring, Tuning và Day-2 Operations, chúng ta sẽ setup monitoring chi tiết và tuning cho production.