Goal
After this lesson, you will understand:
- What Patroni is and how it works_
- DCS (Distributed Configuration Store) - etcd/Consul/ZooKeeper
- Consensus algorithm (Raft)
- Leader election & Failover mechanism
- Split-brain problem and solution
1. What is Patroni?
Introduction
Patroni is an open source HA (High Availability) template for PostgreSQL, developed by Zalando. It automates PostgreSQL cluster management, including:_
- _Leader election: Automatically select primary node_
- Automatic failover: Project transition Automatic backup when primary fails
- Configuration management: Centralized configuration management
- Health checking: Monitor the health of related nodes continued
Patroni Architecture

How Patroni Works
- Start: Each Patroni instance connects to the DCS (etcd)
- Leader election: Nodes compete to become the leader in DCS
- Role assignment: Nodes that win the leader lock will promote PostgreSQL to primary
- Health monitoring: Patroni continuously checks:
- PostgreSQL process health
- Replication status
- DCS connectivity
- Auto failover: If leader fails, Patroni automatically:
- Detect problem
- Select most suitable replica
- Promote new replica to primary
- Update remaining replicas_
Components main
Patroni daemon
- Runs on each PostgreSQL node_
- Manage lifecycle of PostgreSQL_
- Implement health checks
- Interaction with DCS
REST API_
- Endpoint for health checks:
http://node:8008/health - Endpoint for read-only:
http://node:8008/read-only - Endpoint for primary:
http://node:8008/master(deprecated) or/primary
patronictl
- CLI tool for cluster management
- Commands: list, switchover, failover, reinit, restart, reload
2. DCS - Distributed Configuration Store
DCS Role
DCS is the coordination center for the Patroni cluster, storing:
- Leader key: Information about which node is the leader (TTL-based)
- Configuration: Configuring PostgreSQL and Patroni
- Member information: List of nodes in the cluster
- Failover/Switchover state: Switching status_
Compare common DCSs variable
___HTMLTAG_18 4___| Calculation function | etcd | Consul | ZooKeepe r |
|---|---|---|---|
| Language language | Go | Go | _Java |
| Consensus | Raft_ | _Raft_ | ZAB (Paxos-like) |
| API_ | gRPC, HTTP | HTTP, DNS | Custom protocol_ |
| _Setup_ | Simple_ | Central average | Complex miscellaneous |
| Performance___HTMLTAG_223_ __ | High | High | Middle average_ |
| _Documents_ | Good | Very good | Average |
| Usage | Kubernetes, Patroni | Service mesh, HA | Hadoop, Kafka |
Recommended: etcd for most cases because of simplicity and high performance.
etcd - Distributed Key-Value Store_
Features main:
- Strongly consistent (CAP theorem: CP)_
- Distributed and highly available_
- Fast (sub-millisecond latency)
- Simple API
- Watch mechanism for real-time updates
Data structure in etcd for Patroni:
/service/postgres/
├── config # Cấu hình cluster
├── initialize # Bootstrap token
├── leader # Leader lock (TTL: 30s)
├── members/
│ ├── node1 # Thông tin node1
│ ├── node2 # Thông tin node2
│ └── node3 # Thông tin node3
├── optime/
│ └── leader # LSN của leader
└── failover # Failover/switchover instructions
3. Consensus Algorithm - Raft
What is Raft?
Raft is a consensus algorithm designed to be easier to understand than Paxos, ensuring say:
- Safety: Never return false results
- Liveness: Always progress (when majority nodes active)
- Consistency: All nodes see the same state_
Roles in Raft
- Leader:_
- Process all client requests_
- Replicate incoming log entries followers
- Unique in a term
- Follower:
- Passive, only receive requests from leader
- If not receiving heartbeat, become candidate_
- _Candidate_:_
- Follower timeout to candidate
- Request votes from other nodes
- If you win the election → Leader
Leader Process Election

Election details__HTMLTAG_354___:_
- Follower not receiving heartbeat during election timeout (150-300ms random)
- Convert to Candidate, increase term number
- Vote for yourself_
- Send RequestVote RPC to all nodes
- If received majority votes (n/2 + 1):
- Become Leader_
- Send heartbeat immediately ie
- If timeout or lose:
- Return to Follower or start election new
Quorum and Majority_
Quorum: Minimum number of nodes needed for the system to operate dynamic
Cluster size | Quorum | Tolerated failures
-------------|--------|-------------------
1 | 1 | 0
3 | 2 | 1
5 | 3 | 2
7 | 4 | 3
Formula: Quorum = floor(n/2) + 1
Example with 3 nodes:
- ✅ 3 nodes active: Cluster healthy
- ✅ 2 nodes active: Cluster works (quorum met)
- ❌ 1 active node: Cluster stops (no quorum)_
_Recommendation: Always use an odd number of nodes (3, 5, 7) to optimize faults tolerance.
4. Leader Election in Patroni
Leader Lock mechanism
Patroni uses DCS to implement distributed lock:
Leader Lock Properties:
Key: /service/postgres/leader
Value:
{
"version": "3.0.2",
"conn_url": "postgres://node1:5432/postgres",
"api_url": "http://node1:8008/patroni",
"xlog_location": 123456789,
"timeline": 2
}
TTL: 30 seconds
Leader Election Process
Step 1: Race Condition
Time: T0 - Leader crashes
Node1: Check DCS → No leader key exists
Node2: Check DCS → No leader key exists
Node3: Check DCS → No leader key exists
Step 2: Acquire Lock Attempt_
Time: T0 + 100ms
Node1: Try acquire lock → SUCCESS (first to write)
Node2: Try acquire lock → FAILED (key exists)
Node3: Try acquire lock → FAILED (key exists)
Step 3: Role Assignment_
Node1: Promote PostgreSQL to Primary
Node2: Configure as Replica, point to Node1
Node3: Configure as Replica, point to Node1
Step 4: Maintenance
Every 10 seconds: Node1 (Leader): - Renew lock (TTL extension) - Update xlog_location - Send heartbeatNode2/3 (Followers):
- Monitor leader key
- Check replication lag
Ready to take over
Best selection criteria Replica_
When failover, Patroni chooses replica based on:
- Replication state:
streaming>in archive recovery
- Timeline: Higher Timeline takes priority_
- XLog position:
- Replica has LSN closest to primary
- Less data loss most
- No replication lag:
pg_stat_replication.replay_lag = 0
- Explicit candidate: Set in configuration
Priority tag:
tags:
nofailover: false
noloadbalance: false
clonefrom: false
nosync: false
Wallet example:
Primary fails at LSN: 0/3000000Replica1: LSN=0/3000000, lag=0s ← BEST CHOICE Replica2: LSN=0/2FFFFFF, lag=1s Replica3: LSN=0/2FFFFFE, lag=2s
→ Patroni promotes Replica1
5. Failover Mechanism
Automatic Failover Process
Timeline details details:

Detailed failover steps details
Step 1: Detect failure
# Patroni health check loop while True: if not check_postgresql_health(): log.error("PostgreSQL unhealthy") stop_renewing_leader_lock()if not check_dcs_connectivity(): log.error("Lost connection to DCS") demote_if_leader() sleep(10)
Step 2: Leader lock expires
# In etcd $ etcdctl get /service/postgres/leaderAfter TTL: Key not found
Patroni logs on former leader
WARN: Could not renew leader lock INFO: Demoting PostgreSQL to standby
Step 3: Replica promotion_
# Patroni on promoted replica
INFO: No leader found
INFO: Attempting to acquire leader lock
INFO: Lock acquired successfully
INFO: Promoting PostgreSQL instance
INFO: Updating configuration
INFO: Notifying other members
Step 4: Reconfiguration
-- On promoted replica SELECT pg_promote();
-- Changes primary_conninfo to null -- Restarts as read-write
Step 5: Followers repoint_
# Other replicas
INFO: New leader detected: node2
INFO: Updating primary_conninfo
INFO: Restarting replication
Monitoring Failover_
Important Metrics_:
patroni_primary_timeline: Detect timeline changespatroni_xlog_location: Track WAL positionpatroni_replication_lag: Lag before failoverpatroni_failover_count: Count the number of times failover
6. Split-Brain Problem_
What is Split-Brain?
Definition_: Situation where ≥2 nodes think they are Primary, recording different data → Data divergence.
Cause

- Network Partition
- DCS partition: etcd cluster split_
- Slow network: Heartbeat timeout but node still live
Consequences of Split-Brain

Patroni's Split-Brain Prevention
Mechanism 1: DCS-based Lock (Primary)
def maintain_leader_lock(): while is_leader: # Must renew within TTL success = dcs.renew_lock(ttl=30)if not success: log.critical("Lost leader lock!") # Immediate demotion demote_to_standby() stop_accepting_writes() break sleep(10)
Mechanism 2: Leader Key Verification_
def before_handle_write(): leader_key = dcs.get("/service/postgres/leader")if leader_key.owner != my_node_name: # I'm not the real leader! raise Exception("Not leader anymore") demote_immediately()
Mechanism 3: Timeline Divergence Detection
-- PostgreSQL timeline SELECT timeline_id FROM pg_control_checkpoint();
-- If timelines diverge: -- Node1: timeline=5 -- Node2: timeline=6 -- → Data inconsistency detected -- → Requires pg_rewind or rebuild
Quorum requirement
etcd with 3 nodes:
Scenario 1: Network partition 1-2 split Partition A: Node1 (1 node) - Cannot get quorum (1 < 2) - Cannot write to etcd - Demotes to standby ✓Partition B: Node2, Node3 (2 nodes) - Has quorum (2 ≥ 2) - Can elect leader - Node2 becomes primary ✓
Result: Only 1 primary exists ✓
Scenario 2: Complete isolation_
Node1: Isolated, loses DCS
- Tries to renew lock → FAIL
- Demotes PostgreSQL immediately
- Stops accepting connections
Node2/3: See Node1 gone
- Elect new leader
Only 1 primary in cluster ✓
Watchdog Timer (Advanced Protection)
Hardware watchdog:
# patroni.yml
watchdog:
mode: required # or automatic, off
device: /dev/watchdog
safety_margin: 5
Active dynamic:
- Patroni kicks watchdog device every 10s_
- If Patroni hangs or loses DCS → stops kicking
- After timeout → Watchdog reboots entire node
- Prevents "zombie primary" scenario
Best Practices to avoid Split-Brain
- Deploy separate DCS: etcd cluster in different AZ_
- Monitor DCS health: Alert when etcd is not healthy
- Network redundancy: Multiple network paths between nodes
- Proper timeouts_:
patroni:
ttl: 30 # Leader lock TTL
loop_wait: 10 # Check interval
retry_timeout: 10 # DCS operation timeout
- Enable watchdog: Hardware protection layer
- Monitoring:
# Check for timeline divergence patronictl listExpected: All nodes same timeline
Cluster: postgres (7001234567890123456) ----+----+-----------+ | Member | Host | Role | State | TL | Lag in MB | +--------+--------------+---------+---------+----+-----------+ | node1 | 10.0.1.1:5432| Leader | running | 5 | | | node2 | 10.0.1.2:5432| Replica | running | 5 | 0 | | node3 | 10.0.1.3:5432| Replica | running | 5 | 0 | +--------+--------------+---------+---------+----+-----------+
Recovery from Split-Brain
If split-brain occurs:
Step 1: Identify
# Check timeline patronictl listnode1: timeline=5
node2: timeline=6 ← DIVERGED!
Step 2: Choose primary_
- Select the node with important data more
- Or node with higher timeline
Step 3: Rebuild diverged replica
# Option 1: pg_rewind (if safe) patronictl reinit postgres node2Option 2: Full rebuild
patronictl remove postgres node2
Then: reinitialize from scratch
Step 4: Verify_
patronictl listAll nodes same timeline ✓
7. Summary
Key Takeaways
✅ Patroni: HA template automates PostgreSQL management cluster
✅ DCS (etcd): Distributed coordination, store configuration and leader lock_
✅ Raft consensus: Ensure consistency and leader election in etcd
✅ Leader election: Automatic, fast (~30-40s), based on TTL locks
✅ Failover: Automatically promote the best replica when primary fails
✅ Split-brain prevention: DCS quorum + TTL locks + watchdog_
General architecture

Review Questions_
- How is Patroni different from pure Streaming Replication?
- Why do you need DCS? Can't use a database to store state?
- What is the Quorum in a cluster of 5 nodes?
- Patroni chooses which replica to promote when failover?
- Split-brain happens and how does Patroni prevent it? What does HTMLTAG_740_
- Timeline in PostgreSQL mean?
- What does TTL 30 seconds mean? Why not set TTL = 5 seconds?
Preparing for the next lesson_
Lesson 4 will guide you on preparing the infrastructure:
- Setup 3 VMs/Servers
- Network configuration, firewall
- SSH keys, time sync_
- Necessary Dependencies__HTMLTAG_758___