Chuyển đến nội dung chính

Lesson 3: Introducing Patroni and etcd

Understand how Patroni works, the role of DCS (etcd/Consul/ZooKeeper), Raft consensus algorithm and automatic leader election mechanism.

Goal

After this lesson, you will understand:

  • What Patroni is and how it works_
  • DCS (Distributed Configuration Store) - etcd/Consul/ZooKeeper
  • Consensus algorithm (Raft)
  • Leader election & Failover mechanism
  • Split-brain problem and solution

1. What is Patroni?

Introduction

Patroni is an open source HA (High Availability) template for PostgreSQL, developed by Zalando. It automates PostgreSQL cluster management, including:_

  • _Leader election: Automatically select primary node_
  • Automatic failover: Project transition Automatic backup when primary fails
  • Configuration management: Centralized configuration management
  • Health checking: Monitor the health of related nodes continued

Patroni Architecture

The Patroni architecture is a popular choice for managing PostgreSQL clusters.

How Patroni Works

  1. Start: Each Patroni instance connects to the DCS (etcd)
  2. Leader election: Nodes compete to become the leader in DCS
  3. Role assignment: Nodes that win the leader lock will promote PostgreSQL to primary
  4. Health monitoring: Patroni continuously checks:
    • PostgreSQL process health
    • Replication status
    • DCS connectivity
  5. Auto failover: If leader fails, Patroni automatically:
    • Detect problem
    • Select most suitable replica
    • Promote new replica to primary
    • Update remaining replicas_

Components main

Patroni daemon

  • Runs on each PostgreSQL node_
  • Manage lifecycle of PostgreSQL_
  • Implement health checks
  • Interaction with DCS

REST API_

  • Endpoint for health checks: http://node:8008/health
  • Endpoint for read-only: http://node:8008/read-only
  • Endpoint for primary: http://node:8008/master (deprecated) or /primary

patronictl

  • CLI tool for cluster management
  • Commands: list, switchover, failover, reinit, restart, reload

2. DCS - Distributed Configuration Store

DCS Role

DCS is the coordination center for the Patroni cluster, storing:

  • Leader key: Information about which node is the leader (TTL-based)
  • Configuration: Configuring PostgreSQL and Patroni
  • Member information: List of nodes in the cluster
  • Failover/Switchover state: Switching status_

Compare common DCSs variable

___HTMLTAG_18 4___
Calculation functionetcdConsulZooKeepe r
Language languageGoGo_Java
ConsensusRaft__Raft_ZAB (Paxos-like)
API_gRPC, HTTPHTTP, DNSCustom protocol_
_Setup_Simple_Central averageComplex miscellaneous
Performance___HTMLTAG_223_ __HighHighMiddle average_
_Documents_GoodVery goodAverage
UsageKubernetes, PatroniService mesh, HAHadoop, Kafka

Recommended: etcd for most cases because of simplicity and high performance.

etcd - Distributed Key-Value Store_

Features main:

  • Strongly consistent (CAP theorem: CP)_
  • Distributed and highly available_
  • Fast (sub-millisecond latency)
  • Simple API
  • Watch mechanism for real-time updates

Data structure in etcd for Patroni:

/service/postgres/
├── config          # Cấu hình cluster
├── initialize      # Bootstrap token
├── leader          # Leader lock (TTL: 30s)
├── members/
│   ├── node1      # Thông tin node1
│   ├── node2      # Thông tin node2
│   └── node3      # Thông tin node3
├── optime/
│   └── leader     # LSN của leader
└── failover       # Failover/switchover instructions

3. Consensus Algorithm - Raft

What is Raft?

Raft is a consensus algorithm designed to be easier to understand than Paxos, ensuring say:

  • Safety: Never return false results
  • Liveness: Always progress (when majority nodes active)
  • Consistency: All nodes see the same state_

Roles in Raft

  1. Leader:_
    • Process all client requests_
    • Replicate incoming log entries followers
    • Unique in a term
  2. Follower:
    • Passive, only receive requests from leader
    • If not receiving heartbeat, become candidate_
  3. _Candidate_:_
    • Follower timeout to candidate
    • Request votes from other nodes
    • If you win the election → Leader

Leader Process Election

The Leader Election Process is crucial for ensuring the consistency and availability of the distributed system.

Election details__HTMLTAG_354___:_

  1. Follower not receiving heartbeat during election timeout (150-300ms random)
  2. Convert to Candidate, increase term number
  3. Vote for yourself_
  4. Send RequestVote RPC to all nodes
  5. If received majority votes (n/2 + 1):
    • Become Leader_
    • Send heartbeat immediately ie
  6. If timeout or lose:
    • Return to Follower or start election new

Quorum and Majority_

Quorum: Minimum number of nodes needed for the system to operate dynamic

Cluster size | Quorum | Tolerated failures
-------------|--------|-------------------
     1       |   1    |        0
     3       |   2    |        1
     5       |   3    |        2
     7       |   4    |        3

Formula: Quorum = floor(n/2) + 1

Example with 3 nodes:

  • ✅ 3 nodes active: Cluster healthy
  • ✅ 2 nodes active: Cluster works (quorum met)
  • ❌ 1 active node: Cluster stops (no quorum)_

_Recommendation: Always use an odd number of nodes (3, 5, 7) to optimize faults tolerance.

4. Leader Election in Patroni

Leader Lock mechanism

Patroni uses DCS to implement distributed lock:

Leader Lock Properties:

Key: /service/postgres/leader
Value: 
  {
    "version": "3.0.2",
    "conn_url": "postgres://node1:5432/postgres",
    "api_url": "http://node1:8008/patroni",
    "xlog_location": 123456789,
    "timeline": 2
  }
TTL: 30 seconds

Leader Election Process

Step 1: Race Condition

Time: T0 - Leader crashes
Node1: Check DCS → No leader key exists
Node2: Check DCS → No leader key exists  
Node3: Check DCS → No leader key exists

Step 2: Acquire Lock Attempt_

Time: T0 + 100ms
Node1: Try acquire lock → SUCCESS (first to write)
Node2: Try acquire lock → FAILED (key exists)
Node3: Try acquire lock → FAILED (key exists)

Step 3: Role Assignment_

Node1: Promote PostgreSQL to Primary
Node2: Configure as Replica, point to Node1
Node3: Configure as Replica, point to Node1

Step 4: Maintenance

Every 10 seconds:
Node1 (Leader): 
  - Renew lock (TTL extension)
  - Update xlog_location
  - Send heartbeat

Node2/3 (Followers):

  • Monitor leader key
  • Check replication lag
  • Ready to take over

Best selection criteria Replica_

When failover, Patroni chooses replica based on:

  1. Replication state:
    • streaming > in archive recovery
  2. Timeline: Higher Timeline takes priority_
  3. XLog position:
    • Replica has LSN closest to primary
    • Less data loss most
  4. No replication lag:
    • pg_stat_replication.replay_lag = 0
  5. Explicit candidate: Set in configuration

Priority tag:

tags:
nofailover: false
noloadbalance: false
clonefrom: false
nosync: false

Wallet example:

Primary fails at LSN: 0/3000000

Replica1: LSN=0/3000000, lag=0s ← BEST CHOICE Replica2: LSN=0/2FFFFFF, lag=1s Replica3: LSN=0/2FFFFFE, lag=2s

→ Patroni promotes Replica1

5. Failover Mechanism

Automatic Failover Process

Timeline details details:

Automatic Failover Process_

Detailed failover steps details

Step 1: Detect failure

# Patroni health check loop
while True:
if not check_postgresql_health():
log.error("PostgreSQL unhealthy")
stop_renewing_leader_lock()

if not check_dcs_connectivity():
    log.error("Lost connection to DCS")
    demote_if_leader()

sleep(10)

Step 2: Leader lock expires

# In etcd
$ etcdctl get /service/postgres/leader

After TTL: Key not found

Patroni logs on former leader

WARN: Could not renew leader lock INFO: Demoting PostgreSQL to standby

Step 3: Replica promotion_

# Patroni on promoted replica
INFO: No leader found
INFO: Attempting to acquire leader lock
INFO: Lock acquired successfully
INFO: Promoting PostgreSQL instance
INFO: Updating configuration
INFO: Notifying other members

Step 4: Reconfiguration

-- On promoted replica
SELECT pg_promote();

-- Changes primary_conninfo to null -- Restarts as read-write

Step 5: Followers repoint_

# Other replicas
INFO: New leader detected: node2
INFO: Updating primary_conninfo
INFO: Restarting replication

Monitoring Failover_

Important Metrics_:

  • patroni_primary_timeline: Detect timeline changes
  • patroni_xlog_location: Track WAL position
  • patroni_replication_lag: Lag before failover
  • patroni_failover_count: Count the number of times failover

6. Split-Brain Problem_

What is Split-Brain?

Definition_: Situation where ≥2 nodes think they are Primary, recording different data → Data divergence.

Cause

Network Partition
  1. Network Partition
  2. DCS partition: etcd cluster split_
  3. Slow network: Heartbeat timeout but node still live

Consequences of Split-Brain

Consequences of Split-Brain

Patroni's Split-Brain Prevention

Mechanism 1: DCS-based Lock (Primary)

def maintain_leader_lock():
while is_leader:
# Must renew within TTL
success = dcs.renew_lock(ttl=30)

    if not success:
        log.critical("Lost leader lock!")
        # Immediate demotion
        demote_to_standby()
        stop_accepting_writes()
        break
    
    sleep(10)

Mechanism 2: Leader Key Verification_

def before_handle_write():
leader_key = dcs.get("/service/postgres/leader")

if leader_key.owner != my_node_name:
    # I'm not the real leader!
    raise Exception("Not leader anymore")
    demote_immediately()

Mechanism 3: Timeline Divergence Detection

-- PostgreSQL timeline
SELECT timeline_id FROM pg_control_checkpoint();

-- If timelines diverge: -- Node1: timeline=5 -- Node2: timeline=6 -- → Data inconsistency detected -- → Requires pg_rewind or rebuild

Quorum requirement

etcd with 3 nodes:

Scenario 1: Network partition 1-2 split
Partition A: Node1 (1 node)
- Cannot get quorum (1 < 2)
- Cannot write to etcd
- Demotes to standby ✓

Partition B: Node2, Node3 (2 nodes) - Has quorum (2 ≥ 2) - Can elect leader - Node2 becomes primary ✓

Result: Only 1 primary exists ✓

Scenario 2: Complete isolation_

Node1: Isolated, loses DCS

  • Tries to renew lock → FAIL
  • Demotes PostgreSQL immediately
  • Stops accepting connections

Node2/3: See Node1 gone

  • Elect new leader
  • Only 1 primary in cluster ✓

Watchdog Timer (Advanced Protection)

Hardware watchdog:

# patroni.yml
watchdog:
mode: required  # or automatic, off
device: /dev/watchdog
safety_margin: 5

Active dynamic:

  1. Patroni kicks watchdog device every 10s_
  2. If Patroni hangs or loses DCS → stops kicking
  3. After timeout → Watchdog reboots entire node
  4. Prevents "zombie primary" scenario

Best Practices to avoid Split-Brain

  1. Deploy separate DCS: etcd cluster in different AZ_
  2. Monitor DCS health: Alert when etcd is not healthy
  3. Network redundancy: Multiple network paths between nodes
  4. Proper timeouts_:
patroni:
ttl: 30              # Leader lock TTL
loop_wait: 10        # Check interval
retry_timeout: 10    # DCS operation timeout
  1. Enable watchdog: Hardware protection layer
  2. Monitoring:
# Check for timeline divergence
patronictl list

Expected: All nodes same timeline

  • Cluster: postgres (7001234567890123456) ----+----+-----------+ | Member | Host | Role | State | TL | Lag in MB | +--------+--------------+---------+---------+----+-----------+ | node1 | 10.0.1.1:5432| Leader | running | 5 | | | node2 | 10.0.1.2:5432| Replica | running | 5 | 0 | | node3 | 10.0.1.3:5432| Replica | running | 5 | 0 | +--------+--------------+---------+---------+----+-----------+

Recovery from Split-Brain

If split-brain occurs:

Step 1: Identify

# Check timeline
patronictl list

node1: timeline=5

node2: timeline=6 ← DIVERGED!

Step 2: Choose primary_

  • Select the node with important data more
  • Or node with higher timeline

Step 3: Rebuild diverged replica

# Option 1: pg_rewind (if safe)
patronictl reinit postgres node2

Option 2: Full rebuild

patronictl remove postgres node2

Then: reinitialize from scratch

Step 4: Verify_

patronictl list

All nodes same timeline ✓

7. Summary

Key Takeaways

✅ Patroni: HA template automates PostgreSQL management cluster

✅ DCS (etcd): Distributed coordination, store configuration and leader lock_

✅ Raft consensus: Ensure consistency and leader election in etcd

✅ Leader election: Automatic, fast (~30-40s), based on TTL locks

✅ Failover: Automatically promote the best replica when primary fails

✅ Split-brain prevention: DCS quorum + TTL locks + watchdog_

General architecture

General architecture case_

Review Questions_

  1. How is Patroni different from pure Streaming Replication?
  2. Why do you need DCS? Can't use a database to store state?
  3. What is the Quorum in a cluster of 5 nodes?
  4. Patroni chooses which replica to promote when failover?
  5. Split-brain happens and how does Patroni prevent it? What does HTMLTAG_740_
  6. Timeline in PostgreSQL mean?
  7. What does TTL 30 seconds mean? Why not set TTL = 5 seconds?

Preparing for the next lesson_

Lesson 4 will guide you on preparing the infrastructure:

  • Setup 3 VMs/Servers
  • Network configuration, firewall
  • SSH keys, time sync_
  • Necessary Dependencies__HTMLTAG_758___