目的_
このレッスンを終えると、次のことができるようになります:
- Patroni のエラー検出メカニズムを理解する
- リーダー選出プロセスを理解する
- フェールオーバー タイムラインを追跡する詳細_
- 多くのシナリオで自動フェイルオーバーをテスト
- フェイルオーバーの問題のトラブルシューティング
- フェイルオーバー速度の最適化
1。自動フェイルオーバーの概要
1.1。フェイルオーバーとは何ですか?
自動フェイルオーバー = プロセス 自動 プライマリが表示されたときにレプリカをプライマリに昇格させます失敗.
特別なポイント:
- ⚡ 自動: 介入は必要ありません手動
- 🚨 計画外: 試行
- ⏱️ 高速: 30 ~ 60 秒(構成可能)
- 🎯 目標: ダウンタイムを最小限に抑える
いつ発生するかフェイルオーバーしますか?
- プライマリ サーバーがクラッシュ
- PostgreSQL プロセスが停止_
- ネットワーク パーティション
- ハードウェア失敗
- DCS 接続が失われました
- ディスクがいっぱい
1.2。フェイルオーバーとレプリケーション
WITHOUT Patroni (Manual Failover):
- Primary fails
- DBA gets paged
- DBA investigates (10-30 mins)
- DBA manually promotes replica
- DBA updates application config
- Service restored Total downtime: 30+ minutes ❌
WITH Patroni (Automatic Failover):
Primary fails
Patroni detects (10 seconds)
Patroni promotes best replica (20 seconds)
Service restored automatically Total downtime: 30-60 seconds ✅
2。障害検出メカニズム
2.1。ヘルス チェック ループ
Patroni ヘルス チェック コンポーネント:
# Pseudo-code of Patroni's main loop while True:1. Check PostgreSQL health
if not check_postgresql_running(): log.error("PostgreSQL is down!") handle_postgres_failure()
2. Check DCS connectivity
if not can_connect_to_dcs(): log.error("Lost DCS connection!") demote_if_leader()
3. Update status in DCS
update_member_status_in_dcs()
4. Check leader lock (if I'm leader)
if is_leader: renew_leader_lock()
5. Sleep until next check
sleep(loop_wait) # Default: 10 seconds
2.2。 PostgreSQL ヘルスチェック_
Patroni は複数のチェックを実行:
A。プロセスチェック
# Check if postgres process exists ps aux | grep postgresCheck if accepting connections
pg_isready -h localhost -p 5432
B。接続チェック
# Try to connect to PostgreSQL
try:
conn = psycopg2.connect("host=localhost port=5432 dbname=postgres")
conn.close()
except:
# Connection failed!
mark_unhealthy()
C。レプリケーション チェック (レプリカ上)
-- Check if replication is active SELECT status, received_lsn, replay_lsn FROM pg_stat_wal_receiver;
-- If no data or status != 'streaming' → Problem!
D。タイムラインチェック_
-- Ensure timeline matches cluster
SELECT timeline_id FROM pg_control_checkpoint();
2.3。 DCS 接続チェック
DCS 接続が重要な理由:
If node loses DCS connection:
- Cannot renew leader lock
- Cannot read cluster state
- MUST demote to avoid split-brain
Even if PostgreSQL is healthy!
DCS チェック例:_
# Check etcd health etcdctl endpoint healthTry to read/write
etcdctl get /service/postgres/leader etcdctl put /service/postgres/members/node1 "healthy"
2.4。リーダー ロック TTL
TTL (存続時間)メカニズム_:
# In patroni.yml
bootstrap:
dcs:
ttl: 30 # Leader lock expires after 30 seconds
loop_wait: 10 # Check every 10 seconds
タイムライン:_
T+0s: Leader acquires lock (TTL=30s) T+10s: Leader renews lock (TTL extended to T+40s) T+20s: Leader renews lock (TTL extended to T+50s) T+30s: Leader tries to renew but FAILS (crashed) T+40s: Lock expires in DCS T+41s: Replicas detect no leader T+42s: Replica election begins T+45s: New leader elected
Total detection time: ~35-40 seconds
3。リーダー選出プロセス
3.1。選挙トリガー_
リーダー選挙は:
Condition 1: Leader lock expired in DCS /service/postgres/leader → key not foundCondition 2: No active leader for > loop_wait All replicas see: no leader heartbeat
Condition 3: Explicit failover patronictl failover command
3.2のときに開始されます。候補者の選択基準
パトローニが選択するのはst レプリカに基づく:
優先度 1: レプリケーション状態
-- Prefer streaming over archive recovery SELECT state FROM pg_stat_wal_receiver;
streaming > in archive recovery > stopped
優先度 2: レプリケーションの遅延
-- Replica with lowest lag wins SELECT pg_wal_lsn_diff(pg_last_wal_receive_lsn(), pg_last_wal_replay_lsn()) AS lag_bytes;
-- Example: -- node2: lag = 0 bytes ← BEST -- node3: lag = 1048576 bytes (1MB)
優先度 3:タイムライン
-- Higher timeline = more recent SELECT timeline_id FROM pg_control_checkpoint();
-- node2: timeline = 3 ← BEST -- node3: timeline = 2
優先度 4: タグ_
# In patroni.yml
tags:
nofailover: false # true = never promote this node
noloadbalance: false
priority: 100 # Higher = preferred (0-999)
例:
CODEBLOCK_15__優先度 5: 同期状態_
-- Synchronous replica preferred over async SELECT sync_state FROM pg_stat_replication;
sync > potential > async
3.3。競合状態とロックの取得
複数のレプリカが競合:
Scenario: Primary fails, 2 replicas competeT+0s: node2 and node3 both detect no leader T+0.1s: Both try to acquire lock simultaneously
In etcd (atomic operation): node2 tries: PUT /service/postgres/leader "node2" if_not_exists node3 tries: PUT /service/postgres/leader "node3" if_not_exists
Result: Only ONE succeeds (etcd atomic guarantee) node2: SUCCESS → becomes leader node3: FAILED → remains replica
DCS保証:
- 原子性: 1 つのノードのみがロックを取得
- 一貫性: すべてのノードが同じように見えるリーダー
- 孤立: スプリットブレインは不可能
3.4。プロモーション プロセス_
勝者ノードが実行:
Step 1: Acquire leader lock in DCS etcdctl put /service/postgres/leader '{"node": "node2", ...}'Step 2: Run pre_promote callback (if configured) /var/lib/postgresql/callbacks/pre_promote.sh
Step 3: Promote PostgreSQL Method A: pg_ctl promote -D /var/lib/postgresql/18/data Method B: SELECT pg_promote(); Method C: Create trigger file (old method)
Step 4: Wait for promotion complete Check: SELECT pg_is_in_recovery(); Should return: false (not in recovery = primary)
Step 5: Update timeline Timeline increments: 1 → 2
Step 6: Run post_promote callback Update DNS, load balancer, send notifications
Step 7: Run on_role_change callback /var/lib/postgresql/callbacks/on_role_change.sh master
Step 8: Update DCS with new primary info xlog_location, timeline, conn_url
Step 9: Start accepting writes PostgreSQL now in read-write mode
4。フェイルオーバー タイムラインの詳細
4.1。完全なフェイルオーバー フロー
Timeline of Automatic FailoverT+0s: NORMAL OPERATION Primary (node1): Healthy, serving requests Replica (node2): Streaming from node1, lag=0 Replica (node3): Streaming from node1, lag=0
T+1s: PRIMARY FAILS node1: PostgreSQL crashes / server dies node2: Still streaming (buffered data) node3: Still streaming (buffered data)
T+5s: REPLICATION BROKEN node2: WAL receiver error "connection lost" node3: WAL receiver error "connection lost" node1: Still holds leader lock (TTL not expired yet)
T+10s: HEALTH CHECK CYCLE 1 node2: Check replication → FAILED, wait... node3: Check replication → FAILED, wait... node1: Cannot renew lock (crashed)
T+20s: HEALTH CHECK CYCLE 2 node2: Still cannot connect to node1 node3: Still cannot connect to node1
T+30s: LEADER LOCK EXPIRES DCS: /service/postgres/leader TTL expired → key deleted node2: Detects no leader key node3: Detects no leader key
T+31s: CANDIDATE ELECTION BEGINS node2: Check eligibility → YES (lag=0, priority=100) node3: Check eligibility → YES (lag=1MB, priority=100)
T+32s: RACE FOR LOCK node2: PUT /service/postgres/leader "node2" → SUCCESS node3: PUT /service/postgres/leader "node3" → FAILED
T+33s: NODE2 PROMOTES node2: Run pre_promote callback node2: pg_promote() executed node2: Timeline: 1 → 2
T+35s: PROMOTION COMPLETE node2: pg_is_in_recovery() → false node2: Now accepting writes node2: Run post_promote & on_role_change callbacks
T+36s: NODE3 RECONFIGURES node3: Detects new leader = node2 node3: Update primary_conninfo → node2:5432 node3: Restart WAL receiver
T+38s: REPLICATION RESTORED node3: Connected to node2 node3: Streaming at timeline 2
T+40s: CLUSTER OPERATIONAL Primary: node2 (was replica) Replica: node3 (following node2) Failed: node1 (needs manual intervention)
Total Downtime: ~35-40 seconds ✅
4.2。フェイルオーバー速度に影響する要因
構成パラメータ:HTMLTAG_246__CODEBLOCK_20
トレードオフ:
| パラメータ_ | _下限値 | 上限値_ |
|---|---|---|
| _TTL_ | 高速フェイルオーバー | 詳細安定 |
| 誤検知が増加_ | 遅いフェイルオーバー_ | |
| _loop_wait | _高速検出_ | DCS の削減トラフィック_ |
| CPU/ネットワークの増加_ | 反応が遅い |
_一般的な構成:
# Conservative (stable, slower) ttl: 30 loop_wait: 10 → Failover: ~40-50sBalanced (recommended)
ttl: 20 loop_wait: 10 → Failover: ~30-40s
Aggressive (fast, sensitive)
ttl: 15 loop_wait: 5 → Failover: ~20-30s
5。自動フェイルオーバーのテスト
5.1。テスト シナリオ 1: PostgreSQL プロセスの強制終了
PostgreSQL のクラッシュをシミュレート:_
# On current primary (node1) sudo -u postgres psql -c "SELECT pg_backend_pid();"Returns: 12345
sudo kill -9 12345 # Kill PostgreSQL
Or kill all postgres processes
sudo pkill -9 postgres
_Monitorフェイルオーバー:
# Terminal 1: Watch cluster status watch -n 1 "patronictl list postgres"Terminal 2: Monitor logs
sudo journalctl -u patroni -f
Terminal 3: Test connectivity
while true; do psql -h 10.0.1.11 -U app_user -d myapp -c "SELECT 1;" 2>&1 | grep -q "ERROR" && echo "$(date): DOWN" || echo "$(date): UP" sleep 1 done
予想されるタイムライン:
00:00 - Cluster healthy
00:01 - Kill postgres on node1
00:02-00:30 - Patroni detecting failure
00:31 - node2 elected as new primary
00:35 - Cluster operational (node2 = primary)
00:36+ - Connections working again
5.2。テスト シナリオ 2: ネットワーク パーティション
ネットワーク パーティションのシミュレーション:
# On primary node, block traffic to other nodes sudo iptables -A INPUT -s 10.0.1.12 -j DROP sudo iptables -A INPUT -s 10.0.1.13 -j DROP sudo iptables -A OUTPUT -d 10.0.1.12 -j DROP sudo iptables -A OUTPUT -d 10.0.1.13 -j DROPOr block etcd access specifically
sudo iptables -A OUTPUT -p tcp --dport 2379 -j DROP
観察:
_CODEBLOCK_26 _リカバリ:
# Restore network on node1 sudo iptables -Fnode1 should automatically rejoin as replica
patronictl list postgres
5.3。テスト シナリオ 3: サーバーの再起動
サーバーのクラッシュをシミュレート:
# On primary node sudo rebootOr immediate crash
echo c | sudo tee /proc/sysrq-trigger
予想される動作: シナリオ 1 と同じですが、完全にノード化されます。利用できません。
5.4。テスト シナリオ 4: ディスク フル
ディスク フルをシミュレート:
# Fill up disk on primary dd if=/dev/zero of=/var/lib/postgresql/bigfile bs=1M count=10000PostgreSQL will fail when cannot write WAL
Patroni が検出 PostgreSQL の異常 → トリガーフェイルオーバー。
5.5。テスト シナリオ 5: DCS の失敗
すべてのノードで etcd を停止:
# On all 3 etcd nodes
sudo systemctl stop etcd
予想通り動作_:
- All Patroni nodes lose DCS connection
- Current primary DEMOTES (safety mechanism)
- Cluster enters "read-only" state
- NO failover possible (no DCS consensus)
Recovery:
- Restart etcd cluster
- Patroni auto-recovers
Leader election happens
6。フェイルオーバーの成功を確認
6.1。クラスターのステータス
# List cluster members patronictl list postgresExpected after failover:
+ Cluster: postgres (7001234567890123456) ----+----+-----------+
| Member | Host | Role | State | TL | Lag in MB |
+--------+---------------+---------+---------+----+-----------+
| node1 | 10.0.1.11:5432| Replica | stopped | 1 | | ← Old primary
| node2 | 10.0.1.12:5432| Leader | running | 2 | | ← NEW primary
| node3 | 10.0.1.13:5432| Replica | running | 2 | 0 |
+--------+---------------+---------+---------+----+-----------+
Note timeline changed: 1 → 2
6.2 を確認します。新しいプライマリ
# Check primary role sudo -u postgres psql -h 10.0.1.12 -c "SELECT pg_is_in_recovery();"pg_is_in_recovery
------------------
f ← false = PRIMARY
Check timeline
sudo -u postgres psql -h 10.0.1.12 -c "SELECT timeline_id FROM pg_control_checkpoint();"
timeline_id
------------
2
Check replication from new primary
sudo -u postgres psql -h 10.0.1.12 -c "SELECT * FROM pg_stat_replication;"
Should show node3 replicating from node2
6.3 を確認します。書き込み操作をテストします_
# Insert data on new primary sudo -u postgres psql -h 10.0.1.12 -d testdb -c " INSERT INTO test_table (data) VALUES ('After failover at ' || NOW()); "Verify on replica
sudo -u postgres psql -h 10.0.1.13 -d testdb -c " SELECT * FROM test_table ORDER BY id DESC LIMIT 5; "
Should see new data replicated
6.4。フェールオーバー履歴
# View history via REST API curl -s http://10.0.1.12:8008/history | jqOutput:
[
[1, 67108864, "no recovery target specified", "2024-11-25T10:00:00+00:00"],
[2, 134217728, "no recovery target specified", "2024-11-25T11:30:15+00:00"]
]
↑ Timeline 2 = Failover event
Check Patroni logs
sudo journalctl -u patroni --since "30 minutes ago" | grep -i "promote|failover|leader"
7 を確認します。フェイルオーバーの問題のトラブルシューティング
7.1。問題: フェイルオーバーが発生しない
症状: プライマリがダウンしているが昇格なし。
考えられる原因:
A。すべてのレプリカは nofailover
# Check tags patronictl show-config postgres | grep -A5 "tags:"If all replicas have nofailover: true
Solution: Remove tag from at least one replica
patronictl edit-config postgres
Set: nofailover: false
B とタグ付けされています。レプリケーションの遅延が大きすぎます
# Check maximum_lag_on_failover patronictl show-config postgres | grep maximum_lag_on_failoverIf replica lag > threshold, won't promote
Solution: Increase threshold or wait for lag to decrease
patronictl edit-config postgres
Set: maximum_lag_on_failover: 10485760 # 10MB
C。 DCS_
# Check etcd health etcdctl endpoint health --clusterIf etcd cluster has no quorum (< 2 of 3 healthy)
Solution: Fix etcd cluster first
sudo systemctl restart etcd
D にクォーラムがありません。 synchronous_mode_strict が有効
# If enabled and no sync replica available synchronous_mode: true synchronous_mode_strict: true # ← Problem!Primary cannot be demoted, replicas cannot be promoted
Solution: Disable strict mode
patronictl edit-config postgres
Set: synchronous_mode_strict: false
7.2。問題: 複数のフェイルオーバー (フラッピング)
症状: クラスターが繰り返しフェイルオーバーを繰り返します。
考えられる原因:
A。ネットワークが不安定
# Check network between nodes ping -c 100 10.0.1.12High packet loss → false failovers
Solution: Fix network or increase TTL
patronictl edit-config postgres
Set: ttl: 40 # More tolerant
B。 TTL が攻撃的すぎます_
# ttl: 10 ← Too low!Every small network blip causes failover
Solution: Increase TTL
ttl: 30 # More stable
C。リソースの枯渇_
# Check CPU/Memory on nodes top free -hIf resources exhausted, health checks timeout
Solution: Scale up resources or reduce load
7.3。問題: フェールオーバーが遅い
症状: フェールオーバーに 60 秒以上かかります。
診断:
_CODEBLOCK_43 _最適化:
# Reduce TTL and loop_wait bootstrap: dcs: ttl: 20 # Was 30 loop_wait: 5 # Was 10Expected failover: ~30-35 seconds
7.4。問題: フェイルオーバー後のデータ損失
症状: 最近のトランザクションがいくつか欠落しています。
原因: 非同期レプリケーション + レプリケーションラグ。
検証:
___CODEBLOCK_45 ___予防:
# Enable synchronous replication bootstrap: dcs: synchronous_mode: true synchronous_mode_strict: false # Allow degradationpostgresql: parameters: synchronous_commit: 'on'
8。メトリクスとモニタリング
8.1。主要なフェールオーバー メトリック_
-- Time since last failover SELECT timeline_id, pg_postmaster_start_time(), now() - pg_postmaster_start_time() AS uptime FROM pg_control_checkpoint();-- Replication lag (pre-failover indicator) SELECT application_name, pg_wal_lsn_diff(pg_current_wal_lsn(), replay_lsn) AS lag_bytes, replay_lag FROM pg_stat_replication;
-- Failed connection attempts (indicator of downtime) SELECT datname, numbackends, xact_commit, xact_rollback FROM pg_stat_database;
8.2。アラート ルール
Prometheus アラートの例:
groups:
name: patroni_failover rules:-
alert: PatroniFailoverDetected expr: increase(patroni_timeline[5m]) > 0 labels: severity: warning annotations: summary: "Patroni failover detected" description: "Timeline changed, indicating failover"
-
alert: PatroniNoLeader expr: count(patroni_patroni_info{role="master"}) == 0 for: 30s labels: severity: critical annotations: summary: "No Patroni leader" description: "Cluster has no primary"
alert: PatroniHighReplicationLag expr: patroni_replication_lag_bytes > 10485760 # 10MB for: 2m labels: severity: warning annotations: summary: "High replication lag" description: "Replica lag > 10MB, risk of data loss on failover"
-
9。ベスト プラクティス
✅ DO
- フェイルオーバーを定期的にテストする - ステージングでは毎月、運用環境では四半期に一度
- レプリケーションを監視するlag - ラグが > の場合に警告します。 1MB
- 同期レプリケーションを使用 データ損失ゼロ_
- synchronous_mode_strict を設定: false - 許可劣化_
- 適切な TTL を構成_ - 速度と安定性のバランスを取る (20 ~ 30 秒)
- レプリカが 2 つ以上ある - レプリカが 1 つであってもフェイルオーバーを許可するレプリカのダウン
- DCS の正常性を監視 - etcd クラスターが正常である必要があります
- ランブックを文書化 - 手動の手順介入
- フェイルオーバー イベントのログ - パターンと問題の追跡
- 容量計画 - レプリカはプライマリを処理する必要がありますロード_
❌ しないでください
- 単一レプリカを使用しない - フェイルオーバーオプションなし
- 無視しないでくださいラグ - ラグが大きい = データ損失のリスク
- TTL を低く設定しすぎないでください (<15 秒) - 誤検知
- スキップしないでくださいtesting_ - テストされていないフェイルオーバー = ダウンタイムのリスク
- 自動フェイルオーバー中は を手動で昇格させないでください - Patroni に処理させます
- 古いものを忘れないでくださいプライマリ_ - 再結合/再構築が必要
- 監視せずに実行しない - フェイルオーバーがいつ発生するかを把握する必要がある
- 過負荷にしないでくださいDCS_ - 個別の etcd クラスターを推奨_
10。ラボ演習
ラボ 1: 基本フェイルオーバー テスト
タスク: 1. ベースラインを記録します: patronictl list 2.プライマリを停止します: sudo systemctl stop patroni 3. watch -n 1 patronictl list 4 を使用してフェイルオーバーの時間を計測します。ダウンタイムの期間を文書化する 5. 新しいプライマリが書き込みを受け入れることを確認する 6. 古いプライマリを再起動し、再参加を確認する_
ラボ 2: ネットワーク パーティション テスト
タスク: 1. iptables を使用してクラスター パーティションからプライマリに接続する2. DCS の動作を観察します。 3. パーティションの後にプライマリが 1 つだけ存在することを確認します。 4. ネットワークを復元し、自動回復を確認します
ラボ 3: フェールオーバー速度の最適化_
タスク: 1. デフォルト設定 (TTL=30) でのベースライン テスト 2. 削減TTL を 20 に、再度テストします。 3. 15 に減らし、再度テストします。 4. フェイルオーバー時間を比較します。 5. トレードオフを評価します (速度と誤検知)
ラボ 4: 負荷時のフェイルオーバー
タスク: 1. 次のコマンドで負荷を生成します。 pgbench: pgbench -c 10 -T 300 2.ロード中にプライマリを停止します。 3. pgbench 出力で接続エラーをカウントします。 4. 可用性の割合を計算します。 5. ユーザーへの影響を文書化します
11。概要
重要な概念
✅ 自動フェイルオーバー = 手動なしの自己修復介入_
✅ 検出 = ヘルスチェック + DCS 接続 + TTL有効期限
✅ 選出 = ラグ、タイムライン、タグに基づく最適なレプリカ
✅ プロモーション = pg_promote() + タイムラインの増分 + ロールの変更
✅ Timeline = フェイルオーバーカウンター、防止発散
✅ TTL = 速度と安定性の間のトレードオフ
フェイルオーバーチェックリスト
- 主な失敗検出
- DCS でリーダー ロックの有効期限が切れました
- 最良のレプリカが特定されました
- リーダー ロックを取得
- PostgreSQL が昇格されました正常に
- タイムラインが増加しました_
- コールバックが実行されました_
- 他のレプリカが再構成されました_
- レプリケーション復元済み
- クラスタ運用可能
次のステップ
レッスン 14 で説明します 切り替え計画中計画済み:
- 計画されたメンテナンス シナリオ
- ゼロダウンタイムのスイッチオーバー プロセス_
- スムーズなスイッチオーバーと即時スイッチオーバー
- 計画されたメンテナンスのベスト プラクティスフェイルオーバー_