目標
このレッスンの後、次のことを学習します:
- スイッチオーバーとフェイルオーバーを区別する_
- 計画されたスイッチオーバーを完全に安全に実装する
- グレースフルと即時を理解するスイッチオーバー
- メンテナンス時のダウンタイムを最小限に抑える
- ローリングアップデートのためのスイッチオーバーを自動化
- 運用環境でのスイッチオーバーを処理
1。切り替えの概要
1.1。スイッチオーバーとは何ですか?
スイッチオーバー = 計画 レプリカをプライマリにプロモートします。
比較フェイルオーバー:
___HTMLTA G_151___| アスペクト | _フェイルオーバー | _スイッチオーバー___HTMLTAG_1 08___ |
|---|---|---|
| _トリガー | 主な障害 (計画外) | 手動/スケジュール済み(予定) |
| ダウンタイム | 30-60秒 | 0-10秒_ |
| データ損失 | 可能(非同期の場合) | ゼロ(制御) |
| コントロール___HTMLTAG_145_ __ | 自動_ | 手動/スクリプト |
| タイミング | 非発令可能_ | _予定 |
1.2.いつ切り替える必要がありますか?_
一般的なシナリオ:
A。ハードウェアのメンテナンス_
Scenario: Need to replace failing disk on primary server
→ Switchover to replica
→ Perform maintenance on old primary
→ Keep as replica or switchover back
B。ソフトウェアのアップグレード_
Scenario: OS kernel update requires reboot
→ Switchover to replica
→ Update & reboot old primary
→ Verify, then switchover back (optional)
C。データベースの移行_
Scenario: Move database to larger server
→ Add new server as replica
→ Switchover to new server
→ Remove old server
D。データセンターの移行
Scenario: Move from DC1 to DC2
→ Setup replicas in DC2
→ Switchover primary to DC2
→ Decommission DC1 nodes
E。
Scenario: Test HA readiness before production
→ Perform switchover in staging
→ Validate application behavior
→ Measure downtime
1.3をテストしています。切り替えのメリット
✅ データ損失ゼロ - 切り替え前にコミットされたすべてのトランザクション
✅ 制御済みタイミング_ - メンテナンス期間中
✅ リスクが低い - 調整されテストされたプロセス
✅ 最小限ダウンタイム - フェイルオーバーの場合は 0 ~ 10 秒、30 ~ 60 秒
✅ リバーシブル - 問題が発生した場合は元に戻すことができます
2。スイッチオーバーのタイプ
2.1。グレースフル スイッチオーバー (デフォルト)
プロセス:
___CODEBLOCK_5 ___コマンド:
patronictl switchover postgres
2.2。即時切り替え
プロセス:
___CODEBLOCK_7 ___コマンド:
patronictl switchover postgres --force
2.3。スケジュールされたスイッチオーバー
プロセス:
___CODEBLOCK_ 9___コマンド:
patronictl switchover postgres --scheduled 2024-11-25T02:00:00
3。切り替えの前提条件
3.1。クラスターのヘルスチェック
# 1. Verify all nodes running patronictl list postgresExpected:
+ Cluster: postgres (7001234567890123456) ----+----+-----------+
| Member | Host | Role | State | TL | Lag in MB |
+--------+---------------+---------+---------+----+-----------+
| node1 | 10.0.1.11:5432| Leader | running | 2 | |
| node2 | 10.0.1.12:5432| Replica | running | 2 | 0 | ✅
| node3 | 10.0.1.13:5432| Replica | running | 2 | 0 | ✅
+--------+---------------+---------+---------+----+-----------+
All nodes must be:
- State: running ✅
- Lag: 0 or very low ✅
- Same timeline ✅
3.2。レプリケーションラグチェック
# Check lag on all replicas sudo -u postgres psql -h 10.0.1.11 -c " SELECT application_name, client_addr, state, pg_wal_lsn_diff(pg_current_wal_lsn(), replay_lsn) AS lag_bytes, replay_lag FROM pg_stat_replication ORDER BY lag_bytes DESC; "Desired:
application_name | client_addr | state | lag_bytes | replay_lag
-----------------+-------------+-----------+-----------+------------
node2 | 10.0.1.12 | streaming | 0 | 00:00:00 ✅
node3 | 10.0.1.13 | streaming | 0 | 00:00:00 ✅
3.3。ターゲット候補チェック_
# Check if target has nofailover tag patronictl show-config postgres | grep -A10 "tags:"Target node should have:
tags: nofailover: false # ✅ Can be promoted priority: 100 # Higher = preferred
NOT:
tags: nofailover: true # ❌ Cannot be promoted
3.4。接続の可用性
# Test connection to target psql -h 10.0.1.12 -U postgres -c "SELECT 1;"Test application user
psql -h 10.0.1.12 -U app_user -d myapp -c "SELECT 1;"
4。スイッチオーバーの実行
4.1。対話型スイッチオーバー (推奨)
ステップバイステップ:
_ CODEBLOCK_16_# 1. Initiate switchover patronictl switchover postgresPatroni prompts:
出力:
2024-11-25 10:30:00.123 UTC [INFO]: Switching over from node1 to node2 2024-11-25 10:30:02.456 UTC [INFO]: Waiting for replica node2 to catch up... 2024-11-25 10:30:02.789 UTC [INFO]: Replica node2 lag: 0 bytes ✅ 2024-11-25 10:30:03.012 UTC [INFO]: Promoting node2... 2024-11-25 10:30:05.234 UTC [INFO]: node2 promoted successfully 2024-11-25 10:30:06.567 UTC [INFO]: Demoting node1... 2024-11-25 10:30:08.890 UTC [INFO]: node1 reconfigured as replica 2024-11-25 10:30:10.123 UTC [INFO]: Switchover completed ✅
Total time: 10 seconds
4.2。非対話型スイッチオーバー
直接コマンド:
# Specify master and candidate explicitly patronictl switchover postgres
--master node1
--candidate node2
--force--force: Skip confirmation prompt
4.3。スケジュールされたスイッチオーバー
メンテナンス期間のスケジュール:
# Schedule switchover at 2 AM patronictl switchover postgres
--master node1
--candidate node2
--scheduled "2024-11-25T02:00:00"Patroni will automatically execute at scheduled time
スケジュールを確認スイッチオーバー_:HTMLTAG_272__CODEBLOCK_20
スケジュールされたスイッチオーバーをキャンセル:HTMLTAG_276__CODEBLOCK_21
4.4。 REST API による切り替え_
API 経由のトリガー:
# POST to current leader curl -X POST http://10.0.1.11:8008/switchover
-H "Content-Type: application/json"
-d '{ "leader": "node1", "candidate": "node2" }'Response:
{
"status": "ok",
"message": "Switchover scheduled"
}
5。切り替えタイムライン_
5.1。詳細なフロー_
T+0s: INITIATE SWITCHOVER Command: patronictl switchover postgres --master node1 --candidate node2T+0.5s: PRE-CHECKS ✓ node1 is current leader ✓ node2 is healthy replica ✓ node2 replication lag: 0 bytes ✓ node2 timeline matches: 2
T+1s: PREPARE OLD PRIMARY (node1)
- Checkpoint: CHECKPOINT;
- Flush WAL
- Set session_replication_role = 'replica' (prevent writes soon)
T+2s: WAIT FOR LAG = 0
- Monitor: pg_stat_replication.replay_lag
- node2 lag: 0 bytes ✅
- All WAL replayed
T+3s: PAUSE OLD PRIMARY
- Set: pg_catalog.pg_pause_wal_replay() on replicas (not needed, they're already replaying)
- Actually: Just ensure all WAL consumed
T+4s: DEMOTE OLD PRIMARY (node1)
- Remove leader lock from DCS
- Stop accepting new connections (pg_ctl reload with max_connections=0)
- Wait for active transactions (timeout: 30s default)
T+5s: PROMOTE NEW PRIMARY (node2)
- Acquire leader lock in DCS
- Execute: SELECT pg_promote();
- Timeline: 2 → 3
- Run callbacks: on_role_change, post_promote
T+7s: VERIFY NEW PRIMARY
- pg_is_in_recovery() → false ✅
- Accepting connections
- Timeline = 3
T+8s: RECONFIGURE OLD PRIMARY (node1)
- Update primary_conninfo → node2:5432
- Update recovery.signal
- Restart PostgreSQL in recovery mode
- Timeline: 2 → 3
T+10s: REPLICATION RESTORED
- node1 now streaming from node2
- node3 updated to stream from node2
- All replicas timeline = 3
T+10s: SWITCHOVER COMPLETE ✅ Primary: node2 (was replica) Replica: node1 (was primary) Replica: node3
Total downtime: ~5-10 seconds Data loss: None ✅
5.2。アクティブな接続はどうなりますか?_
切り替え中:
Client connections to old primary (node1):
Option A: Graceful (default)
- New connections: REJECTED
- Active queries: ALLOWED TO COMPLETE (timeout: 30s)
- Idle connections: TERMINATED after queries done
Option B: Force (--force)
- All connections: TERMINATED IMMEDIATELY
- Active queries: ROLLBACK
Faster but risky ⚠️
アプリケーションの動作:
# Well-written application with retry logic import psycopg2
def execute_query(): retries = 3 for i in range(retries): try: conn = psycopg2.connect("host=10.0.1.11 ...") cursor = conn.cursor() cursor.execute("SELECT * FROM users;") return cursor.fetchall() except psycopg2.OperationalError as e: if i < retries - 1: time.sleep(1) # Wait and retry continue raise
6。切り替え後の検証
6.1。クラスターのステータス_
patronictl list postgresExpected:
+ Cluster: postgres (7001234567890123456) ----+----+-----------+
| Member | Host | Role | State | TL | Lag in MB |
+--------+---------------+---------+---------+----+-----------+
| node1 | 10.0.1.11:5432| Replica | running | 3 | 0 | ← Was Leader
| node2 | 10.0.1.12:5432| Leader | running | 3 | | ← Was Replica
| node3 | 10.0.1.13:5432| Replica | running | 3 | 0 |
+--------+---------------+---------+---------+----+-----------+
Check:
✅ node2 is now Leader
✅ Timeline changed: 2 → 3
✅ All nodes running
✅ Replication lag = 0
6.2。レプリケーションステータス
# On new primary (node2) sudo -u postgres psql -h 10.0.1.12 -c " SELECT application_name, client_addr, state, sync_state FROM pg_stat_replication; "Expected:
application_name | client_addr | state | sync_state
-----------------+-------------+-----------+------------
node1 | 10.0.1.11 | streaming | async
node3 | 10.0.1.13 | streaming | async
Both replicas should be streaming from node2 ✅
6.3。テスト_
# Insert on new primary sudo -u postgres psql -h 10.0.1.12 -d testdb -c " INSERT INTO test_table (data, created_at) VALUES ('After switchover', NOW()) RETURNING *; "Verify on replicas
sudo -u postgres psql -h 10.0.1.11 -d testdb -c " SELECT * FROM test_table ORDER BY created_at DESC LIMIT 1; "
sudo -u postgres psql -h 10.0.1.13 -d testdb -c " SELECT * FROM test_table ORDER BY created_at DESC LIMIT 1; "
Should see the new row on both replicas ✅
6.4を書き込みます。タイムラインの確認
# Check timeline on all nodes for node in 10.0.1.11 10.0.1.12 10.0.1.13; do echo "=== $node ===" sudo -u postgres psql -h $node -c " SELECT timeline_id, pg_is_in_recovery() AS is_replica FROM pg_control_checkpoint(); " doneAll should report:
timeline_id | is_replica
------------+------------
3 | t/f
7。切り替えのベスト プラクティス_
7.1。切り替え前チェックリスト
#!/bin/bashpre-switchover-check.sh
echo "=== Pre-Switchover Checks ==="
1. Cluster health
echo "1. Checking cluster health..." patronictl list postgres | grep -q "running" || { echo "❌ Not all nodes running"; exit 1; } echo "✅ All nodes running"
2. Replication lag
echo "2. Checking replication lag..." lag=$(sudo -u postgres psql -h 10.0.1.11 -t -c " SELECT COALESCE(MAX(pg_wal_lsn_diff(pg_current_wal_lsn(), replay_lsn)), 0) FROM pg_stat_replication; ") if [ "$lag" -gt 1048576 ]; then # 1MB echo "❌ Lag too high: $lag bytes" exit 1 fi echo "✅ Lag acceptable: $lag bytes"
3. Target candidate available
echo "3. Checking target candidate..." patronictl list postgres | grep node2 | grep -q "running" || { echo "❌ node2 not available"; exit 1; } echo "✅ Target candidate available"
4. No scheduled maintenance
echo "4. Checking scheduled actions..." curl -s http://10.0.1.11:8008/patroni | jq -e '.scheduled_switchover == null' > /dev/null || { echo "⚠️ Another switchover already scheduled" }
echo "" echo "✅ All pre-checks passed. Safe to proceed."
7.2。ダウンタイムを最小限に抑える戦略_
A。接続プーラー
Use PgBouncer/HAProxy between app and database:App → PgBouncer → Primary ↓ Replicas
During switchover:
- PgBouncer detects primary change
- Reconnects to new primary automatically
Application sees minimal disruption
B。リードレプリカのルーティング
Route read queries to replicas during switchover:
- Write queries: Wait for new primary
- Read queries: Continue on replicas (may be slightly stale)
Result: Partial availability during switchover
C。アプリケーションレベルの再試行_
# Implement exponential backoff
def execute_with_retry(query, max_retries=3):
for i in range(max_retries):
try:
return execute_query(query)
except OperationalError:
if i == max_retries - 1:
raise
time.sleep(2 ** i) # 1s, 2s, 4s
7.3。コミュニケーション プラン_
切り替え前上:
T-24h: Announce maintenance window
- Email: ops@, dev@, stakeholders
- Slack: #incidents, #ops
- Status page: Update with scheduled maintenance
T-1h: Reminder notification
- Final checks
- Confirm go/no-go
T-5min: Begin maintenance
- Start switchover
Monitor dashboards
切り替え中:
- Real-time updates in ops channel
Monitor metrics (latency, error rate)
Have rollback plan ready
後切り替え_:_
- Verify all systems operational
Post-switchover validation
Update documentation
Send completion notification
8。スイッチオーバーのトラブルシューティング
8.1。問題: スイッチオーバー コマンドがハングします
症状: patronictl スイッチオーバー _
診断:HTMLTAG_346__CODEBLOCK_3 7
解決策:
# Option 1: Wait for lag to catch up (recommended)Option 2: Use --force to skip wait (risk data loss)
Option 3: Cancel and reschedule
Ctrl+C # Cancel current switchover attempt
8.2。問題: 候補者は資格がありません
症状: エラー「候補者は資格がありません」
診断:
CODEBLOCK_3 9解決策:
# Remove nofailover tag patronictl edit-config postgresEdit:
tags: nofailover: false # Change to false
Restart Patroni on node2
sudo systemctl restart patroni
8.3。問題: 古いプライマリが降格しない
症状: 切り替えが失敗し、古いプライマリが引き続きリーダーになります。
診断:__HTMLTAG_374__CODEBLOCK_41 __
解決策:
# Force demote via REST API curl -X POST http://10.0.1.11:8008/restartOr manually:
sudo -u postgres psql -h 10.0.1.11 -c " SELECT pg_terminate_backend(pid) FROM pg_stat_activity WHERE pid != pg_backend_pid(); "
sudo systemctl restart patroni
8.4。問題: スイッチオーバー後にレプリケーションが中断されました
症状: 古いプライマリが新しいプライマリからレプリケートされません。
診断:HTMLTAG_388__CODEBLOCK_4 3
解決策:
# A. Restart Patroni (usually auto-fixes) sudo systemctl restart patroniB. Manual reinit if needed
patronictl reinit postgres node1
Patroni will:
1. Stop PostgreSQL on node1
2. Remove data directory
3. pg_basebackup from node2
4. Start as replica
9。スイッチオーバーの自動化
9.1。スクリプトによるスイッチオーバー
#!/bin/bashautomated-switchover.sh
set -e
CLUSTER="postgres" OLD_PRIMARY="node1" NEW_PRIMARY="node2"
echo "=== Starting Automated Switchover ===" echo "From: $OLD_PRIMARY → To: $NEW_PRIMARY"
Pre-checks
echo "Running pre-checks..." ./pre-switchover-check.sh || exit 1
Perform switchover
echo "Executing switchover..." patronictl switchover $CLUSTER
--master $OLD_PRIMARY
--candidate $NEW_PRIMARY
--forceWait for completion
echo "Waiting for switchover to complete..." sleep 15
Post-checks
echo "Running post-checks..." new_leader=$(patronictl list $CLUSTER | grep Leader | awk '{print $2}') if [ "$new_leader" == "$NEW_PRIMARY" ]; then echo "✅ Switchover successful!" echo "New leader: $new_leader" else echo "❌ Switchover failed!" echo "Current leader: $new_leader" exit 1 fi
Verify replication
echo "Verifying replication..." patronictl list $CLUSTER
echo "=== Switchover Complete ==="
9.2。 Ansible プレイブック_
# switchover.yml
name: Perform Patroni switchover hosts: localhost gather_facts: no vars: cluster_name: postgres old_primary: node1 new_primary: node2
tasks:
-
name: Pre-check cluster health command: patronictl list {{ cluster_name }} register: cluster_status changed_when: false
-
name: Verify all nodes running assert: that: - "'running' in cluster_status.stdout" fail_msg: "Not all nodes are running"
-
name: Execute switchover command: > patronictl switchover {{ cluster_name }} --master {{ old_primary }} --candidate {{ new_primary }} --force register: switchover_result
-
name: Wait for switchover completion pause: seconds: 15
-
name: Verify new leader command: patronictl list {{ cluster_name }} register: final_status changed_when: false
-
name: Display result debug: msg: "{{ final_status.stdout_lines }}"
name: Verify leadership assert: that: - "'{{ new_primary }}' in final_status.stdout" - "'Leader' in final_status.stdout" fail_msg: "Switchover failed" success_msg: "Switchover successful"
-
実行:_
ansible-playbook switchover.yml
9.3。 CI/CD 統合
# .github/workflows/db-maintenance.yml name: Database Maintenance Switchoveron: schedule: - cron: '0 2 * * 0' # Every Sunday at 2 AM workflow_dispatch: # Manual trigger
jobs: switchover: runs-on: self-hosted steps: - name: Notify start run: | curl -X POST ${{ secrets.SLACK_WEBHOOK }}
-d '{"text": "Starting scheduled database switchover"}'- name: Pre-checks run: ./scripts/pre-switchover-check.sh - name: Execute switchover run: | patronictl switchover postgres \ --master node1 \ --candidate node2 \ --force - name: Verify run: ./scripts/post-switchover-verify.sh - name: Notify completion if: always() run: | curl -X POST ${{ secrets.SLACK_WEBHOOK }} \ -d '{"text": "Switchover completed: ${{ job.status }}"}'
10。スイッチオーバーによるローリング アップデート
10.1。更新戦略
シナリオ: PostgreSQL を 17 から更新 → 18.
手順:
1. Update replica node3 (least critical)
- Stop Patroni
- Upgrade PostgreSQL
- Start Patroni
- Verify replication
Update replica node2
- Stop Patroni
- Upgrade PostgreSQL
- Start Patroni
- Verify replication
Switchover to node2 (now updated)
- patronictl switchover --master node1 --candidate node2
Update old primary node1
- Stop Patroni
- Upgrade PostgreSQL
- Start Patroni (now replica)
- Verify replication
Optionally switchover back to node1
- patronictl switchover --master node2 --candidate node1
Result: Zero-downtime upgrade ✅
10.2。カーネル更新例_
#!/bin/bashrolling-kernel-update.sh
NODES=("node1" "node2" "node3") PRIMARY=$(patronictl list postgres | grep Leader | awk '{print $2}')
echo "Current primary: $PRIMARY"
Update replicas first
for node in "${NODES[@]}"; do if [ "$node" == "$PRIMARY" ]; then continue # Skip primary for now fi
echo "=== Updating $node ===" ssh $node 'sudo yum update -y kernel && sudo reboot'
echo "Waiting for $node to come back..." sleep 60
Wait for node to rejoin
until patronictl list postgres | grep $node | grep -q "running"; do echo "Waiting for $node..." sleep 10 done
echo "✅ $node updated and rejoined" done
Now switchover from primary
NEW_PRIMARY=${NODES[1]} # Pick a replica if [ "$NEW_PRIMARY" == "$PRIMARY" ]; then NEW_PRIMARY=${NODES[2]} fi
echo "=== Switching over from $PRIMARY to $NEW_PRIMARY ===" patronictl switchover postgres
--master $PRIMARY
--candidate $NEW_PRIMARY
--forcesleep 15
Update old primary
echo "=== Updating $PRIMARY ===" ssh $PRIMARY 'sudo yum update -y kernel && sudo reboot'
echo "Waiting for $PRIMARY to rejoin as replica..." sleep 60
until patronictl list postgres | grep $PRIMARY | grep -q "running"; do echo "Waiting for $PRIMARY..." sleep 10 done
echo "✅ All nodes updated!" patronictl list postgres
11。ラボ演習_
ラボ 1: 基本的な切り替え
タスク:
- 現在のプライマリを確認する:
patronictl list_ - スイッチオーバーの実行:
patronictl switchover postgres_ - 連続クエリ ループでダウンタイムを測定
- 新規確認トポロジ
- 文書の観察
ラボ 2: スケジュールされた切り替え
タスク:
- スケジュール2の切り替え今から数分_
- 待機期間中のログの監視
- 自動実行の観察
- スケジュールされたスイッチオーバーのキャンセル (繰り返しおよびテストキャンセル)
ラボ 3: 強制 vsグレースフル
タスク:
- 長時間実行クエリの作成:
SELECT pg_sleep(300); - 正常なスイッチオーバーを試行 (待機を観察)
- --force を使用してキャンセルして再試行_
- 動作とダウンタイム
ラボ 4: ローリングアップデートのシミュレーション
タスク:
- 3 ノードから開始クラスター
- 「更新」ノード 3 (再起動によるシミュレート)
- 「更新」ノード 2
- ノード 2 に切り替え_
- 「更新」ノード 1_
- すべてのノードを確認運用
ラボ 5: 負荷時のスイッチオーバー
タスク:
- 開始pgbench:
pgbench -c 10 -T 300 - ロード中にスイッチオーバーを実行
- pgbench 出力のエラーを分析____HTMLTAG_511__HTMLTAG_512___成功の計算レート
- 接続プーラー (PgBouncer) を使用したテスト
12。概要
主要な概念
✅ スイッチオーバー = 計画された制御された役割Change_
✅ Graceful = トランザクションを待機します (遅い、より安全)
✅ 即時 = 強制終了 (高速、リスクが高い)
✅ スケジュール済み = 特定の時間に自動化
✅ ダウンタイムゼロ =適切なアーキテクチャで実現可能
スイッチオーバーとフェイルオーバー
___HTMLT AG_553______HTML TAG_572___| _アスペクト_ | _スイッチオーバー | _フェイルオーバー |
|---|---|---|
| 計画中 | 予定 | 計画外___HTMLTAG_562 ___ |
| コントロール | 手動_ | _自動 |
| ダウンタイム | 0-10代 | 30-60代 |
| データ損失 | なし | 可能 |
| リバーシブル | はい | いいえ |
ベスト プラクティス
- ✅ 最初にステージングでテスト
- ✅ トラフィックの少ないウィンドウでスケジュールを設定_
- ✅ グレースフル モードを使用する(デフォルト)
- ✅ 切り替え前にラグ = 0 を確認
- ✅ プロセス中の監視
- ✅ ロールバック計画がある
- ✅ と通信する関係者_
- ✅ 文書化手順
次のステップ_
レッスン 15 ではリカバリの失敗について説明しますノード_:
- フェイルオーバー後に古いプライマリに再参加
- _pg_rewind の使用法とシナリオ_
- _pg_basebackup による完全な再構築
- _タイムライン相違解決_
- スプリットブレイン回復