Chuyển đến nội dung chính

LESSON 15: CEPH MONITORING, TUNING AND TROUBLESHOOTING

Prometheus metrics for Ceph, Grafana dashboards, performance tuning parameters, OSD tuning, scrubbing, recovery settings, and troubleshooting common issues.

🔒 DevSecOps — Lesson 15 LESSON 15: CEPH MONITORING, TUNING AND TROUBLESHOOTING

Deploy Microservices On-Premises with Kubernetes HA

Part 3: Distributed Storage — Rook-Ceph

xdev.asia

🎯 LESSON OBJECTIVE__HTMLTAG_68___

After completing this lesson, you will:

  • ✅ Setup Prometheus monitoring for Ceph
  • ✅ Import Grafana dashboards for Ceph
  • ✅ Tuning OSD performance parameters
  • ✅ Managing scrubbing and recovery
  • ✅ Troubleshoot HEALTH_WARN and common issues__HTMLTAG_81___

PART 1: PROMETHEUS METRICS

1.1. Ceph Prometheus Module

# Prometheus module đã enable ở CephCluster CRD:
# modules: [{name: prometheus, enabled: true}]

Verify metrics endpoint:

kubectl -n rook-ceph exec deploy/rook-ceph-tools --
ceph mgr services

{

"dashboard": "https://rook-ceph-mgr-dashboard:8443/",

"prometheus": "http://rook-ceph-mgr:9283/"

}

Test scrape metrics:

kubectl -n rook-ceph port-forward svc/rook-ceph-mgr 9283:9283 curl -s http://localhost:9283/metrics | head -20

# HELP ceph_health_status Cluster health status

# TYPE ceph_health_status gauge

ceph_health_status 0

1.2. ServiceMonitor (for Prometheus Operator)

# ceph-servicemonitor.yaml:
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
  name: rook-ceph-mgr
  namespace: rook-ceph
  labels:
    team: rook
spec:
  namespaceSelector:
    matchNames:
      - rook-ceph
  selector:
    matchLabels:
      app: rook-ceph-mgr
      rook_cluster: rook-ceph
  endpoints:
    - port: http-metrics
      path: /metrics
      interval: 15s

1.3. Key Metrics to Monitor

Metric Meaning__HTMLTAG_99___ Alert threshold
ceph_health_status 0=OK, 1=WARN, 2=ERR > 0 for 5m
ceph_osd_up OSD up status != 1
ceph_osd_in OSD in cluster != 1
ceph_cluster_total_used_raw_bytes Raw usage > 80% capacity__HTMLTAG_135___
ceph_osd_op_r_latency_sum Read latency__HTMLTAG_141___ > 50ms p99
ceph_osd_op_w_latency_sum Write latency > 100ms p99
ceph_pg_degraded Degraded PGs > 0 for 5m
ceph_pool_stored_raw Pool raw usage Near full

PART 2: GRAFANA DASHBOARDS__HTMLTAG_174___
# Ceph cung cấp built-in Grafana dashboards:
# Dashboard IDs cho Grafana:
# - Ceph Cluster Overview: 2842
# - Ceph OSD Performance: 5336
# - Ceph Pool Overview: 5342
# - Ceph RBD Overview: 7845

# Import qua Grafana UI:
# 1. Grafana → + → Import
# 2. Nhập Dashboard ID
# 3. Chọn Prometheus datasource
# 4. Import

# Hoặc tạo ConfigMap cho auto-import (Bài 33):

PART 3: PERFORMANCE TUNING

3.1. OSD Tuning

# Các OSD config quan trọng:
kubectl -n rook-ceph exec deploy/rook-ceph-tools -- bash

Xem current config:

ceph config dump | grep osd

Tuning cho SSD:

ceph config set osd osd_op_num_threads_per_shard 2 ceph config set osd osd_op_num_shards 8 ceph config set osd bluestore_min_alloc_size_ssd 4096 # 4K cho SSD ceph config set osd bluestore_cache_size_ssd 3221225472 # 3GB cache ceph config set osd osd_memory_target 4294967296 # 4GB per OSD

Scrub scheduling:

ceph config set osd osd_scrub_begin_hour 2 # Start 2 AM ceph config set osd osd_scrub_end_hour 6 # End 6 AM ceph config set osd osd_scrub_sleep 0.1 # Sleep between scrubs ceph config set osd osd_deep_scrub_interval 604800 # Deep scrub weekly

Recovery throttling:

ceph config set osd osd_recovery_max_active 3 # Max concurrent recoveries ceph config set osd osd_recovery_sleep 0.1 # Throttle recovery ceph config set osd osd_max_backfills 1 # Max backfill per OSD

3.2. Pool Tuning

# PG autoscaler:
ceph osd pool set replicapool pg_autoscale_mode on

PG count (nếu manual):

PG count formula: (Target_PGs_per_OSD × OSD_count) / size

(100 × 3) / 3 = 100 PGs

ceph osd pool set replicapool pg_num 128 # Power of 2

Compression:

ceph osd pool set replicapool compression_algorithm zstd ceph osd pool set replicapool compression_mode aggressive


PART 4: TROUBLESHOOTING

4.1. Common HEALTH_WARN

# 1. HEALTH_WARN: x pgs degraded
# → Một số PGs không đủ replicas
ceph health detail
ceph pg dump_stuck degraded
# Fix: Đợi recovery tự động, hoặc:
ceph pg repair <PG_ID>

2. HEALTH_WARN: OSD near full

ceph osd df

Fix: Thêm OSD mới, hoặc giảm data:

ceph osd reweight-by-utilization

3. HEALTH_WARN: clock skew detected

→ NTP không sync giữa MON nodes

Fix: Kiểm tra chrony/NTP sync trên tất cả nodes

4. HEALTH_WARN: slow requests

ceph daemon osd.0 dump_ops_in_flight

Fix: Check disk I/O, network latency, OSD config

5. OSD pod CrashLoopBackOff

kubectl -n rook-ceph logs rook-ceph-osd-0-xxxxx-xxxxx

Common causes:

- Disk permission issues

- Previous data on disk (wipefs -a)

- Insufficient memory

4.2. Replace Failed OSD

# Scenario: disk /dev/sdb trên worker2 bị hỏng

1. Mark OSD out:

ceph osd out osd.1

→ Data tự migrate sang OSD khác

2. Wait for recovery:

ceph -w

Đợi "recovery" events hoàn tất

3. Remove OSD:

ceph osd purge osd.1 --yes-i-really-mean-it

4. Replace physical disk

5. Update CephCluster CRD (hoặc Rook tự detect disk mới)

→ Rook Operator sẽ tạo OSD mới tự động

6. Verify:

ceph osd tree

4.3. Pool Full Emergency

# ⚠️ Khi pool đạt near_full (85%) hoặc full (95%):
ceph osd pool set replicapool full_ratio 0.97           # Tạm tăng ratio
ceph osd pool set replicapool nearfull_ratio 0.90

Cleanup:

1. Xóa data không cần thiết

2. Thêm OSD mới

3. Giảm replication size tạm thời (⚠️ risky)


PART 5: STORAGE BENCHMARKING__HTMLTAG_193___
# RADOS bench (raw RADOS performance):
kubectl -n rook-ceph exec deploy/rook-ceph-tools -- \
  rados bench -p replicapool 60 write --no-cleanup
# Output:
# Total time run:       60.000000
# Total writes made:    15000
# Write size:           4194304
# Object size:          4194304
# Bandwidth (MB/sec):   1000.00
# Average IOPS:         250

kubectl -n rook-ceph exec deploy/rook-ceph-tools -- \
  rados bench -p replicapool 60 seq
# -> Sequential read benchmark

# Cleanup bench data:
kubectl -n rook-ceph exec deploy/rook-ceph-tools -- \
  rados -p replicapool cleanup

# fio benchmark (from pod):
# Deploy fio pod with ceph-block PVC → benchmark IOPS, throughput, latency

💡 KEY TAKEAWAYS

  1. Prometheus + Grafana: Monitor Ceph health, latency, IOPS, capacity
  2. Scrub scheduling: Off-peak hours (2-6 AM) to minimize impact
  3. Recovery throttling: osd_recovery_max_active, osd_max_backfills
  4. PG autoscaler: Automatically adjust PG count
  5. OSD replace: mark out → wait recovery → purge → replace disk
  6. Benchmark before deployment: rados bench for baseline

🎯 EXERCISE

Exercise 1: Monitoring

  • Verify Prometheus scraping Ceph metrics__HTMLTAG_230___
  • Check ceph health detail and resolve warnings

Exercise 2: Benchmark

  • Run rados bench write/read
  • Deploy fio in pod with ceph-block PVC
  • Compare IOPS and latency__HTMLTAG_244___

Exercise 3: OSD Failure Simulation__HTMLTAG_247___
  • Mark 1 OSD out, observe recovery
  • Mark OSD print, observe rebalancing

📚 NEXT POST

In Lesson 16: PostgreSQL HA Architecture with Patroni and CloudNativePG, we will start Part 4 — database HA for microservices.