Chuyển đến nội dung chính

第 15 課:指標 — Prometheus 和 Grafana

RED方法、USE方法、Prometheus架構、基本PromQL、Kubernetes中的ServiceMonitor、Grafana儀表板設計、警報規則和Alertmanager配置。

🏗️ 建築 — 第 15 課 第 15 課:指標 — Prometheus 和 Grafana

雲端原生微服務架構

第 5 部分:可觀察性 — 三大支柱

亞洲開發網

第 15 課:指標 — Prometheus 和 Grafana

簡介

在微服務系統中,您無法像傳統除錯那樣「查看」正在運行的服務的內部。 Metrics 是 Observability 的第一個平台 — 提供衡量系統健康狀況和效能隨時間變化的數值資料。

Prometheus + Grafana 是雲原生生態系中用於指標收集和視覺化的事實上的標準組合。


1. 測量方法

1.1 RED 方法 — 對於服務

Tom Wilkie 的 RED 方法 重點關注每項服務的三個最重要的指標:

R — Rate     : Số lượng request service xử lý mỗi giây
E — Errors   : Tỷ lệ phần trăm request thất bại
D — Duration : Thời gian phân phối (latency percentiles)

訂單服務儀表板範例:

┌─────────────────────────────────────────────────────────┐
│                  Order Service — RED                    │
├──────────────────┬──────────────────┬───────────────────┤
│ Rate             │ Error Rate       │ Duration (p99)    │
│ 1,247 req/s      │ 0.3%             │ 145ms             │
│ ↑ +12% vs 1h ago│ ↓ Normal         │ → Within SLO      │
└──────────────────┴──────────────────┴───────────────────┘

1.2 USE 方法-對於資源

使用 Brendan Gregg 的方法取得基礎設施資源:

U — Utilization : Tỷ lệ thời gian resource đang busy (%)
S — Saturation  : Mức độ "hàng đợi" extra work (queue length)
E — Errors      : Số lượng lỗi từ resource

適用於每個資源:

資源利用率飽和度錯誤
中央處理器rate(cpu_seconds[5m])平均負載 > 核心機器檢查錯誤
內存1 - mem_free/mem_total掉期率、OOM 殺戮
磁碟 I/Orate(disk_io_time[5m])磁碟佇列長度I/O 錯誤
網路rate(net_bytes[5m]) / capacity封包遺失/重傳網卡錯誤

1.3 黃金訊號-端到端

Google SRE 定義了四個黃金訊號:

  1. 延遲 — 請求處理時間(區分成功與錯誤延遲)
  2. 流量 - 對系統的需求(請求數/秒,同時使用者數)
  3. Errors-請求失敗率(明確5xx,隱式錯誤資料)
  4. 飽和度 — 系統資源飽和度

2. 普羅米修斯架構

2.1 概述

┌─────────────────────────────────────────────────────────────┐
│                      Prometheus Server                      │
│                                                             │
│  ┌──────────────┐   ┌──────────────┐   ┌────────────────┐  │
│  │ Retrieval    │   │ TSDB         │   │ HTTP Server    │  │
│  │ (Scraper)    │   │ (Time-Series │   │ (Query/API)    │  │
│  │              │   │  Database)   │   │                │  │
│  └──────┬───────┘   └──────────────┘   └────────┬───────┘  │
│         │                                        │          │
└─────────┼────────────────────────────────────────┼──────────┘
          │ /metrics scrape                        │ PromQL
          ▼                                        ▼
 ┌──────────────────┐                    ┌─────────────────┐
 │ Targets:         │                    │   Grafana       │
 │ - App /metrics   │                    │   Alertmanager  │
 │ - Node Exporter  │                    │   Other clients │
 │ - cAdvisor       │                    └─────────────────┘
 │ - kube-state-mts │
 └──────────────────┘

基於拉取的模型:Prometheus 每 15-30 秒主動從目標中抓取指標。與基於推送的(InfluxDB、StatsD)不同。

2.2 指標類型

# Counter — Chỉ tăng, dùng cho counts và rates
http_requests_total{method="POST", status="200"} 1027

# Gauge — Tăng giảm tự do, dùng cho current values
memory_usage_bytes 153344000
active_connections 42

# Histogram — Phân phối latency, tạo buckets tự động
http_request_duration_seconds_bucket{le="0.1"} 8521
http_request_duration_seconds_bucket{le="0.5"} 9812
http_request_duration_seconds_sum 1234.6789
http_request_duration_seconds_count 9987

# Summary — Như Histogram nhưng tính percentile phía client
http_request_duration_seconds{quantile="0.99"} 0.145

2.3 說明格式

每個服務都公開指標 /metrics 端點:

// Spring Boot — thêm dependency
// actuator + micrometer-registry-prometheus

// Sau đó endpoint tự động có:
// GET /actuator/prometheus
// Go — prometheus/client_golang
import "github.com/prometheus/client_golang/prometheus"
import "github.com/prometheus/client_golang/prometheus/promauto"

var (
    httpRequestsTotal = promauto.NewCounterVec(
        prometheus.CounterOpts{
            Name: "http_requests_total",
            Help: "Total HTTP requests",
        },
        []string{"method", "path", "status"},
    )

    httpDuration = promauto.NewHistogramVec(
        prometheus.HistogramOpts{
            Name:    "http_request_duration_seconds",
            Buckets: prometheus.DefBuckets,
        },
        []string{"method", "path"},
    )
)

3.Kubernetes 集成

3.1 kube-prometheus-stack

最簡單的方法:安裝套件 kube-prometheus-stack Helm 圖表 — 包括 Prometheus、Grafana、Alertmanager 和必要的導出器:

helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo update

helm upgrade --install kube-prometheus-stack \
  prometheus-community/kube-prometheus-stack \
  --namespace monitoring \
  --create-namespace \
  --set grafana.adminPassword=<your-password> \
  --set prometheus.prometheusSpec.retention=15d

3.2 服務監控

ServiceMonitor 是 Prometheus Operator 的 CRD(自訂資源定義)-定義如何從服務中取得指標:

apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
  name: order-service
  namespace: services-prod
  labels:
    # Phải match với selector của Prometheus
    release: kube-prometheus-stack
spec:
  namespaceSelector:
    matchNames:
      - services-prod
  selector:
    matchLabels:
      app: order-service
  endpoints:
    - port: http
      path: /actuator/prometheus
      interval: 15s
      scrapeTimeout: 10s

3.3 PodMonitor

當服務沒有 Kubernetes Service(只有直接 Pod):

apiVersion: monitoring.coreos.com/v1
kind: PodMonitor
metadata:
  name: worker-pods
spec:
  selector:
    matchLabels:
      app: async-worker
  podMetricsEndpoints:
    - port: metrics
      path: /metrics
      interval: 30s

4. PromQL — Prometheus 查詢語言

4.1 基本文法

# Instant vector — giá trị tại thời điểm hiện tại
http_requests_total

# Range vector — giá trị trong khoảng thời gian
http_requests_total[5m]

# Filtering bằng labels
http_requests_total{job="order-service", status=~"5.."}

# Operators
http_requests_total{status=~"5.."} / http_requests_total  # ratio

4.2 重要功能

# rate() — tốc độ thay đổi per second (dùng cho Counter)
rate(http_requests_total[5m])

# irate() — instant rate (nhạy hơn với spike ngắn)
irate(http_requests_total[5m])

# increase() — tổng tăng trong khoảng thời gian
increase(http_requests_total[1h])

# histogram_quantile() — percentile từ Histogram
histogram_quantile(0.99, rate(http_request_duration_seconds_bucket[5m]))

# sum() với by/without
sum(rate(http_requests_total[5m])) by (service)

# topk()
topk(5, rate(http_requests_total[5m]))

4.3 實際查詢

# Error rate (5xx) cho tất cả services
sum(rate(http_requests_total{status=~"5.."}[5m])) by (service)
/
sum(rate(http_requests_total[5m])) by (service)

# p99 latency theo service
histogram_quantile(
  0.99,
  sum(rate(http_request_duration_seconds_bucket[5m])) by (service, le)
)

# CPU usage trên các pods
sum(rate(container_cpu_usage_seconds_total{container!=""}[5m])) by (pod, namespace)

# Memory usage (bytes)
sum(container_memory_rss{container!=""}) by (pod, namespace)

# Pod restart count
kube_pod_container_status_restarts_total{namespace="services-prod"}

# Service availability (based on successful requests)
1 - (
  sum(rate(http_requests_total{status=~"5.."}[5m])) by (service)
  /
  sum(rate(http_requests_total[5m])) by (service)
)

5.Grafana 儀表板

5.1 良好的儀表板結構

儀表板設計原則:

Level 1 — Overview (Top Row)
  Tổng quan toàn bộ hệ thống:
  - Total request rate
  - Overall error rate
  - System availability
  - Active alerts count

Level 2 — Service Overview
  Metrics per service:
  - Rate, Errors, Duration (RED)
  - Service status

Level 3 — Drill-down
  Chi tiết khi có vấn đề:
  - Request breakdown by endpoint
  - Latency percentiles (p50, p95, p99)
  - Error messages
  - Dependencies

5.2 儀表板作為程式碼

將儀表板以 JSON 格式儲存到 Git:

# Grafana ConfigMap trong Kubernetes
apiVersion: v1
kind: ConfigMap
metadata:
  name: order-service-dashboard
  namespace: monitoring
  labels:
    grafana_dashboard: "1"  # Auto-discovered bởi Grafana sidecar
data:
  order-service.json: |
    {
      "title": "Order Service",
      "uid": "order-service",
      "panels": [ ... ]
    }

6. 警報

6.1 普羅米修斯規則

警報規則定義:

apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: order-service-alerts
  namespace: services-prod
  labels:
    release: kube-prometheus-stack
spec:
  groups:
    - name: order-service.rules
      interval: 30s
      rules:
        # Error rate cao
        - alert: HighErrorRate
          expr: |
            sum(rate(http_requests_total{service="order-service", status=~"5.."}[5m]))
            /
            sum(rate(http_requests_total{service="order-service"}[5m]))
            > 0.05
          for: 2m
          labels:
            severity: critical
            team: backend
          annotations:
            summary: "High error rate on order-service"
            description: "Error rate {{ $value | humanizePercentage }} > 5%"

        # Latency cao
        - alert: HighLatency
          expr: |
            histogram_quantile(0.99,
              sum(rate(http_request_duration_seconds_bucket{service="order-service"}[5m]))
              by (le)
            ) > 1.0
          for: 5m
          labels:
            severity: warning
          annotations:
            summary: "High p99 latency on order-service"
            description: "p99 latency is {{ $value | humanizeDuration }}"

        # Pod down
        - alert: PodDown
          expr: |
            kube_deployment_status_replicas_available{
              namespace="services-prod",
              deployment="order-service"
            } < kube_deployment_spec_replicas{
              namespace="services-prod",
              deployment="order-service"
            }
          for: 1m
          labels:
            severity: critical
          annotations:
            summary: "Order service pod(s) down"

6.2 警報管理器

Alertmanager接收來自Prometheus的警報並發送通知:

# alertmanager.yaml
global:
  slack_api_url: 'https://hooks.slack.com/services/...'

route:
  group_by: ['alertname', 'service']
  group_wait: 30s
  group_interval: 5m
  repeat_interval: 1h
  receiver: 'default'
  routes:
    - match:
        severity: critical
      receiver: 'pagerduty-critical'
      continue: true
    - match:
        severity: warning
      receiver: 'slack-warnings'

receivers:
  - name: 'default'
    slack_configs:
      - channel: '#alerts'
        text: '{{ .CommonAnnotations.summary }}'

  - name: 'pagerduty-critical'
    pagerduty_configs:
      - service_key: '<key>'

  - name: 'slack-warnings'
    slack_configs:
      - channel: '#alerts-warning'
        send_resolved: true

7.SLI/SLO/SLA

將指標與服務等級目標結合:

# SLO Definition
service: order-service
slo:
  # 99.9% requests thành công trong 30 ngày
  - name: availability
    target: 99.9%
    indicator:
      ratio:
        good_events: http_requests_total{status!~"5.."}
        total_events: http_requests_total

  # 95% requests hoàn thành trong 200ms
  - name: latency
    target: 95%
    indicator:
      ratio:
        good_events: http_request_duration_seconds_bucket{le="0.2"}
        total_events: http_request_duration_seconds_count

錯誤預算:

  • SLO 99.9% → 錯誤預算 = 0.1% = 43.8 分鐘/月
  • 當錯誤預算耗盡→凍結功能部署,專注於可靠性

8. 最佳實踐

正確命名指標:

# Format: <namespace>_<subsystem>_<name>_<unit>
http_request_duration_seconds
database_queries_total
cache_hit_ratio
background_jobs_processed_total

基數控制:

# ĐÚNG — cardinality thấp, controllable labels
http_requests_total{method, status_code, service}

# SAI — cardinality explode làm Prometheus OOM
http_requests_total{user_id, request_id, ip_address}

預先計算昂貴查詢的記錄規則:

groups:
  - name: recording_rules
    rules:
      - record: job:http_requests_total:rate5m
        expr: sum(rate(http_requests_total[5m])) by (job)

總結

概念目的
紅色方法衡量每項服務的健康狀況
使用方法衡量基礎設施資源的健康狀況
服務監控將 Prometheus 與 Kubernetes 整合
PromQL查詢與計算指標
普羅米修斯規則警報規則定義
警報管理器路由與發送通知
SLO/SLI量化可靠性目標

下一篇文章:日誌記錄 - 結構化日誌記錄、Loki 和 ELK Stack