
簡介
在微服務系統中,您無法像傳統除錯那樣「查看」正在運行的服務的內部。 Metrics 是 Observability 的第一個平台 — 提供衡量系統健康狀況和效能隨時間變化的數值資料。
Prometheus + Grafana 是雲原生生態系中用於指標收集和視覺化的事實上的標準組合。
1. 測量方法
1.1 RED 方法 — 對於服務
Tom Wilkie 的 RED 方法 重點關注每項服務的三個最重要的指標:
R — Rate : Số lượng request service xử lý mỗi giây
E — Errors : Tỷ lệ phần trăm request thất bại
D — Duration : Thời gian phân phối (latency percentiles)
訂單服務儀表板範例:
┌─────────────────────────────────────────────────────────┐
│ Order Service — RED │
├──────────────────┬──────────────────┬───────────────────┤
│ Rate │ Error Rate │ Duration (p99) │
│ 1,247 req/s │ 0.3% │ 145ms │
│ ↑ +12% vs 1h ago│ ↓ Normal │ → Within SLO │
└──────────────────┴──────────────────┴───────────────────┘
1.2 USE 方法-對於資源
使用 Brendan Gregg 的方法取得基礎設施資源:
U — Utilization : Tỷ lệ thời gian resource đang busy (%)
S — Saturation : Mức độ "hàng đợi" extra work (queue length)
E — Errors : Số lượng lỗi từ resource
適用於每個資源:
| 資源 | 利用率 | 飽和度 | 錯誤 |
|---|---|---|---|
| 中央處理器 | rate(cpu_seconds[5m]) | 平均負載 > 核心 | 機器檢查錯誤 |
| 內存 | 1 - mem_free/mem_total | 掉期率、OOM 殺戮 | |
| 磁碟 I/O | rate(disk_io_time[5m]) | 磁碟佇列長度 | I/O 錯誤 |
| 網路 | rate(net_bytes[5m]) / capacity | 封包遺失/重傳 | 網卡錯誤 |
1.3 黃金訊號-端到端
Google SRE 定義了四個黃金訊號:
- 延遲 — 請求處理時間(區分成功與錯誤延遲)
- 流量 - 對系統的需求(請求數/秒,同時使用者數)
- Errors-請求失敗率(明確5xx,隱式錯誤資料)
- 飽和度 — 系統資源飽和度
2. 普羅米修斯架構
2.1 概述
┌─────────────────────────────────────────────────────────────┐
│ Prometheus Server │
│ │
│ ┌──────────────┐ ┌──────────────┐ ┌────────────────┐ │
│ │ Retrieval │ │ TSDB │ │ HTTP Server │ │
│ │ (Scraper) │ │ (Time-Series │ │ (Query/API) │ │
│ │ │ │ Database) │ │ │ │
│ └──────┬───────┘ └──────────────┘ └────────┬───────┘ │
│ │ │ │
└─────────┼────────────────────────────────────────┼──────────┘
│ /metrics scrape │ PromQL
▼ ▼
┌──────────────────┐ ┌─────────────────┐
│ Targets: │ │ Grafana │
│ - App /metrics │ │ Alertmanager │
│ - Node Exporter │ │ Other clients │
│ - cAdvisor │ └─────────────────┘
│ - kube-state-mts │
└──────────────────┘
基於拉取的模型:Prometheus 每 15-30 秒主動從目標中抓取指標。與基於推送的(InfluxDB、StatsD)不同。
2.2 指標類型
# Counter — Chỉ tăng, dùng cho counts và rates
http_requests_total{method="POST", status="200"} 1027
# Gauge — Tăng giảm tự do, dùng cho current values
memory_usage_bytes 153344000
active_connections 42
# Histogram — Phân phối latency, tạo buckets tự động
http_request_duration_seconds_bucket{le="0.1"} 8521
http_request_duration_seconds_bucket{le="0.5"} 9812
http_request_duration_seconds_sum 1234.6789
http_request_duration_seconds_count 9987
# Summary — Như Histogram nhưng tính percentile phía client
http_request_duration_seconds{quantile="0.99"} 0.145
2.3 說明格式
每個服務都公開指標 /metrics 端點:
// Spring Boot — thêm dependency
// actuator + micrometer-registry-prometheus
// Sau đó endpoint tự động có:
// GET /actuator/prometheus
// Go — prometheus/client_golang
import "github.com/prometheus/client_golang/prometheus"
import "github.com/prometheus/client_golang/prometheus/promauto"
var (
httpRequestsTotal = promauto.NewCounterVec(
prometheus.CounterOpts{
Name: "http_requests_total",
Help: "Total HTTP requests",
},
[]string{"method", "path", "status"},
)
httpDuration = promauto.NewHistogramVec(
prometheus.HistogramOpts{
Name: "http_request_duration_seconds",
Buckets: prometheus.DefBuckets,
},
[]string{"method", "path"},
)
)
3.Kubernetes 集成
3.1 kube-prometheus-stack
最簡單的方法:安裝套件 kube-prometheus-stack Helm 圖表 — 包括 Prometheus、Grafana、Alertmanager 和必要的導出器:
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo update
helm upgrade --install kube-prometheus-stack \
prometheus-community/kube-prometheus-stack \
--namespace monitoring \
--create-namespace \
--set grafana.adminPassword=<your-password> \
--set prometheus.prometheusSpec.retention=15d
3.2 服務監控
ServiceMonitor 是 Prometheus Operator 的 CRD(自訂資源定義)-定義如何從服務中取得指標:
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: order-service
namespace: services-prod
labels:
# Phải match với selector của Prometheus
release: kube-prometheus-stack
spec:
namespaceSelector:
matchNames:
- services-prod
selector:
matchLabels:
app: order-service
endpoints:
- port: http
path: /actuator/prometheus
interval: 15s
scrapeTimeout: 10s
3.3 PodMonitor
當服務沒有 Kubernetes Service(只有直接 Pod):
apiVersion: monitoring.coreos.com/v1
kind: PodMonitor
metadata:
name: worker-pods
spec:
selector:
matchLabels:
app: async-worker
podMetricsEndpoints:
- port: metrics
path: /metrics
interval: 30s
4. PromQL — Prometheus 查詢語言
4.1 基本文法
# Instant vector — giá trị tại thời điểm hiện tại
http_requests_total
# Range vector — giá trị trong khoảng thời gian
http_requests_total[5m]
# Filtering bằng labels
http_requests_total{job="order-service", status=~"5.."}
# Operators
http_requests_total{status=~"5.."} / http_requests_total # ratio
4.2 重要功能
# rate() — tốc độ thay đổi per second (dùng cho Counter)
rate(http_requests_total[5m])
# irate() — instant rate (nhạy hơn với spike ngắn)
irate(http_requests_total[5m])
# increase() — tổng tăng trong khoảng thời gian
increase(http_requests_total[1h])
# histogram_quantile() — percentile từ Histogram
histogram_quantile(0.99, rate(http_request_duration_seconds_bucket[5m]))
# sum() với by/without
sum(rate(http_requests_total[5m])) by (service)
# topk()
topk(5, rate(http_requests_total[5m]))
4.3 實際查詢
# Error rate (5xx) cho tất cả services
sum(rate(http_requests_total{status=~"5.."}[5m])) by (service)
/
sum(rate(http_requests_total[5m])) by (service)
# p99 latency theo service
histogram_quantile(
0.99,
sum(rate(http_request_duration_seconds_bucket[5m])) by (service, le)
)
# CPU usage trên các pods
sum(rate(container_cpu_usage_seconds_total{container!=""}[5m])) by (pod, namespace)
# Memory usage (bytes)
sum(container_memory_rss{container!=""}) by (pod, namespace)
# Pod restart count
kube_pod_container_status_restarts_total{namespace="services-prod"}
# Service availability (based on successful requests)
1 - (
sum(rate(http_requests_total{status=~"5.."}[5m])) by (service)
/
sum(rate(http_requests_total[5m])) by (service)
)
5.Grafana 儀表板
5.1 良好的儀表板結構
儀表板設計原則:
Level 1 — Overview (Top Row)
Tổng quan toàn bộ hệ thống:
- Total request rate
- Overall error rate
- System availability
- Active alerts count
Level 2 — Service Overview
Metrics per service:
- Rate, Errors, Duration (RED)
- Service status
Level 3 — Drill-down
Chi tiết khi có vấn đề:
- Request breakdown by endpoint
- Latency percentiles (p50, p95, p99)
- Error messages
- Dependencies
5.2 儀表板作為程式碼
將儀表板以 JSON 格式儲存到 Git:
# Grafana ConfigMap trong Kubernetes
apiVersion: v1
kind: ConfigMap
metadata:
name: order-service-dashboard
namespace: monitoring
labels:
grafana_dashboard: "1" # Auto-discovered bởi Grafana sidecar
data:
order-service.json: |
{
"title": "Order Service",
"uid": "order-service",
"panels": [ ... ]
}
6. 警報
6.1 普羅米修斯規則
警報規則定義:
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: order-service-alerts
namespace: services-prod
labels:
release: kube-prometheus-stack
spec:
groups:
- name: order-service.rules
interval: 30s
rules:
# Error rate cao
- alert: HighErrorRate
expr: |
sum(rate(http_requests_total{service="order-service", status=~"5.."}[5m]))
/
sum(rate(http_requests_total{service="order-service"}[5m]))
> 0.05
for: 2m
labels:
severity: critical
team: backend
annotations:
summary: "High error rate on order-service"
description: "Error rate {{ $value | humanizePercentage }} > 5%"
# Latency cao
- alert: HighLatency
expr: |
histogram_quantile(0.99,
sum(rate(http_request_duration_seconds_bucket{service="order-service"}[5m]))
by (le)
) > 1.0
for: 5m
labels:
severity: warning
annotations:
summary: "High p99 latency on order-service"
description: "p99 latency is {{ $value | humanizeDuration }}"
# Pod down
- alert: PodDown
expr: |
kube_deployment_status_replicas_available{
namespace="services-prod",
deployment="order-service"
} < kube_deployment_spec_replicas{
namespace="services-prod",
deployment="order-service"
}
for: 1m
labels:
severity: critical
annotations:
summary: "Order service pod(s) down"
6.2 警報管理器
Alertmanager接收來自Prometheus的警報並發送通知:
# alertmanager.yaml
global:
slack_api_url: 'https://hooks.slack.com/services/...'
route:
group_by: ['alertname', 'service']
group_wait: 30s
group_interval: 5m
repeat_interval: 1h
receiver: 'default'
routes:
- match:
severity: critical
receiver: 'pagerduty-critical'
continue: true
- match:
severity: warning
receiver: 'slack-warnings'
receivers:
- name: 'default'
slack_configs:
- channel: '#alerts'
text: '{{ .CommonAnnotations.summary }}'
- name: 'pagerduty-critical'
pagerduty_configs:
- service_key: '<key>'
- name: 'slack-warnings'
slack_configs:
- channel: '#alerts-warning'
send_resolved: true
7.SLI/SLO/SLA
將指標與服務等級目標結合:
# SLO Definition
service: order-service
slo:
# 99.9% requests thành công trong 30 ngày
- name: availability
target: 99.9%
indicator:
ratio:
good_events: http_requests_total{status!~"5.."}
total_events: http_requests_total
# 95% requests hoàn thành trong 200ms
- name: latency
target: 95%
indicator:
ratio:
good_events: http_request_duration_seconds_bucket{le="0.2"}
total_events: http_request_duration_seconds_count
錯誤預算:
- SLO 99.9% → 錯誤預算 = 0.1% = 43.8 分鐘/月
- 當錯誤預算耗盡→凍結功能部署,專注於可靠性
8. 最佳實踐
正確命名指標:
# Format: <namespace>_<subsystem>_<name>_<unit>
http_request_duration_seconds
database_queries_total
cache_hit_ratio
background_jobs_processed_total
基數控制:
# ĐÚNG — cardinality thấp, controllable labels
http_requests_total{method, status_code, service}
# SAI — cardinality explode làm Prometheus OOM
http_requests_total{user_id, request_id, ip_address}
預先計算昂貴查詢的記錄規則:
groups:
- name: recording_rules
rules:
- record: job:http_requests_total:rate5m
expr: sum(rate(http_requests_total[5m])) by (job)
總結
| 概念 | 目的 |
|---|---|
| 紅色方法 | 衡量每項服務的健康狀況 |
| 使用方法 | 衡量基礎設施資源的健康狀況 |
| 服務監控 | 將 Prometheus 與 Kubernetes 整合 |
| PromQL | 查詢與計算指標 |
| 普羅米修斯規則 | 警報規則定義 |
| 警報管理器 | 路由與發送通知 |
| SLO/SLI | 量化可靠性目標 |
下一篇文章:日誌記錄 - 結構化日誌記錄、Loki 和 ELK Stack