Xây dựng Grafana unified dashboards, SLI/SLO/Error Budget monitoring, dashboard-as-code với Grafonnet, alerting workflows, và production-grade observability stack.
Deploy Microservices On-Premises với Kubernetes HA
Phần 8: Observability — Prometheus, Loki, Tempo
xdev.asia
🎯 MỤC TIÊU BÀI HỌC
- ✅ Grafana unified dashboard design cho microservices
- ✅ SLI/SLO/Error Budget concepts và implementation
- ✅ Dashboard-as-Code với ConfigMap/Grafonnet
- ✅ Alert workflows và on-call routing
- ✅ Production observability checklist
PHẦN 1: SLI / SLO / ERROR BUDGET
| Concept | Definition | Example |
| SLI (Service Level Indicator) | Metric that measures service quality | 99.2% requests < 500ms |
| SLO (Service Level Objective) | Target value for an SLI | 99.9% availability per month |
| Error Budget | Allowed downtime = 1 - SLO | 99.9% → 43.2 min/month |
| SLA (Service Level Agreement) | Contract with consequences | 99.9% or refund credit |
Error Budget Calculation:
SLO = 99.9% availability
Error Budget = 100% - 99.9% = 0.1%
Per month (30 days):
0.1% × 30 × 24 × 60 = 43.2 minutes downtime allowed
Budget consumption tracking:
┌────────────────────────────────────────┐
│ Error Budget: 43.2 min/month │
│ ████████████████░░░░ 75% remaining │
│ Used: 10.8 min | Left: 32.4 min │
│ Burn rate: 1.2x (slightly fast) │
└────────────────────────────────────────┘
PHẦN 2: SLO IMPLEMENTATION VỚI PROMETHEUS
# PrometheusRule for SLO:
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: slo-rules
namespace: monitoring
spec:
groups:
- name: slo-availability
rules:
# SLI: Request success rate
- record: sli:http_requests:availability
expr: |
sum(rate(http_server_request_duration_seconds_count{http_status_code!~"5.."}[5m])) by (service)
/
sum(rate(http_server_request_duration_seconds_count[5m])) by (service)
# SLI: Latency P99 < 500ms
- record: sli:http_requests:latency_ok
expr: |
sum(rate(http_server_request_duration_seconds_bucket{le="0.5"}[5m])) by (service)
/
sum(rate(http_server_request_duration_seconds_count[5m])) by (service)
# Error budget remaining (30-day window):
- record: slo:error_budget:remaining
expr: |
1 - (
(1 - sli:http_requests:availability)
/
(1 - 0.999)
)
- name: slo-alerts
rules:
# Burn rate alert (fast burn):
- alert: SLOBurnRateFast
expr: |
(
1 - sli:http_requests:availability
) / (1 - 0.999) > 14.4
for: 2m
labels:
severity: critical
annotations:
summary: "{{ $labels.service }} SLO burn rate 14.4x (budget exhausted in 1h)"
# Burn rate alert (slow burn):
- alert: SLOBurnRateSlow
expr: |
(
1 - sli:http_requests:availability
) / (1 - 0.999) > 3
for: 30m
labels:
severity: warning
annotations:
summary: "{{ $labels.service }} SLO burn rate 3x (budget exhausted in 10h)"
# Error budget exhausted:
- alert: SLOErrorBudgetExhausted
expr: slo:error_budget:remaining < 0
for: 5m
labels:
severity: critical
annotations:
summary: "{{ $labels.service }} has exhausted its error budget"
PHẦN 3: GRAFANA DASHBOARD DESIGN
Dashboard Hierarchy:
Level 1: Platform Overview
┌──────────────────────────────────────┐
│ K8s Cluster Health │ SLO Status │
│ Nodes: 7/7 ✅ │ Order: 99.95│
│ Pods: 142/150 │ Payment:99.9│
│ CPU: 62% Mem: 71% │ User: 99.99 │
└──────────────────────────────────────┘
│ click service
▼
Level 2: Service Dashboard (RED Method)
┌──────────────────────────────────────┐
│ Rate: 1.2K req/s │
│ Errors: 0.1% (SLO: < 0.1%) │
│ Duration: P50=12ms P99=180ms │
│ ───────────────────────────── │
│ Top Endpoints │ Error Breakdown │
│ Recent Traces │ Log Errors │
└──────────────────────────────────────┘
│ click endpoint
▼
Level 3: Request Detail
┌──────────────────────────────────────┐
│ Trace: abc123 │
│ order-svc (12ms) → payment (80ms) │
│ Logs: 3 entries │
│ DB queries: 2 (5ms total) │
└──────────────────────────────────────┘
# Dashboard as ConfigMap:
apiVersion: v1
kind: ConfigMap
metadata:
name: service-dashboard
namespace: monitoring
labels:
grafana_dashboard: "true"
data:
service-overview.json: |
{
"dashboard": {
"title": "Service Overview - RED Method",
"tags": ["microservices", "red"],
"templating": {
"list": [
{
"name": "service",
"type": "query",
"query": "label_values(http_server_request_duration_seconds_count, service)",
"refresh": 2
},
{
"name": "namespace",
"type": "query",
"query": "label_values(kube_pod_info, namespace)",
"refresh": 2
}
]
},
"panels": [
{
"title": "Request Rate",
"type": "timeseries",
"targets": [{
"expr": "sum(rate(http_server_request_duration_seconds_count{service=\"$service\"}[5m]))"
}]
},
{
"title": "Error Rate",
"type": "stat",
"targets": [{
"expr": "1 - sli:http_requests:availability{service=\"$service\"}"
}],
"fieldConfig": {
"defaults": {
"thresholds": {
"steps": [
{"color": "green", "value": 0},
{"color": "yellow", "value": 0.001},
{"color": "red", "value": 0.01}
]
},
"unit": "percentunit"
}
}
},
{
"title": "Latency P99",
"type": "timeseries",
"targets": [{
"expr": "histogram_quantile(0.99, sum(rate(http_server_request_duration_seconds_bucket{service=\"$service\"}[5m])) by (le))"
}]
}
]
}
}
PHẦN 4: ALERTING WORKFLOWS
# Alertmanager routing:
apiVersion: monitoring.coreos.com/v1alpha1
kind: AlertmanagerConfig
metadata:
name: alert-routing
namespace: monitoring
spec:
route:
groupBy: ['alertname', 'service']
groupWait: 30s
groupInterval: 5m
repeatInterval: 4h
receiver: default-slack
routes:
- match:
severity: critical
receiver: pagerduty-oncall
repeatInterval: 5m
- match:
severity: warning
receiver: team-slack
repeatInterval: 1h
receivers:
- name: pagerduty-oncall
pagerdutyConfigs:
- routingKey:
name: pagerduty-secret
key: routing-key
severity: critical
- name: team-slack
slackConfigs:
- apiURL:
name: slack-webhook
key: url
channel: '#alerts-warning'
title: '{{ .GroupLabels.alertname }}'
text: >-
*Service:* {{ .GroupLabels.service }}
*Summary:* {{ range .Alerts }}{{ .Annotations.summary }}{{ end }}
PHẦN 5: PRODUCTION OBSERVABILITY CHECKLIST
| Category | Item | Tool |
| Metrics | RED method per service | Prometheus |
| Metrics | USE method per node | node-exporter |
| Metrics | SLO/Error budget | Recording rules |
| Logs | Centralized structured logs | Loki + Promtail |
| Logs | Log-based alerts | Loki ruler |
| Traces | Distributed tracing | Tempo + OTel |
| Traces | Trace-log-metric correlation | Grafana |
| Alerts | Multi-burn-rate SLO alerts | Alertmanager |
| Alerts | On-call routing | PagerDuty |
| Dashboards | 3-level drill-down | Grafana |
💡 KEY TAKEAWAYS
- SLO: Define error budget, alert on burn rate (not just threshold)
- RED Method: Rate, Errors, Duration per service
- Dashboard hierarchy: Platform → Service → Request detail
- Dashboard-as-Code: ConfigMap + sidecar auto-provision
- Alert routing: Critical → PagerDuty, Warning → Slack
- Correlation: Metrics → Traces → Logs = full picture
🎯 BÀI TẬP
Bài tập 1: SLO Setup
- Define SLIs cho sample service (availability + latency)
- Create recording rules + burn rate alerts
- Build Error Budget dashboard
Bài tập 2: Unified Dashboard
- Create 3-level dashboard hierarchy
- Configure trace-log-metric linking
- Simulate incident, use observability stack to find root cause
📚 BÀI TIẾP THEO
Trong Bài 36: RBAC & Pod Security Standards, chúng ta sẽ bắt đầu Section 9 — Security Hardening cho K8s cluster.