Chuyển đến nội dung chính

BÀI 35: GRAFANA DASHBOARDS & SLO MONITORING

Xây dựng Grafana unified dashboards, SLI/SLO/Error Budget monitoring, dashboard-as-code với Grafonnet, alerting workflows, và production-grade observability stack.

🔒 DevSecOps — Bài 35 BÀI 35: GRAFANA DASHBOARDS & SLO MONITORING

Deploy Microservices On-Premises với Kubernetes HA

Phần 8: Observability — Prometheus, Loki, Tempo

xdev.asia

🎯 MỤC TIÊU BÀI HỌC

  • ✅ Grafana unified dashboard design cho microservices
  • ✅ SLI/SLO/Error Budget concepts và implementation
  • ✅ Dashboard-as-Code với ConfigMap/Grafonnet
  • ✅ Alert workflows và on-call routing
  • ✅ Production observability checklist

PHẦN 1: SLI / SLO / ERROR BUDGET

ConceptDefinitionExample
SLI (Service Level Indicator)Metric that measures service quality99.2% requests < 500ms
SLO (Service Level Objective)Target value for an SLI99.9% availability per month
Error BudgetAllowed downtime = 1 - SLO99.9% → 43.2 min/month
SLA (Service Level Agreement)Contract with consequences99.9% or refund credit

Error Budget Calculation:

SLO = 99.9% availability
Error Budget = 100% - 99.9% = 0.1%

Per month (30 days):
  0.1% × 30 × 24 × 60 = 43.2 minutes downtime allowed

Budget consumption tracking:
┌────────────────────────────────────────┐
│ Error Budget: 43.2 min/month           │
│ ████████████████░░░░ 75% remaining     │
│ Used: 10.8 min  |  Left: 32.4 min      │
│ Burn rate: 1.2x (slightly fast)        │
└────────────────────────────────────────┘

PHẦN 2: SLO IMPLEMENTATION VỚI PROMETHEUS

# PrometheusRule for SLO:
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: slo-rules
  namespace: monitoring
spec:
  groups:
    - name: slo-availability
      rules:
        # SLI: Request success rate
        - record: sli:http_requests:availability
          expr: |
            sum(rate(http_server_request_duration_seconds_count{http_status_code!~"5.."}[5m])) by (service)
            /
            sum(rate(http_server_request_duration_seconds_count[5m])) by (service)

        # SLI: Latency P99 < 500ms
        - record: sli:http_requests:latency_ok
          expr: |
            sum(rate(http_server_request_duration_seconds_bucket{le="0.5"}[5m])) by (service)
            /
            sum(rate(http_server_request_duration_seconds_count[5m])) by (service)

        # Error budget remaining (30-day window):
        - record: slo:error_budget:remaining
          expr: |
            1 - (
              (1 - sli:http_requests:availability)
              /
              (1 - 0.999)
            )

    - name: slo-alerts
      rules:
        # Burn rate alert (fast burn):
        - alert: SLOBurnRateFast
          expr: |
            (
              1 - sli:http_requests:availability
            ) / (1 - 0.999) > 14.4
          for: 2m
          labels:
            severity: critical
          annotations:
            summary: "{{ $labels.service }} SLO burn rate 14.4x (budget exhausted in 1h)"

        # Burn rate alert (slow burn):
        - alert: SLOBurnRateSlow
          expr: |
            (
              1 - sli:http_requests:availability
            ) / (1 - 0.999) > 3
          for: 30m
          labels:
            severity: warning
          annotations:
            summary: "{{ $labels.service }} SLO burn rate 3x (budget exhausted in 10h)"

        # Error budget exhausted:
        - alert: SLOErrorBudgetExhausted
          expr: slo:error_budget:remaining < 0
          for: 5m
          labels:
            severity: critical
          annotations:
            summary: "{{ $labels.service }} has exhausted its error budget"

PHẦN 3: GRAFANA DASHBOARD DESIGN


Dashboard Hierarchy:

Level 1: Platform Overview
┌──────────────────────────────────────┐
│  K8s Cluster Health  │  SLO Status  │
│  Nodes: 7/7 ✅       │  Order: 99.95│
│  Pods: 142/150       │  Payment:99.9│
│  CPU: 62%  Mem: 71%  │  User: 99.99 │
└──────────────────────────────────────┘
        │ click service
        ▼
Level 2: Service Dashboard (RED Method)
┌──────────────────────────────────────┐
│  Rate: 1.2K req/s                    │
│  Errors: 0.1% (SLO: < 0.1%)         │
│  Duration: P50=12ms P99=180ms        │
│  ─────────────────────────────       │
│  Top Endpoints │ Error Breakdown     │
│  Recent Traces │ Log Errors          │
└──────────────────────────────────────┘
        │ click endpoint
        ▼
Level 3: Request Detail
┌──────────────────────────────────────┐
│  Trace: abc123                       │
│  order-svc (12ms) → payment (80ms)   │
│  Logs: 3 entries                     │
│  DB queries: 2 (5ms total)           │
└──────────────────────────────────────┘
# Dashboard as ConfigMap:
apiVersion: v1
kind: ConfigMap
metadata:
  name: service-dashboard
  namespace: monitoring
  labels:
    grafana_dashboard: "true"
data:
  service-overview.json: |
    {
      "dashboard": {
        "title": "Service Overview - RED Method",
        "tags": ["microservices", "red"],
        "templating": {
          "list": [
            {
              "name": "service",
              "type": "query",
              "query": "label_values(http_server_request_duration_seconds_count, service)",
              "refresh": 2
            },
            {
              "name": "namespace",
              "type": "query", 
              "query": "label_values(kube_pod_info, namespace)",
              "refresh": 2
            }
          ]
        },
        "panels": [
          {
            "title": "Request Rate",
            "type": "timeseries",
            "targets": [{
              "expr": "sum(rate(http_server_request_duration_seconds_count{service=\"$service\"}[5m]))"
            }]
          },
          {
            "title": "Error Rate",
            "type": "stat",
            "targets": [{
              "expr": "1 - sli:http_requests:availability{service=\"$service\"}"
            }],
            "fieldConfig": {
              "defaults": {
                "thresholds": {
                  "steps": [
                    {"color": "green", "value": 0},
                    {"color": "yellow", "value": 0.001},
                    {"color": "red", "value": 0.01}
                  ]
                },
                "unit": "percentunit"
              }
            }
          },
          {
            "title": "Latency P99",
            "type": "timeseries",
            "targets": [{
              "expr": "histogram_quantile(0.99, sum(rate(http_server_request_duration_seconds_bucket{service=\"$service\"}[5m])) by (le))"
            }]
          }
        ]
      }
    }

PHẦN 4: ALERTING WORKFLOWS

# Alertmanager routing:
apiVersion: monitoring.coreos.com/v1alpha1
kind: AlertmanagerConfig
metadata:
  name: alert-routing
  namespace: monitoring
spec:
  route:
    groupBy: ['alertname', 'service']
    groupWait: 30s
    groupInterval: 5m
    repeatInterval: 4h
    receiver: default-slack
    routes:
      - match:
          severity: critical
        receiver: pagerduty-oncall
        repeatInterval: 5m
      - match:
          severity: warning
        receiver: team-slack
        repeatInterval: 1h

  receivers:
    - name: pagerduty-oncall
      pagerdutyConfigs:
        - routingKey:
            name: pagerduty-secret
            key: routing-key
          severity: critical
    
    - name: team-slack
      slackConfigs:
        - apiURL:
            name: slack-webhook
            key: url
          channel: '#alerts-warning'
          title: '{{ .GroupLabels.alertname }}'
          text: >-
            *Service:* {{ .GroupLabels.service }}
            *Summary:* {{ range .Alerts }}{{ .Annotations.summary }}{{ end }}

PHẦN 5: PRODUCTION OBSERVABILITY CHECKLIST

CategoryItemTool
MetricsRED method per servicePrometheus
MetricsUSE method per nodenode-exporter
MetricsSLO/Error budgetRecording rules
LogsCentralized structured logsLoki + Promtail
LogsLog-based alertsLoki ruler
TracesDistributed tracingTempo + OTel
TracesTrace-log-metric correlationGrafana
AlertsMulti-burn-rate SLO alertsAlertmanager
AlertsOn-call routingPagerDuty
Dashboards3-level drill-downGrafana

💡 KEY TAKEAWAYS

  1. SLO: Define error budget, alert on burn rate (not just threshold)
  2. RED Method: Rate, Errors, Duration per service
  3. Dashboard hierarchy: Platform → Service → Request detail
  4. Dashboard-as-Code: ConfigMap + sidecar auto-provision
  5. Alert routing: Critical → PagerDuty, Warning → Slack
  6. Correlation: Metrics → Traces → Logs = full picture

🎯 BÀI TẬP

Bài tập 1: SLO Setup

  • Define SLIs cho sample service (availability + latency)
  • Create recording rules + burn rate alerts
  • Build Error Budget dashboard

Bài tập 2: Unified Dashboard

  • Create 3-level dashboard hierarchy
  • Configure trace-log-metric linking
  • Simulate incident, use observability stack to find root cause

📚 BÀI TIẾP THEO

Trong Bài 36: RBAC & Pod Security Standards, chúng ta sẽ bắt đầu Section 9 — Security Hardening cho K8s cluster.