Chuyển đến nội dung chính

LESSON 32: PROMETHEUS STACK — MONITORING INFRASTRUCTURE

Deploy kube-prometheus-stack (Prometheus, Grafana, Alertmanager), ServiceMonitor, recording rules, alerting rules, long-term storage with Thanos, and custom metrics.

🔒 DevSecOps — Lesson 32 LESSON 32: PROMETHEUS STACK — MONITORING INFRASTRUCTURE

Deploy Microservices On-Premises with Kubernetes HA

Part 8: Observability — Prometheus, Loki, Tempo

xdev.asia

🎯 LESSON OBJECTIVE__HTMLTAG_68___
  • ✅ Deploy kube-prometheus-stack (Prometheus + Grafana + Alertmanager)
  • ✅ ServiceMonitor and PodMonitor for service discovery
  • ✅ Recording rules for pre-computed metrics__HTMLTAG_75___
  • ✅ Alerting rules and Alertmanager routing__HTMLTAG_77___
  • ✅ Thanos for long-term storage__HTMLTAG_79___
  • ✅ Custom application metrics__HTMLTAG_81___

PART 1: OBSERVABILITY ARCHITECTURE

graph TD
    subgraph PILLARS["📊 THREE PILLARS OF OBSERVABILITY"]
        METRICS["📈 METRICS<br/>Prometheus<br/>Bài 32"]
        LOGS["📝 LOGS<br/>Loki<br/>Bài 33"]
        TRACES["🔗 TRACES<br/>Tempo<br/>Bài 34"]
    end

    METRICS --> GRAFANA["📊 Grafana<br/>Dashboards"]
    LOGS --> GRAFANA
    TRACES --> GRAFANA

    METRICS --> ALERT["🔔 Alertmanager<br/>PagerDuty / Slack"]

    style PILLARS fill:#0f172a,stroke:#3b82f6,color:#e2e8f0
    style METRICS fill:#15803d,stroke:#22c55e,color:#e2e8f0
    style LOGS fill:#1e3a5f,stroke:#3b82f6,color:#e2e8f0
    style TRACES fill:#7c3aed,stroke:#a78bfa,color:#e2e8f0
    style GRAFANA fill:#f59e0b,stroke:#fbbf24,color:#0f172a
    style ALERT fill:#dc2626,stroke:#ef4444,color:#e2e8f0

PART 2: INSTALL KUBE-POMETHEUS-STACK

# Create namespace:
kubectl create namespace monitoring

# Install kube-prometheus-stack:
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo update

helm install prometheus prometheus-community/kube-prometheus-stack \
  --namespace monitoring \
  -f prometheus-values.yaml
# prometheus-values.yaml:
prometheus:
  prometheusSpec:
    replicas: 2
    retention: 30d
    retentionSize: 50GB
    
    resources:
      requests:
        cpu: "1"
        memory: 2Gi
      limits:
        cpu: "4"
        memory: 8Gi
    
    storageSpec:
      volumeClaimTemplate:
        spec:
          storageClassName: ceph-block
          accessModes: ["ReadWriteOnce"]
          resources:
            requests:
              storage: 100Gi
    
    # Scrape all ServiceMonitors in any namespace:
    serviceMonitorSelectorNilUsesHelmValues: false
    podMonitorSelectorNilUsesHelmValues: false
    ruleSelectorNilUsesHelmValues: false

    # External labels (for Thanos):
    externalLabels:
      cluster: production
      environment: production

grafana:
  replicas: 2
  persistence:
    enabled: true
    storageClassName: ceph-block
    size: 10Gi
  
  adminPassword: ""   # Use existing secret
  admin:
    existingSecret: grafana-admin-secret
  
  dashboardProviders:
    dashboardproviders.yaml:
      apiVersion: 1
      providers:
        - name: default
          orgId: 1
          folder: ""
          type: file
          disableDeletion: false
          editable: true
          options:
            path: /var/lib/grafana/dashboards/default
  
  datasources:
    datasources.yaml:
      apiVersion: 1
      datasources:
        - name: Prometheus
          type: prometheus
          url: http://prometheus-kube-prometheus-prometheus:9090
          isDefault: true
        - name: Loki
          type: loki
          url: http://loki.monitoring:3100
        - name: Tempo
          type: tempo
          url: http://tempo.monitoring:3100

alertmanager:
  alertmanagerSpec:
    replicas: 3
    storage:
      volumeClaimTemplate:
        spec:
          storageClassName: ceph-block
          resources:
            requests:
              storage: 5Gi
  
  config:
    global:
      resolve_timeout: 5m
    
    route:
      receiver: 'slack-default'
      group_by: ['alertname', 'namespace']
      group_wait: 30s
      group_interval: 5m
      repeat_interval: 4h
      routes:
        - match:
            severity: critical
          receiver: 'pagerduty-critical'
          group_wait: 10s
        - match:
            severity: warning
          receiver: 'slack-warnings'
    
    receivers:
      - name: 'slack-default'
        slack_configs:
          - api_url: 'https://hooks.slack.com/services/xxx'
            channel: '#alerts'
            title: '{{ .GroupLabels.alertname }}'
            text: '{{ range .Alerts }}{{ .Annotations.summary }}{{ end }}'
      
      - name: 'slack-warnings'
        slack_configs:
          - api_url: 'https://hooks.slack.com/services/xxx'
            channel: '#warnings'
      
      - name: 'pagerduty-critical'
        pagerduty_configs:
          - service_key: 'xxx'
# Verify:
kubectl -n monitoring get pods
# prometheus-kube-prometheus-prometheus-0    2/2   Running
# prometheus-kube-prometheus-prometheus-1    2/2   Running
# prometheus-grafana-xxx                     3/3   Running
# alertmanager-kube-prometheus-alertmanager-0  2/2 Running
# prometheus-kube-state-metrics-xxx          1/1   Running
# prometheus-node-exporter-xxx (per node)    1/1   Running

PART 3: SERVICEMONITOR & PODMONITOR

# Custom ServiceMonitor cho ứng dụng:
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
  name: order-service-monitor
  namespace: monitoring
  labels:
    release: prometheus    # Must match Prometheus selector
spec:
  namespaceSelector:
    matchNames:
      - default
  selector:
    matchLabels:
      app: order-service
  endpoints:
    - port: metrics
      interval: 15s
      path: /metrics
      scrapeTimeout: 10s

PART 4: ALERTING RULES

# custom-alerts.yaml:
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: microservice-alerts
  namespace: monitoring
spec:
  groups:
    - name: microservice.rules
      rules:
        # High error rate:
        - alert: HighErrorRate
          expr: |
            sum(rate(http_requests_total{status=~"5.."}[5m])) by (service)
            /
            sum(rate(http_requests_total[5m])) by (service)
            > 0.05
          for: 5m
          labels:
            severity: critical
          annotations:
            summary: "High error rate on {{ $labels.service }}"
            description: "Error rate > 5% for 5 minutes"

        # High latency:
        - alert: HighLatencyP99
          expr: |
            histogram_quantile(0.99,
              sum(rate(http_request_duration_seconds_bucket[5m])) by (le, service)
            ) > 2
          for: 5m
          labels:
            severity: warning
          annotations:
            summary: "P99 latency > 2s on {{ $labels.service }}"

        # Pod restarts:
        - alert: PodRestartLoop
          expr: rate(kube_pod_container_status_restarts_total[15m]) > 0
          for: 10m
          labels:
            severity: warning
          annotations:
            summary: "Pod {{ $labels.pod }} restart loop"

        # Disk pressure:
        - alert: NodeDiskPressure
          expr: |
            (node_filesystem_avail_bytes{mountpoint="/"} / 
             node_filesystem_size_bytes{mountpoint="/"}) < 0.1
          for: 5m
          labels:
            severity: critical
          annotations:
            summary: "Node {{ $labels.instance }} disk < 10%"

    - name: sla.rules
      rules:
        # SLO: 99.9% availability:
        - record: slo:availability:ratio
          expr: |
            1 - (
              sum(rate(http_requests_total{status=~"5.."}[30d]))
              /
              sum(rate(http_requests_total[30d]))
            )

        - alert: SLOBudgetBurning
          expr: slo:availability:ratio < 0.999
          for: 1h
          labels:
            severity: critical
          annotations:
            summary: "SLO budget burning: {{ $value | humanizePercentage }}"

PART 5: GRAFANA DASHBOARDS

# Essential Grafana dashboards to import:
# 1. K8s Cluster Overview:     ID: 315
# 2. Node Exporter Full:       ID: 1860  
# 3. K8s Pod Monitoring:       ID: 6417
# 4. NGINX Ingress:            ID: 9614
# 5. CoreDNS:                  ID: 5926
# 6. etcd:                     ID: 3070
# USE (Utilization, Saturation, Errors) Dashboard panels:
# 
# CPU Utilization:
# sum(rate(container_cpu_usage_seconds_total{namespace="default"}[5m])) by (pod)
#
# Memory Utilization:
# container_memory_working_set_bytes{namespace="default"} / 
# container_spec_memory_limit_bytes{namespace="default"}
#
# Network Errors:
# sum(rate(container_network_receive_errors_total[5m])) by (pod)
#
# Request Rate (RED method):
# sum(rate(http_requests_total{namespace="default"}[5m])) by (service)
#
# Error Rate:
# sum(rate(http_requests_total{status=~"5.."}[5m])) by (service)
# / sum(rate(http_requests_total[5m])) by (service)
#
# Duration:
# histogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket[5m])) by (le, service))

PART 6: THANOS (LONG-TERM STORAGE)

# Thanos sidecar cho Prometheus:
prometheus:
  prometheusSpec:
    thanos:
      image: quay.io/thanos/thanos:v0.35.0
      objectStorageConfig:
        existingSecret:
          name: thanos-objstore-config
          key: objstore.yml

# thanos-objstore-config:
apiVersion: v1
kind: Secret
metadata:
  name: thanos-objstore-config
  namespace: monitoring
stringData:
  objstore.yml: |
    type: S3
    config:
      bucket: thanos-metrics
      endpoint: ceph-rgw.storage:8080
      access_key: thanos
      secret_key: thanos-secret
      insecure: true

💡 KEY TAKEAWAYS

  1. kube-prometheus-stack: Complete monitoring in one Helm chart
  2. ServiceMonitor: Auto-discover targets, no manual scrape config
  3. RED method: Rate, Error, Duration for all services
  4. Alertmanager: Route alerts by severity to Slack/PagerDuty
  5. Recording rules: Pre-compute expensive queries
  6. Thanos: Long-term storage (months/years) on object storage

🎯 EXERCISES__HTMLTAG_132___

Exercise 1: Monitoring Setup__HTMLTAG_134___
  • Deploy kube-prometheus-stack
  • Create ServiceMonitor for app
  • Build RED method Grafana dashboard__HTMLTAG_141___

Exercise 2: Alerting

  • Create alert rules (error rate, latency, pod restart)
  • Configure Alertmanager → Slack
  • Trigger alert, verify notification

📚 NEXT POST

In Lesson 33: Loki — Centralized Logging, we will setup centralized log aggregation with Grafana Loki.