Chuyển đến nội dung chính

LESSON 29: PROMETHEUS AND GRAFANA

Prometheus Operator and kube-prometheus-stack. ServiceMonitor, PodMonitor, PrometheusRule. Grafana dashboards for Kubernetes clusters. AlertManager: routes, receivers (Slack, PagerDuty). Recording rules and PromQL best practices.

🔒 DevSecOps — Lesson 29 LESSON 29: PROMETHEUS AND GRAFANA

KUBERNETES: FROM BASIC TO ADVANCED

Module 7: Observability & Monitoring

xdev.asia

Prometheus and Grafana__HTMLTAG_66___

Prometheus and Grafana are an inseparable duo in the Kubernetes ecosystem. Prometheus is responsible for collecting, storing, and querying metrics, while Grafana provides a powerful visualization layer. With the introduction of Prometheus Operator, managing Prometheus on Kubernetes becomes declarative and fully automated.

Prometheus Data Model

Before going into Prometheus Operator, you need to understand Prometheus's data model to write effective queries.

Time Series and Labels__HTMLTAG_74___

All data in Prometheus is time series — a series of values (float64) over time, uniquely identified by a metric name and a set of key-value labels. For example:

http_requests_total{method="GET", status="200", job="api-server", instance="10.0.0.1:8080"}
http_requests_total{method="POST", status="500", job="api-server", instance="10.0.0.1:8080"}

Labels is the main tool to filter, aggregate and join data in PromQL. Good label design is important — don't use high-cardinality labels (like user ID or request ID) because it will create millions of time series and slow down Prometheus.

Metric Types

Prometheus defines four basic types of metrics:

  • Counter: value only increases, never decreases (reset to 0 when restarting). Used for: total requests, total errors, bytes sent. Query often used with rate() or increase().
  • Gauge: value can be increased or decreased freely. Used for: current memory usage, current number of pods, queue size.
  • Histogram: measures the distribution of observations (usually request duration, response size). Create time series with suffix _bucket, _sum, _count. Used to calculate percentiles with histogram_quantile().
  • Summary: similar to Histogram but calculates percentiles on the client-side. Less flexible than Histogram, not recommended for new metrics.

Prometheus Operator

Prometheus Operator helps manage Prometheus and AlertManager on Kubernetes in a declarative way. Instead of writing config files and reloading Prometheus manually, you create Kubernetes resources and the Operator will automatically update the configuration.

CRDs of Prometheus Operator

Prometheus Operator provides the following CRDs:

  • Prometheus: defines a Prometheus instance
  • AlertManager: AlertManager cluster definition
  • ServiceMonitor: defines how to scrape metrics from Services
  • PodMonitor: defines how to scrape metrics from Pods
  • PrometheusRule: defines alerting and recording rules
  • Probe: defines blackbox monitoring targets

Prometheus CRD

apiVersion: monitoring.coreos.com/v1
kind: Prometheus
metadata:
  name: prometheus
  namespace: monitoring
spec:
  replicas: 2
  retention: 15d
  serviceAccountName: prometheus
  serviceMonitorSelector:
    matchLabels:
      team: frontend
  serviceMonitorNamespaceSelector:
    matchLabels:
      monitoring: enabled
  ruleSelector:
    matchLabels:
      prometheus: kube-prometheus
  storage:
    volumeClaimTemplate:
      spec:
        storageClassName: fast-ssd
        resources:
          requests:
            storage: 100Gi
  resources:
    requests:
      memory: 2Gi
      cpu: 500m
    limits:
      memory: 4Gi
      cpu: 2000m

Prometheus Operator automatically discovers ServiceMonitors and PodMonitors based on label selectors — no need to restart or reload Prometheus when adding new targets.

ServiceMonitor — Scrape Metrics From Services

ServiceMonitor is the most popular way to add scrape targets to Prometheus. It defines how Prometheus finds and scrapes metrics from a group of Services.

apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
  name: api-server
  namespace: production
  labels:
    team: frontend
    app: api-server
spec:
  selector:
    matchLabels:
      app: api-server
  namespaceSelector:
    matchNames:
      - production
  endpoints:
    - port: metrics
      path: /metrics
      interval: 30s
      scrapeTimeout: 10s
      relabelings:
        - sourceLabels: [__meta_kubernetes_pod_name]
          targetLabel: pod
        - sourceLabels: [__meta_kubernetes_namespace]
          targetLabel: namespace

Service needs to expose metrics port with correct name:

apiVersion: v1
kind: Service
metadata:
  name: api-server
  namespace: production
  labels:
    app: api-server
spec:
  ports:
    - name: http
      port: 8080
    - name: metrics      # tên port phải match với ServiceMonitor
      port: 9090
  selector:
    app: api-server

PodMonitor — Scrape Metrics Directly From Pods

PodMonitor is used when you want to scrape directly from Pods without going through the Service, or when each Pod needs to be scraped independently (for example, each Pod exposes different metrics).

apiVersion: monitoring.coreos.com/v1
kind: PodMonitor
metadata:
  name: worker-pods
  namespace: production
spec:
  selector:
    matchLabels:
      app: worker
  podMetricsEndpoints:
    - port: metrics
      path: /metrics
      interval: 60s
  namespaceSelector:
    matchNames:
      - production
      - staging

PromQL — Prometheus Query Language

PromQL (Prometheus Query Language) is a powerful tool for querying time series data. Below are the most common patterns.

Rate and Increase

With Counter metrics, you always need to use rate() or increase() to get a meaningful value:

# Requests per second (5 minute rate)
rate(http_requests_total[5m])

# Total requests trong 1 giờ qua
increase(http_requests_total[1h])

# Error rate percentage
rate(http_requests_total{status=~"5.."}[5m]) / rate(http_requests_total[5m]) * 100

Histogram Percentiles__HTMLTAG_176___
# P95 request latency
histogram_quantile(0.95, rate(http_request_duration_seconds_bucket[5m]))

# P99 latency theo service
histogram_quantile(0.99,
  sum by (le, service) (
    rate(http_request_duration_seconds_bucket[5m])
  )
)

Kubernetes-Specific Queries

# Pods với unavailable replicas
kube_deployment_status_replicas_unavailable > 0

# Pods đang trong CrashLoopBackOff
kube_pod_container_status_waiting_reason{reason="CrashLoopBackOff"} == 1

# CPU throttling percentage
rate(container_cpu_cfs_throttled_seconds_total[5m]) /
rate(container_cpu_cfs_periods_total[5m]) * 100 > 25

# Memory usage percentage
container_memory_working_set_bytes /
container_spec_memory_limit_bytes * 100

# Node disk pressure
kube_node_status_condition{condition="DiskPressure", status="true"} == 1

Grafana Dashboards__HTMLTAG_180___

Grafana is a visualization layer, connecting with Prometheus (and Loki, Tempo) to create rich dashboards.

Import Dashboard From Grafana.com

Grafana.com has thousands of community dashboards. Some important dashboards for Kubernetes:

  • ID 315: Kubernetes cluster monitoring (basic)
  • ID 12740: Kubernetes monitoring (advanced, requires kube-state-metrics)
  • ID 15661: Kubernetes Node Overview
  • ID 15760: Kubernetes Views — Global
  • ID 14205: Kubernetes — Pod Overview

To import: go to Grafana UI → Dashboards → Import → enter ID → select Prometheus data source.

Dashboard Variables

Variables turns the dashboard into an interactive tool. Common variables for Kubernetes:

# Variable: cluster
Type: Query
Query: label_values(kube_node_info, cluster)

# Variable: namespace
Type: Query
Query: label_values(kube_namespace_labels{cluster="$cluster"}, namespace)

# Variable: pod
Type: Query
Query: label_values(kube_pod_info{cluster="$cluster", namespace="$namespace"}, pod)

With this variable, the user can select cluster → namespace → pod and every panel in the dashboard will automatically filter according to that selection.

Important Panels

The complete Kubernetes Dashboard should have:

  • Node Overview: CPU usage, memory usage, disk I/O, network I/O per node
  • Pod Metrics: CPU/memory request vs limit vs actual usage
  • Deployment Status: desired vs available replicas
  • Container Restarts: restart count in 1h, 24h
  • Error Rate: HTTP 5xx rate per service
  • Latency P50/P95/P99: request duration percentiles

AlertManager

AlertManager receives alerts from Prometheus and processes them: routing to the correct receiver, grouping to reduce noise, inhibition to avoid alert storms, and silencing during maintenance.

Routing Tree

AlertManager routing is configured in a hierarchical tree. Each route matches the labels of the alert and sends it to the corresponding receiver:

apiVersion: monitoring.coreos.com/v1alpha1
kind: AlertmanagerConfig
metadata:
  name: main-routing
  namespace: monitoring
spec:
  route:
    groupBy: ['alertname', 'cluster', 'service']
    groupWait: 30s
    groupInterval: 5m
    repeatInterval: 12h
    receiver: 'default-slack'
    routes:
      - match:
          severity: critical
        receiver: 'pagerduty-critical'
        continue: false
      - match:
          severity: warning
          team: frontend
        receiver: 'slack-frontend'
      - match:
          severity: warning
        receiver: 'slack-platform'
  receivers:
    - name: 'default-slack'
      slackConfigs:
        - apiURL:
            name: slack-secret
            key: webhook-url
          channel: '#alerts'
          title: '{{ .CommonAnnotations.summary }}'
          text: '{{ range .Alerts }}{{ .Annotations.description }}{{ end }}'
    - name: 'pagerduty-critical'
      pagerdutyConfigs:
        - routingKey:
            name: pagerduty-secret
            key: routing-key
          severity: '{{ .CommonLabels.severity }}'

Inhibition Rules__HTMLTAG_256___

Inhibition rules prevent alerts from being sent when another alert is active. For example, when the entire cluster is down, there is no need to send alerts for each service:

inhibitRules:
  - sourceMatch:
      alertname: ClusterDown
    targetMatch:
      severity: warning
    equal: ['cluster']
  - sourceMatch:
      alertname: NodeNotReady
    targetMatch:
      alertname: KubePodNotRunning
    equal: ['node']

PrometheusRule — Alert Rules__HTMLTAG_260___

PrometheusRule defines alerting rules and recording rules. Prometheus Operator automatically loads these rules into Prometheus.

apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: kubernetes-alerts
  namespace: monitoring
  labels:
    prometheus: kube-prometheus
    role: alert-rules
spec:
  groups:
    - name: kubernetes.deployment
      rules:
        - alert: KubeDeploymentReplicasMismatch
          expr: |
            (
              kube_deployment_spec_replicas
              !=
              kube_deployment_status_replicas_available
            ) and (
              changes(kube_deployment_status_replicas_updated[10m]) == 0
            )
          for: 15m
          labels:
            severity: warning
          annotations:
            summary: "Deployment {{ $labels.namespace }}/{{ $labels.deployment }} replica mismatch"
            description: "Deployment {{ $labels.namespace }}/{{ $labels.deployment }} has not matched the expected number of replicas for over 15 minutes."

        - alert: KubePodCrashLooping
          expr: |
            increase(kube_pod_container_status_restarts_total[1h]) > 5
          for: 2m
          labels:
            severity: warning
          annotations:
            summary: "Pod {{ $labels.namespace }}/{{ $labels.pod }} is crash looping"
            description: "Pod {{ $labels.namespace }}/{{ $labels.pod }} container {{ $labels.container }} has restarted {{ $value }} times in the last hour."

    - name: kubernetes.node
      rules:
        - alert: NodeMemoryPressure
          expr: kube_node_status_condition{condition="MemoryPressure", status="true"} == 1
          for: 5m
          labels:
            severity: critical
          annotations:
            summary: "Node {{ $labels.node }} is under memory pressure"

        - alert: NodeHighCPUUsage
          expr: |
            100 - (avg by (node) (
              rate(node_cpu_seconds_total{mode="idle"}[5m])
            ) * 100) > 85
          for: 10m
          labels:
            severity: warning
          annotations:
            summary: "High CPU usage on node {{ $labels.node }}: {{ $value | printf \"%.1f\" }}%"

Recording Rules — Pre-Compute Expensive Queries

Recording rules precompute complex queries and save the results as a new time series. This significantly improves the performance of dashboards and alerts that use heavy queries.

groups:
  - name: kubernetes.recording_rules
    interval: 1m
    rules:
      # Pre-compute request rate per service
      - record: job:http_requests_total:rate5m
        expr: |
          sum by (job, namespace, status) (
            rate(http_requests_total[5m])
          )

      # Pre-compute P99 latency per service
      - record: job:http_request_duration_seconds:p99
        expr: |
          histogram_quantile(0.99,
            sum by (le, job, namespace) (
              rate(http_request_duration_seconds_bucket[5m])
            )
          )

      # Pre-compute CPU usage ratio
      - record: namespace:container_cpu_usage_seconds_total:sum_rate
        expr: |
          sum by (namespace) (
            rate(container_cpu_usage_seconds_total{
              container!="",
              image!=""
            }[5m])
          )

Best practices for recording rules:

  • Name in format level:metric:operations (for example: job:http_requests:rate5m)
  • Only create recording rules for queries used in multiple places__HTMLTAG_277___
  • Evaluation interval of recording rules must be less than scrape interval
  • Do not create recording rules for single-use queries

Best Practices For Production

Prometheus Sizing

Prometheus memory usage is proportional to the number of active time series. Estimate: 1-2 bytes per sample, with 15-second scrape interval, 10,000 time series will use about 1GB of RAM. With a cluster of 100 nodes and hundreds of services, expected 500K-1M time series.

High Availability

Run 2 Prometheus instances with the same configuration. Grafana will deduplicate when querying. With AlertManager, run 3 instances in cluster mode to ensure alerts are not lost.

Long-Term Storage

Prometheus should only keep data for 2-4 weeks. For long-term storage (months/years), use Thanos or Grafana Mimir — both support object storage backends (S3, GCS) at a much lower cost than Prometheus local storage.

Summary

Prometheus Operator has revolutionized the way monitoring is managed on Kubernetes. With ServiceMonitor and PodMonitor, adding new targets is completely declarative and requires no manual intervention. PrometheusRule helps alert rules be managed as code (GitOps-friendly), and AlertManager with flexible routing tree ensures the right people receive the right alerts.

The next lesson will flesh out the rest of the observability stack: Loki for logs and Tempo for distributed tracing.