Prometheus and Grafana__HTMLTAG_66___
Prometheus and Grafana are an inseparable duo in the Kubernetes ecosystem. Prometheus is responsible for collecting, storing, and querying metrics, while Grafana provides a powerful visualization layer. With the introduction of Prometheus Operator, managing Prometheus on Kubernetes becomes declarative and fully automated.
Prometheus Data Model
Before going into Prometheus Operator, you need to understand Prometheus's data model to write effective queries.
Time Series and Labels__HTMLTAG_74___
All data in Prometheus is time series — a series of values (float64) over time, uniquely identified by a metric name and a set of key-value labels. For example:
http_requests_total{method="GET", status="200", job="api-server", instance="10.0.0.1:8080"}
http_requests_total{method="POST", status="500", job="api-server", instance="10.0.0.1:8080"}
Labels is the main tool to filter, aggregate and join data in PromQL. Good label design is important — don't use high-cardinality labels (like user ID or request ID) because it will create millions of time series and slow down Prometheus.
Metric Types
Prometheus defines four basic types of metrics:
- Counter: value only increases, never decreases (reset to 0 when restarting). Used for: total requests, total errors, bytes sent. Query often used with
rate()orincrease(). - Gauge: value can be increased or decreased freely. Used for: current memory usage, current number of pods, queue size.
- Histogram: measures the distribution of observations (usually request duration, response size). Create time series with suffix
_bucket,_sum,_count. Used to calculate percentiles withhistogram_quantile(). - Summary: similar to Histogram but calculates percentiles on the client-side. Less flexible than Histogram, not recommended for new metrics.
Prometheus Operator
Prometheus Operator helps manage Prometheus and AlertManager on Kubernetes in a declarative way. Instead of writing config files and reloading Prometheus manually, you create Kubernetes resources and the Operator will automatically update the configuration.
CRDs of Prometheus Operator
Prometheus Operator provides the following CRDs:
- Prometheus: defines a Prometheus instance
- AlertManager: AlertManager cluster definition
- ServiceMonitor: defines how to scrape metrics from Services
- PodMonitor: defines how to scrape metrics from Pods
- PrometheusRule: defines alerting and recording rules
- Probe: defines blackbox monitoring targets
Prometheus CRD
apiVersion: monitoring.coreos.com/v1
kind: Prometheus
metadata:
name: prometheus
namespace: monitoring
spec:
replicas: 2
retention: 15d
serviceAccountName: prometheus
serviceMonitorSelector:
matchLabels:
team: frontend
serviceMonitorNamespaceSelector:
matchLabels:
monitoring: enabled
ruleSelector:
matchLabels:
prometheus: kube-prometheus
storage:
volumeClaimTemplate:
spec:
storageClassName: fast-ssd
resources:
requests:
storage: 100Gi
resources:
requests:
memory: 2Gi
cpu: 500m
limits:
memory: 4Gi
cpu: 2000m
Prometheus Operator automatically discovers ServiceMonitors and PodMonitors based on label selectors — no need to restart or reload Prometheus when adding new targets.
ServiceMonitor — Scrape Metrics From Services
ServiceMonitor is the most popular way to add scrape targets to Prometheus. It defines how Prometheus finds and scrapes metrics from a group of Services.
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: api-server
namespace: production
labels:
team: frontend
app: api-server
spec:
selector:
matchLabels:
app: api-server
namespaceSelector:
matchNames:
- production
endpoints:
- port: metrics
path: /metrics
interval: 30s
scrapeTimeout: 10s
relabelings:
- sourceLabels: [__meta_kubernetes_pod_name]
targetLabel: pod
- sourceLabels: [__meta_kubernetes_namespace]
targetLabel: namespace
Service needs to expose metrics port with correct name:
apiVersion: v1
kind: Service
metadata:
name: api-server
namespace: production
labels:
app: api-server
spec:
ports:
- name: http
port: 8080
- name: metrics # tên port phải match với ServiceMonitor
port: 9090
selector:
app: api-server
PodMonitor — Scrape Metrics Directly From Pods
PodMonitor is used when you want to scrape directly from Pods without going through the Service, or when each Pod needs to be scraped independently (for example, each Pod exposes different metrics).
apiVersion: monitoring.coreos.com/v1
kind: PodMonitor
metadata:
name: worker-pods
namespace: production
spec:
selector:
matchLabels:
app: worker
podMetricsEndpoints:
- port: metrics
path: /metrics
interval: 60s
namespaceSelector:
matchNames:
- production
- staging
PromQL — Prometheus Query Language
PromQL (Prometheus Query Language) is a powerful tool for querying time series data. Below are the most common patterns.
Rate and Increase
With Counter metrics, you always need to use rate() or increase() to get a meaningful value:
# Requests per second (5 minute rate)
rate(http_requests_total[5m])
# Total requests trong 1 giờ qua
increase(http_requests_total[1h])
# Error rate percentage
rate(http_requests_total{status=~"5.."}[5m]) / rate(http_requests_total[5m]) * 100
Histogram Percentiles__HTMLTAG_176___
# P95 request latency
histogram_quantile(0.95, rate(http_request_duration_seconds_bucket[5m]))
# P99 latency theo service
histogram_quantile(0.99,
sum by (le, service) (
rate(http_request_duration_seconds_bucket[5m])
)
)
# P95 request latency
histogram_quantile(0.95, rate(http_request_duration_seconds_bucket[5m]))
# P99 latency theo service
histogram_quantile(0.99,
sum by (le, service) (
rate(http_request_duration_seconds_bucket[5m])
)
)
Kubernetes-Specific Queries
# Pods với unavailable replicas
kube_deployment_status_replicas_unavailable > 0
# Pods đang trong CrashLoopBackOff
kube_pod_container_status_waiting_reason{reason="CrashLoopBackOff"} == 1
# CPU throttling percentage
rate(container_cpu_cfs_throttled_seconds_total[5m]) /
rate(container_cpu_cfs_periods_total[5m]) * 100 > 25
# Memory usage percentage
container_memory_working_set_bytes /
container_spec_memory_limit_bytes * 100
# Node disk pressure
kube_node_status_condition{condition="DiskPressure", status="true"} == 1
Grafana Dashboards__HTMLTAG_180___
Grafana is a visualization layer, connecting with Prometheus (and Loki, Tempo) to create rich dashboards.
Import Dashboard From Grafana.com
Grafana.com has thousands of community dashboards. Some important dashboards for Kubernetes:
- ID 315: Kubernetes cluster monitoring (basic)
- ID 12740: Kubernetes monitoring (advanced, requires kube-state-metrics)
- ID 15661: Kubernetes Node Overview
- ID 15760: Kubernetes Views — Global
- ID 14205: Kubernetes — Pod Overview
To import: go to Grafana UI → Dashboards → Import → enter ID → select Prometheus data source.
Dashboard Variables
Variables turns the dashboard into an interactive tool. Common variables for Kubernetes:
# Variable: cluster
Type: Query
Query: label_values(kube_node_info, cluster)
# Variable: namespace
Type: Query
Query: label_values(kube_namespace_labels{cluster="$cluster"}, namespace)
# Variable: pod
Type: Query
Query: label_values(kube_pod_info{cluster="$cluster", namespace="$namespace"}, pod)
With this variable, the user can select cluster → namespace → pod and every panel in the dashboard will automatically filter according to that selection.
Important Panels
The complete Kubernetes Dashboard should have:
- Node Overview: CPU usage, memory usage, disk I/O, network I/O per node
- Pod Metrics: CPU/memory request vs limit vs actual usage
- Deployment Status: desired vs available replicas
- Container Restarts: restart count in 1h, 24h
- Error Rate: HTTP 5xx rate per service
- Latency P50/P95/P99: request duration percentiles
AlertManager
AlertManager receives alerts from Prometheus and processes them: routing to the correct receiver, grouping to reduce noise, inhibition to avoid alert storms, and silencing during maintenance.
Routing Tree
AlertManager routing is configured in a hierarchical tree. Each route matches the labels of the alert and sends it to the corresponding receiver:
apiVersion: monitoring.coreos.com/v1alpha1
kind: AlertmanagerConfig
metadata:
name: main-routing
namespace: monitoring
spec:
route:
groupBy: ['alertname', 'cluster', 'service']
groupWait: 30s
groupInterval: 5m
repeatInterval: 12h
receiver: 'default-slack'
routes:
- match:
severity: critical
receiver: 'pagerduty-critical'
continue: false
- match:
severity: warning
team: frontend
receiver: 'slack-frontend'
- match:
severity: warning
receiver: 'slack-platform'
receivers:
- name: 'default-slack'
slackConfigs:
- apiURL:
name: slack-secret
key: webhook-url
channel: '#alerts'
title: '{{ .CommonAnnotations.summary }}'
text: '{{ range .Alerts }}{{ .Annotations.description }}{{ end }}'
- name: 'pagerduty-critical'
pagerdutyConfigs:
- routingKey:
name: pagerduty-secret
key: routing-key
severity: '{{ .CommonLabels.severity }}'
Inhibition Rules__HTMLTAG_256___
Inhibition rules prevent alerts from being sent when another alert is active. For example, when the entire cluster is down, there is no need to send alerts for each service:
inhibitRules:
- sourceMatch:
alertname: ClusterDown
targetMatch:
severity: warning
equal: ['cluster']
- sourceMatch:
alertname: NodeNotReady
targetMatch:
alertname: KubePodNotRunning
equal: ['node']
PrometheusRule — Alert Rules__HTMLTAG_260___
PrometheusRule defines alerting rules and recording rules. Prometheus Operator automatically loads these rules into Prometheus.
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: kubernetes-alerts
namespace: monitoring
labels:
prometheus: kube-prometheus
role: alert-rules
spec:
groups:
- name: kubernetes.deployment
rules:
- alert: KubeDeploymentReplicasMismatch
expr: |
(
kube_deployment_spec_replicas
!=
kube_deployment_status_replicas_available
) and (
changes(kube_deployment_status_replicas_updated[10m]) == 0
)
for: 15m
labels:
severity: warning
annotations:
summary: "Deployment {{ $labels.namespace }}/{{ $labels.deployment }} replica mismatch"
description: "Deployment {{ $labels.namespace }}/{{ $labels.deployment }} has not matched the expected number of replicas for over 15 minutes."
- alert: KubePodCrashLooping
expr: |
increase(kube_pod_container_status_restarts_total[1h]) > 5
for: 2m
labels:
severity: warning
annotations:
summary: "Pod {{ $labels.namespace }}/{{ $labels.pod }} is crash looping"
description: "Pod {{ $labels.namespace }}/{{ $labels.pod }} container {{ $labels.container }} has restarted {{ $value }} times in the last hour."
- name: kubernetes.node
rules:
- alert: NodeMemoryPressure
expr: kube_node_status_condition{condition="MemoryPressure", status="true"} == 1
for: 5m
labels:
severity: critical
annotations:
summary: "Node {{ $labels.node }} is under memory pressure"
- alert: NodeHighCPUUsage
expr: |
100 - (avg by (node) (
rate(node_cpu_seconds_total{mode="idle"}[5m])
) * 100) > 85
for: 10m
labels:
severity: warning
annotations:
summary: "High CPU usage on node {{ $labels.node }}: {{ $value | printf \"%.1f\" }}%"
Recording Rules — Pre-Compute Expensive Queries
Recording rules precompute complex queries and save the results as a new time series. This significantly improves the performance of dashboards and alerts that use heavy queries.
groups:
- name: kubernetes.recording_rules
interval: 1m
rules:
# Pre-compute request rate per service
- record: job:http_requests_total:rate5m
expr: |
sum by (job, namespace, status) (
rate(http_requests_total[5m])
)
# Pre-compute P99 latency per service
- record: job:http_request_duration_seconds:p99
expr: |
histogram_quantile(0.99,
sum by (le, job, namespace) (
rate(http_request_duration_seconds_bucket[5m])
)
)
# Pre-compute CPU usage ratio
- record: namespace:container_cpu_usage_seconds_total:sum_rate
expr: |
sum by (namespace) (
rate(container_cpu_usage_seconds_total{
container!="",
image!=""
}[5m])
)
Best practices for recording rules:
- Name in format
level:metric:operations(for example:job:http_requests:rate5m) - Only create recording rules for queries used in multiple places__HTMLTAG_277___
- Evaluation interval of recording rules must be less than scrape interval
- Do not create recording rules for single-use queries
Best Practices For Production
Prometheus Sizing
Prometheus memory usage is proportional to the number of active time series. Estimate: 1-2 bytes per sample, with 15-second scrape interval, 10,000 time series will use about 1GB of RAM. With a cluster of 100 nodes and hundreds of services, expected 500K-1M time series.
High Availability
Run 2 Prometheus instances with the same configuration. Grafana will deduplicate when querying. With AlertManager, run 3 instances in cluster mode to ensure alerts are not lost.
Long-Term Storage
Prometheus should only keep data for 2-4 weeks. For long-term storage (months/years), use Thanos or Grafana Mimir — both support object storage backends (S3, GCS) at a much lower cost than Prometheus local storage.
Summary
Prometheus Operator has revolutionized the way monitoring is managed on Kubernetes. With ServiceMonitor and PodMonitor, adding new targets is completely declarative and requires no manual intervention. PrometheusRule helps alert rules be managed as code (GitOps-friendly), and AlertManager with flexible routing tree ensures the right people receive the right alerts.
The next lesson will flesh out the rest of the observability stack: Loki for logs and Tempo for distributed tracing.