🎯 练习课的目的
- 使用 kube-prometheus-stack + Loki + Tempo 部署完整的 PLG 堆疊
- 使用 Grafana Alloy 收集日誌
- 使用 OpenTelemetry Operator 的自動儀器應用
- 建立警報規則和 Slack 通知
- 使用相關可觀測性進行調試
實驗 1:部署 kube-prometheus-stack
kubectl create namespace monitoringCài kube-prometheus-stack (Prometheus + Grafana + AlertManager)
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts helm repo update
cat > prom-values.yaml <<EOF grafana: adminPassword: "admin123" persistence: enabled: false service: type: NodePort nodePort: 32000
prometheus: prometheusSpec: retention: 7d resources: requests: memory: 512Mi limits: memory: 1Gi
alertmanager: alertmanagerSpec: resources: requests: memory: 64Mi EOF
helm install kube-prometheus prometheus-community/kube-prometheus-stack
--namespace monitoring
--values prom-values.yamlkubectl rollout status deployment/kube-prometheus-grafana -n monitoring
Truy cập Grafana
NODE_IP=$(kubectl get nodes -o jsonpath='{.items[0].status.addresses[0].address}') echo "Grafana: http://$NODE_IP:32000 (admin/admin123)"
實驗 2:為自訂應用程式建立 ServiceMonitor
kubectl create namespace lab7Deploy app expose Prometheus metrics
cat <<EOF | kubectl apply -f - apiVersion: apps/v1 kind: Deployment metadata: name: metrics-app namespace: lab7 labels: app: metrics-app spec: replicas: 2 selector: matchLabels: app: metrics-app template: metadata: labels: app: metrics-app spec: containers: - name: app # image này expose /metrics endpoint image: prom/prometheus:v2.49.1 args: ["--web.listen-address=:9090", "--config.file=/etc/prometheus/prometheus.yml"] ports: - containerPort: 9090 name: metrics resources: requests: cpu: "100m" memory: "128Mi"
apiVersion: v1 kind: Service metadata: name: metrics-app namespace: lab7 labels: app: metrics-app # Label này quan trọng cho ServiceMonitor spec: selector: app: metrics-app ports:
- name: metrics port: 9090 targetPort: 9090
apiVersion: monitoring.coreos.com/v1 kind: ServiceMonitor metadata: name: metrics-app namespace: lab7 labels: release: kube-prometheus # phải match Prometheus selector spec: namespaceSelector: matchNames: - lab7 selector: matchLabels: app: metrics-app endpoints:
- port: metrics interval: 30s path: /metrics EOF
Verify trong Prometheus UI
Status → Targets → metrics-app
實驗 3:AlertManager — Slack 通知
# Cấu hình AlertManager với Slack
# (Cần Slack webhook URL, dùng dummy URL cho lab)
cat > alertmanager-config.yaml <<EOF
apiVersion: v1
kind: Secret
metadata:
name: alertmanager-kube-prometheus-alertmanager
namespace: monitoring
stringData:
alertmanager.yaml: |
global:
slack_api_url: 'https://hooks.slack.com/services/YOUR/SLACK/WEBHOOK'
resolve_timeout: 5m
route:
group_by: ['job', 'alertname', 'namespace']
group_wait: 30s
group_interval: 5m
repeat_interval: 12h
receiver: 'slack-alerts'
routes:
- match:
severity: critical
receiver: 'slack-critical'
receivers:
- name: 'slack-alerts'
slack_configs:
- channel: '#k8s-alerts'
send_resolved: true
title: '[{{ .Status | toUpper }}] {{ .CommonLabels.alertname }}'
text: |
{{ range .Alerts }}
*Alert:* {{ .Labels.alertname }}
*Severity:* {{ .Labels.severity }}
*Namespace:* {{ .Labels.namespace }}
*Description:* {{ .Annotations.description }}
{{ end }}
- name: 'slack-critical'
slack_configs:
- channel: '#k8s-critical'
send_resolved: true
EOF
kubectl apply -f alertmanager-config.yaml
Tạo PrometheusRule
cat <<EOF | kubectl apply -f -
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: custom-alerts
namespace: lab7
labels:
release: kube-prometheus
spec:
groups:
name: kubernetes.rules rules:- alert: PodCrashLooping expr: rate(kube_pod_container_status_restarts_total{namespace="lab7"}[5m]) > 0 for: 2m labels: severity: warning annotations: description: "Pod {{ $labels.pod }} in {{ $labels.namespace }} is crash looping"
alert: HighMemoryUsage expr: | (container_memory_working_set_bytes{namespace="lab7"} / container_spec_memory_limit_bytes{namespace="lab7"}) * 100 > 80 for: 5m labels: severity: warning annotations: description: "Container {{ $labels.container }} memory > 80%" EOF
實驗室 4:部署 Loki + Grafana 合金
helm repo add grafana https://grafana.github.io/helm-charts helm repo updateCài Loki
helm install loki grafana/loki
--namespace monitoring
--set loki.auth_enabled=false
--set loki.commonConfig.replication_factor=1
--set loki.storage.type=filesystem
--set singleBinary.replicas=1Cài Grafana Alloy (thay Promtail)
cat > alloy-values.yaml <<EOF alloy: configMap: create: true content: | // Thu thập Kubernetes logs discovery.kubernetes "pods" { role = "pod" }
discovery.relabel "pod_logs" { targets = discovery.kubernetes.pods.targets rule { source_labels = ["__meta_kubernetes_namespace"] target_label = "namespace" } rule { source_labels = ["__meta_kubernetes_pod_label_app"] target_label = "app" } rule { source_labels = ["__meta_kubernetes_pod_name"] target_label = "pod" } rule { source_labels = ["__meta_kubernetes_pod_container_name"] target_label = "container" } } loki.source.kubernetes "pods" { targets = discovery.relabel.pod_logs.output forward_to = [loki.write.default.receiver] } loki.write "default" { endpoint { url = "http://loki:3100/loki/api/v1/push" } }EOF
helm install alloy grafana/alloy
--namespace monitoring
--values alloy-values.yamlVerify Alloy đang chạy
kubectl get pods -n monitoring -l app.kubernetes.io/name=alloy
Thêm Loki datasource vào Grafana
Grafana → Configuration → Data Sources → Add Loki
URL: http://loki:3100
實驗 5:寫 LogQL 查詢
# Mở Grafana Explore → Loki datasourceQuery 1: Tất cả logs từ namespace lab7
{namespace="lab7"}
Query 2: Chỉ error logs
{namespace="lab7"} |= "error"
Query 3: JSON logs với level filter
{namespace="monitoring"} | json | level="error"
Query 4: Rate của error logs
rate({namespace="lab7"} |= "error" [5m])
Query 5: Top errors
topk(5, sum by (pod) (count_over_time({namespace="lab7"} |= "error" [10m])))
實驗 6:部署 Tempo 和 OTel 自動儀表
# Cài Tempo helm install tempo grafana/tempo \ --namespace monitoring \ --set tempo.storage.trace.backend=localCài OpenTelemetry Operator
kubectl apply -f https://github.com/open-telemetry/opentelemetry-operator/releases/latest/download/opentelemetry-operator.yaml
Tạo Instrumentation
cat <<EOF | kubectl apply -f - apiVersion: opentelemetry.io/v1alpha1 kind: Instrumentation metadata: name: auto-instrumentation namespace: lab7 spec: exporter: endpoint: http://tempo-collector.monitoring:4317 propagators:
- tracecontext
- baggage sampler: type: parentbased_traceidratio argument: "1.0" # 100% sample cho demo nodejs: image: ghcr.io/open-telemetry/opentelemetry-operator/autoinstrumentation-nodejs:latest EOF
Deploy Node.js app với auto-instrumentation
cat <<EOF | kubectl apply -f - apiVersion: apps/v1 kind: Deployment metadata: name: node-app namespace: lab7 spec: replicas: 1 selector: matchLabels: app: node-app template: metadata: labels: app: node-app annotations: instrumentation.opentelemetry.io/inject-nodejs: "true" spec: containers: - name: app image: node:20-alpine command: ['node', '-e', ' const http = require("http"); const server = http.createServer((req, res) => { console.log(JSON.stringify({level: "info", msg: "request", path: req.url})); res.end("Hello K8s Observability!"); }); server.listen(3000); '] ports: - containerPort: 3000 EOF
Kiểm tra OTel injected env vars
kubectl exec -n lab7 deploy/node-app -- env | grep OTEL
實驗 7:相關儀表板
# Thêm Tempo datasource vào Grafana # URL: http://tempo:3100Cấu hình derived fields trong Loki datasource:
Name: TraceID
Regex: traceID=(\w+)
URL: ${__value.raw}
Internal link → Tempo datasource
Tạo requests để generate traces
for i in $(seq 1 20); do kubectl exec -n lab7 deploy/node-app -- wget -qO- http://localhost:3000 done
Grafana Explore:
1. Query Loki: {namespace="lab7", app="node-app"}
2. Click TraceID link trong log → mở Tempo
3. Xem trace spans
清理
kubectl delete namespace lab7
helm uninstall loki alloy tempo -n monitoring
helm uninstall kube-prometheus -n monitoring
kubectl delete namespace monitoring
總結
- ✅ kube-prometheus-stack:Prometheus + Grafana + AlertManager
- ✅ 用於自訂應用程式指標的 ServiceMonitor
- ✅ AlertManager:路由+Slack 通知
- ✅ Loki + Grafana Alloy:日誌聚合
- ✅ LogQL 查詢範圍從簡單到複雜
- ✅ Tempo + OTel 自動偵測
- ✅ 相關可觀察性:日誌 ↔ 痕跡