Chuyển đến nội dung chính

LESSON 30: LOKI, TEMPO AND DISTRIBUTED TRACING

Loki log aggregation with LogQL queries. Grafana Alloy collects logs from containers. Tempo distributed tracing. OpenTelemetry auto-instrumentation. Correlated observability: see logs, metrics, traces together in Grafana.

🔒 DevSecOps — Lesson 30 LESSON 30: LOKI, TEMPO AND DISTRIBUTED TRACING

KUBERNETES: FROM BASIC TO ADVANCED

Module 7: Observability & Monitoring

xdev.asia

🎯 Lesson Objective

Understand Loki as a log aggregation system that is lighter than Elasticsearch, how Grafana Alloy collects logs, Tempo for distributed tracing, OpenTelemetry auto-instrumentation, and how to combine Logs + Metrics + Traces in Grafana.

1. Loki — Log Aggregation

1.1 Loki vs Elasticsearch

  • Loki: label-based indexing (only index labels, not index log content) → lighter, cheaper, suitable for Kubernetes logs
  • Elasticsearch: full-text indexed → consumes a lot of memory/CPU, suitable when full-text search__HTMLTAG_81___ is needed

For Kubernetes logs (structured, labeled), Loki is the better choice in 2026.

1.2 Loki Architecture

  • Distributor: receive logs from agents, validate and fan-out
  • Ingester: buffer logs in memory, flush to object storage
  • Querier: execute queries, merge results from ingester + storage
  • Compactor: compact chunks, retention management

1.3 LogQL — Query Language

# Cơ bản: filter theo labels
{app="nginx", namespace="production"}

Filter log content

{app="nginx"} |= "error" {app="nginx"} != "health"

JSON parsing

{app="myapp"} | json | level="error"

Regex filter

{app="nginx"} |~ "HTTP/1\.1 [45][0-9]{2}"

Metrics từ logs

rate({app="nginx"} |= "error" [5m]) count_over_time({app="myapp"} | json | level="error" [1h])

Top N errors

topk(10, sum by (message) ( count_over_time({namespace="production"} | json | level="error" [1h]) ) )

2. Grafana Alloy — Unified Collector

Grafana Alloy replaces all individual agents: Promtail, OTel Collector, Prometheus agent mode.

# Alloy config (River DSL)
# /etc/alloy/config.alloy

Thu thập logs từ Kubernetes pods

discovery.kubernetes "pods" { role = "pod" }

discovery.relabel "pod_logs" { targets = discovery.kubernetes.pods.targets rule { source_labels = ["__meta_kubernetes_pod_label_app"] target_label = "app" } rule { source_labels = ["__meta_kubernetes_namespace"] target_label = "namespace" } rule { source_labels = ["__meta_kubernetes_pod_container_name"] target_label = "container" } }

loki.source.kubernetes "pods" { targets = discovery.relabel.pod_logs.output forward_to = [loki.write.default.receiver] }

Forward đến Loki

loki.write "default" { endpoint { url = "http://loki:3100/loki/api/v1/push" } }

Thu thập Prometheus metrics

prometheus.scrape "kubernetes" { targets = discovery.kubernetes.pods.targets forward_to = [prometheus.remote_write.grafana_cloud.receiver] }

prometheus.remote_write "grafana_cloud" { endpoint { url = "http://prometheus:9090/api/v1/write" } }

3. Tempo — Distributed Tracing

Tempo is Grafana's distributed tracing backend, integrating well with Loki and Prometheus.

3.1 Concepts

  • Trace: end-to-end journey of a request (for example, from browser to database)
  • Span: an operation in the trace (eg: HTTP handler, DB query)
  • TraceID: Unique ID to link all spans of a trace
  • SpanID: ID of the individual span

3.2 Tempo Setup__HTMLTAG_136___
apiVersion: apps/v1
kind: Deployment
metadata:
  name: tempo
  namespace: monitoring
spec:
  replicas: 1
  selector:
    matchLabels:
      app: tempo
  template:
    metadata:
      labels:
        app: tempo
    spec:
      containers:
      - name: tempo
        image: grafana/tempo:2.7.0
        args: ["-config.file=/etc/tempo.yaml"]
        ports:
        - containerPort: 3200   # Tempo HTTP API
        - containerPort: 4317   # OTLP gRPC
        - containerPort: 4318   # OTLP HTTP
        volumeMounts:
        - name: config
          mountPath: /etc
        - name: data
          mountPath: /tmp/tempo
      volumes:
      - name: config
        configMap:
          name: tempo-config
      - name: data
        emptyDir: {}

4. OpenTelemetry Auto-instrumentation

OpenTelemetry Operator automatically instruments applications without changing code.

4.1 Install Otel Operator

kubectl apply -f https://github.com/open-telemetry/opentelemetry-operator/releases/latest/download/opentelemetry-operator.yaml

4.2 Instrumentation CRD

apiVersion: opentelemetry.io/v1alpha1
kind: Instrumentation
metadata:
  name: my-instrumentation
  namespace: production
spec:
  exporter:
    endpoint: http://otel-collector:4317
  propagators:
    - tracecontext
    - baggage
  sampler:
    type: parentbased_traceidratio
    argument: "0.1"   # sample 10% requests
  java:
    image: ghcr.io/open-telemetry/opentelemetry-operator/autoinstrumentation-java:latest
  nodejs:
    image: ghcr.io/open-telemetry/opentelemetry-operator/autoinstrumentation-nodejs:latest
  python:
    image: ghcr.io/open-telemetry/opentelemetry-operator/autoinstrumentation-python:latest

4.3 Annotate Pod to auto-instrument

apiVersion: apps/v1
kind: Deployment
metadata:
  name: java-app
  namespace: production
spec:
  template:
    metadata:
      annotations:
        # Bật auto-instrumentation cho Java
        instrumentation.opentelemetry.io/inject-java: "true"
        # Cho Node.js
        # instrumentation.opentelemetry.io/inject-nodejs: "true"
        # Cho Python
        # instrumentation.opentelemetry.io/inject-python: "true"
    spec:
      containers:
      - name: java-app
        image: myapp:v1.2.3
        # OTel Operator tự động inject JAVA_TOOL_OPTIONS, OTEL_SERVICE_NAME, etc.

5. Correlated Observability in Grafana

The real power is when you can go from Alert → Metrics → Logs → Traces within Grafana.

5.1 Derived Fields in Loki

// Grafana Data Source Loki config
{
  "derivedFields": [
    {
      "name": "TraceID",
      "matcherRegex": "traceID=(\\w+)",
      "url": "${__value.raw}",
      "datasourceUid": "tempo-uid"   // link sang Tempo
    }
  ]
}
// Khi log line có "traceID=abc123" → click link → mở trace trong Tempo

5.2 Exemplars — Link from Metrics to Traces

# Prometheus config để enable exemplars
global:
  scrape_interval: 15s
  evaluation_interval: 15s

Application phải expose exemplars trong metrics

http_request_duration_seconds_bucket{le="0.1"} 24054 # {trace_id="abc123"} 0.092

5.3 Grafana Explore — Correlated View

# Workflow debug:
# 1. Xem alert → metric spike ở 15:32
# 2. Grafana Explore → chuyển sang Loki → lọc logs 15:30-15:35
#    Query: {namespace="production", app="api"} |= "error" | json
# 3. Click traceID trong log line → mở Tempo
# 4. Xem trace → span nào chậm? → DB query 3.2s

6. Loki vs EFK Stack — When to use what?

  • Using Loki: Kubernetes logs, structured logs (JSON), no need for full-text search, want low cost
  • Using Elasticsearch (EFK): need full-text search through log content, compliance logging needs long retention, already has current ELK infrastructure__HTMLTAG_167___

Summary

  • Loki: label-based, lightweight, suitable for Kubernetes logs — LogQL query language
  • Grafana Alloy: unified collector (metrics + logs + traces)
  • Tempo: distributed tracing, well integrated with Grafana ecosystem
  • OTel Operator: auto-instrumentation without code change
  • Correlated observability: from alert → metric → log → trace in Grafana