🎯 Lesson Objective
Understand Loki as a log aggregation system that is lighter than Elasticsearch, how Grafana Alloy collects logs, Tempo for distributed tracing, OpenTelemetry auto-instrumentation, and how to combine Logs + Metrics + Traces in Grafana.
1. Loki — Log Aggregation
1.1 Loki vs Elasticsearch
- Loki: label-based indexing (only index labels, not index log content) → lighter, cheaper, suitable for Kubernetes logs
- Elasticsearch: full-text indexed → consumes a lot of memory/CPU, suitable when full-text search__HTMLTAG_81___ is needed
For Kubernetes logs (structured, labeled), Loki is the better choice in 2026.
1.2 Loki Architecture
- Distributor: receive logs from agents, validate and fan-out
- Ingester: buffer logs in memory, flush to object storage
- Querier: execute queries, merge results from ingester + storage
- Compactor: compact chunks, retention management
1.3 LogQL — Query Language
# Cơ bản: filter theo labels {app="nginx", namespace="production"}Filter log content
{app="nginx"} |= "error" {app="nginx"} != "health"
JSON parsing
{app="myapp"} | json | level="error"
Regex filter
{app="nginx"} |~ "HTTP/1\.1 [45][0-9]{2}"
Metrics từ logs
rate({app="nginx"} |= "error" [5m]) count_over_time({app="myapp"} | json | level="error" [1h])
Top N errors
topk(10, sum by (message) ( count_over_time({namespace="production"} | json | level="error" [1h]) ) )
2. Grafana Alloy — Unified Collector
Grafana Alloy replaces all individual agents: Promtail, OTel Collector, Prometheus agent mode.
# Alloy config (River DSL) # /etc/alloy/config.alloyThu thập logs từ Kubernetes pods
discovery.kubernetes "pods" { role = "pod" }
discovery.relabel "pod_logs" { targets = discovery.kubernetes.pods.targets rule { source_labels = ["__meta_kubernetes_pod_label_app"] target_label = "app" } rule { source_labels = ["__meta_kubernetes_namespace"] target_label = "namespace" } rule { source_labels = ["__meta_kubernetes_pod_container_name"] target_label = "container" } }
loki.source.kubernetes "pods" { targets = discovery.relabel.pod_logs.output forward_to = [loki.write.default.receiver] }
Forward đến Loki
loki.write "default" { endpoint { url = "http://loki:3100/loki/api/v1/push" } }
Thu thập Prometheus metrics
prometheus.scrape "kubernetes" { targets = discovery.kubernetes.pods.targets forward_to = [prometheus.remote_write.grafana_cloud.receiver] }
prometheus.remote_write "grafana_cloud" { endpoint { url = "http://prometheus:9090/api/v1/write" } }
3. Tempo — Distributed Tracing
Tempo is Grafana's distributed tracing backend, integrating well with Loki and Prometheus.
3.1 Concepts
- Trace: end-to-end journey of a request (for example, from browser to database)
- Span: an operation in the trace (eg: HTTP handler, DB query)
- TraceID: Unique ID to link all spans of a trace
- SpanID: ID of the individual span
3.2 Tempo Setup__HTMLTAG_136___
apiVersion: apps/v1
kind: Deployment
metadata:
name: tempo
namespace: monitoring
spec:
replicas: 1
selector:
matchLabels:
app: tempo
template:
metadata:
labels:
app: tempo
spec:
containers:
- name: tempo
image: grafana/tempo:2.7.0
args: ["-config.file=/etc/tempo.yaml"]
ports:
- containerPort: 3200 # Tempo HTTP API
- containerPort: 4317 # OTLP gRPC
- containerPort: 4318 # OTLP HTTP
volumeMounts:
- name: config
mountPath: /etc
- name: data
mountPath: /tmp/tempo
volumes:
- name: config
configMap:
name: tempo-config
- name: data
emptyDir: {}
apiVersion: apps/v1
kind: Deployment
metadata:
name: tempo
namespace: monitoring
spec:
replicas: 1
selector:
matchLabels:
app: tempo
template:
metadata:
labels:
app: tempo
spec:
containers:
- name: tempo
image: grafana/tempo:2.7.0
args: ["-config.file=/etc/tempo.yaml"]
ports:
- containerPort: 3200 # Tempo HTTP API
- containerPort: 4317 # OTLP gRPC
- containerPort: 4318 # OTLP HTTP
volumeMounts:
- name: config
mountPath: /etc
- name: data
mountPath: /tmp/tempo
volumes:
- name: config
configMap:
name: tempo-config
- name: data
emptyDir: {}
4. OpenTelemetry Auto-instrumentation
OpenTelemetry Operator automatically instruments applications without changing code.
4.1 Install Otel Operator
kubectl apply -f https://github.com/open-telemetry/opentelemetry-operator/releases/latest/download/opentelemetry-operator.yaml
4.2 Instrumentation CRD
apiVersion: opentelemetry.io/v1alpha1
kind: Instrumentation
metadata:
name: my-instrumentation
namespace: production
spec:
exporter:
endpoint: http://otel-collector:4317
propagators:
- tracecontext
- baggage
sampler:
type: parentbased_traceidratio
argument: "0.1" # sample 10% requests
java:
image: ghcr.io/open-telemetry/opentelemetry-operator/autoinstrumentation-java:latest
nodejs:
image: ghcr.io/open-telemetry/opentelemetry-operator/autoinstrumentation-nodejs:latest
python:
image: ghcr.io/open-telemetry/opentelemetry-operator/autoinstrumentation-python:latest
4.3 Annotate Pod to auto-instrument
apiVersion: apps/v1
kind: Deployment
metadata:
name: java-app
namespace: production
spec:
template:
metadata:
annotations:
# Bật auto-instrumentation cho Java
instrumentation.opentelemetry.io/inject-java: "true"
# Cho Node.js
# instrumentation.opentelemetry.io/inject-nodejs: "true"
# Cho Python
# instrumentation.opentelemetry.io/inject-python: "true"
spec:
containers:
- name: java-app
image: myapp:v1.2.3
# OTel Operator tự động inject JAVA_TOOL_OPTIONS, OTEL_SERVICE_NAME, etc.
5. Correlated Observability in Grafana
The real power is when you can go from Alert → Metrics → Logs → Traces within Grafana.
5.1 Derived Fields in Loki
// Grafana Data Source Loki config
{
"derivedFields": [
{
"name": "TraceID",
"matcherRegex": "traceID=(\\w+)",
"url": "${__value.raw}",
"datasourceUid": "tempo-uid" // link sang Tempo
}
]
}
// Khi log line có "traceID=abc123" → click link → mở trace trong Tempo
5.2 Exemplars — Link from Metrics to Traces
# Prometheus config để enable exemplars global: scrape_interval: 15s evaluation_interval: 15sApplication phải expose exemplars trong metrics
http_request_duration_seconds_bucket{le="0.1"} 24054 # {trace_id="abc123"} 0.092
5.3 Grafana Explore — Correlated View
# Workflow debug:
# 1. Xem alert → metric spike ở 15:32
# 2. Grafana Explore → chuyển sang Loki → lọc logs 15:30-15:35
# Query: {namespace="production", app="api"} |= "error" | json
# 3. Click traceID trong log line → mở Tempo
# 4. Xem trace → span nào chậm? → DB query 3.2s
6. Loki vs EFK Stack — When to use what?
- Using Loki: Kubernetes logs, structured logs (JSON), no need for full-text search, want low cost
- Using Elasticsearch (EFK): need full-text search through log content, compliance logging needs long retention, already has current ELK infrastructure__HTMLTAG_167___
Summary
- Loki: label-based, lightweight, suitable for Kubernetes logs — LogQL query language
- Grafana Alloy: unified collector (metrics + logs + traces)
- Tempo: distributed tracing, well integrated with Grafana ecosystem
- OTel Operator: auto-instrumentation without code change
- Correlated observability: from alert → metric → log → trace in Grafana