🎯 MỤC TIÊU BÀI HỌC
- ✅ Distributed tracing concepts (spans, traces, context propagation)
- ✅ Deploy Grafana Tempo trên K8s
- ✅ OpenTelemetry Collector và SDK instrumentation
- ✅ Trace → Log → Metric correlation
- ✅ Sampling strategies (head, tail, adaptive)
- ✅ TraceQL queries
PHẦN 1: DISTRIBUTED TRACING CONCEPTS
Distributed Trace Flow:
Client Request
│
▼
┌──────────┐ trace_id=abc123 ┌──────────────┐ ┌──────────────┐
│ API GW │────────────────►│ Order Service │─►│Payment Service│
│ span_id=1│ │ span_id=2 │ │ span_id=3 │
└──────────┘ └──────┬────────┘ └──────────────┘
│
┌──────▼────────┐
│ Inventory Svc │
│ span_id=4 │
└───────────────┘
Trace = collection of spans sharing trace_id
Span = single operation (HTTP call, DB query, etc.)
Context Propagation = passing trace_id between services
Headers:
traceparent: 00-abc123-span1-01
tracestate: tempo=true
| Feature | Tempo | Jaeger | Zipkin |
|---|---|---|---|
| Storage Backend | Object storage (S3/GCS) | Elasticsearch/Cassandra | Elasticsearch/MySQL |
| Cost | Very low | High (index everything) | Medium |
| Search | TraceQL (powerful) | Tag-based | Tag-based |
| Integration | Grafana native | Standalone UI | Standalone UI |
| Trace Discovery | Metrics → Traces | Manual search | Manual search |
| Protocol | OTLP, Jaeger, Zipkin | Jaeger, OTLP | Zipkin, OTLP |
PHẦN 2: DEPLOY GRAFANA TEMPO
# Install Tempo distributed:
helm install tempo grafana/tempo-distributed \
--namespace monitoring \
-f tempo-values.yaml
# tempo-values.yaml:
global:
clusterDomain: cluster.local
tempo:
storage:
trace:
backend: s3
s3:
bucket: tempo-traces
endpoint: ceph-rgw.storage:8080
access_key: tempo
secret_key: tempo-secret
insecure: true
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
jaeger:
protocols:
grpc:
endpoint: 0.0.0.0:14250
thrift_http:
endpoint: 0.0.0.0:14268
distributor:
replicas: 2
resources:
requests:
cpu: 100m
memory: 256Mi
limits:
cpu: 500m
memory: 512Mi
ingester:
replicas: 2
persistence:
enabled: true
storageClass: ceph-block
size: 10Gi
querier:
replicas: 2
queryFrontend:
replicas: 2
compactor:
replicas: 1
persistence:
enabled: true
storageClass: ceph-block
size: 10Gi
metricsGenerator:
enabled: true
replicas: 1
config:
storage:
remote_write:
- url: http://prometheus-kube-prometheus-prometheus.monitoring:9090/api/v1/write
PHẦN 3: OPENTELEMETRY COLLECTOR
# otel-collector-values.yaml:
apiVersion: opentelemetry.io/v1beta1
kind: OpenTelemetryCollector
metadata:
name: otel-collector
namespace: monitoring
spec:
mode: deployment
replicas: 2
config:
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
processors:
batch:
send_batch_size: 1000
timeout: 10s
memory_limiter:
check_interval: 1s
limit_mib: 512
spike_limit_mib: 128
tail_sampling:
decision_wait: 10s
policies:
# Always sample errors:
- name: error-policy
type: status_code
status_code:
status_codes: [ERROR]
# Always sample slow requests:
- name: latency-policy
type: latency
latency:
threshold_ms: 1000
# Sample 10% of normal requests:
- name: probabilistic-policy
type: probabilistic
probabilistic:
sampling_percentage: 10
exporters:
otlp/tempo:
endpoint: tempo-distributor.monitoring:4317
tls:
insecure: true
prometheus:
endpoint: 0.0.0.0:8889
resource_to_telemetry_conversion:
enabled: true
service:
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, tail_sampling, batch]
exporters: [otlp/tempo]
metrics:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [prometheus]
PHẦN 4: APPLICATION INSTRUMENTATION
// Go OpenTelemetry setup:
package main
import (
"context"
"go.opentelemetry.io/otel"
"go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc"
"go.opentelemetry.io/otel/sdk/resource"
sdktrace "go.opentelemetry.io/otel/sdk/trace"
semconv "go.opentelemetry.io/otel/semconv/v1.21.0"
)
func initTracer() (*sdktrace.TracerProvider, error) {
exporter, err := otlptracegrpc.New(context.Background(),
otlptracegrpc.WithEndpoint("otel-collector.monitoring:4317"),
otlptracegrpc.WithInsecure(),
)
if err != nil {
return nil, err
}
tp := sdktrace.NewTracerProvider(
sdktrace.WithBatcher(exporter),
sdktrace.WithResource(resource.NewWithAttributes(
semconv.SchemaURL,
semconv.ServiceName("order-service"),
semconv.ServiceVersion("1.0.0"),
semconv.DeploymentEnvironment("production"),
)),
)
otel.SetTracerProvider(tp)
return tp, nil
}
// Trace context in HTTP handler:
func CreateOrder(w http.ResponseWriter, r *http.Request) {
ctx, span := otel.Tracer("order-service").Start(r.Context(), "CreateOrder")
defer span.End()
span.SetAttributes(
attribute.String("order.id", orderID),
attribute.Int("order.items", len(items)),
)
// Call downstream service (context propagated automatically):
resp, err := httpClient.Do(req.WithContext(ctx))
if err != nil {
span.RecordError(err)
span.SetStatus(codes.Error, err.Error())
}
}
PHẦN 5: TRACEQL QUERIES
# TraceQL examples:
# Find traces by service name:
{ resource.service.name = "order-service" }
# Find error traces:
{ status = error }
# Find slow spans (> 1s):
{ duration > 1s }
# Find traces by HTTP route:
{ span.http.route = "/api/v1/orders" && status = error }
# Find traces with specific attribute:
{ span.order_id = "ORD-12345" }
# Complex: errors in payment service called from order service:
{ resource.service.name = "order-service" } >> { resource.service.name = "payment-service" && status = error }
PHẦN 6: TRACE-LOG-METRIC CORRELATION
# Grafana datasource: enable trace-to-log:
apiVersion: 1
datasources:
- name: Tempo
type: tempo
url: http://tempo-query-frontend.monitoring:3100
jsonData:
tracesToLogs:
datasourceUid: loki
filterByTraceID: true
filterBySpanID: true
tracesToMetrics:
datasourceUid: prometheus
queries:
- name: Request rate
query: sum(rate(http_server_request_duration_seconds_count{$$__tags}[5m]))
- name: Error rate
query: sum(rate(http_server_request_duration_seconds_count{$$__tags, http_status_code=~"5.."}[5m]))
serviceMap:
datasourceUid: prometheus
Correlation Flow:
Grafana Dashboard (Metrics)
│ "Error spike on order-service"
│ Click exemplar point
▼
Tempo (Traces)
│ trace_id = abc123
│ "order-service → payment-service TIMEOUT"
│ Click "View Logs"
▼
Loki (Logs)
│ {trace_id="abc123"}
│ "Connection refused to payment-service:8080"
▼
Root Cause: Payment service pod crashed (OOMKilled)
💡 KEY TAKEAWAYS
- Tempo: Trace storage on object storage → cost-efficient
- OpenTelemetry: Vendor-neutral instrumentation standard
- OTel Collector: Central pipeline, tail sampling
- TraceQL: Query traces by attributes, duration, status
- Correlation: Traces ↔ Logs ↔ Metrics = fast root cause
- Sampling: Always keep errors + slow, sample normal
🎯 BÀI TẬP
Bài tập 1: Tempo + OTel Setup
- Deploy Tempo + OTel Collector
- Instrument sample Go/Node.js app
- View traces in Grafana
Bài tập 2: Trace Correlation
- Configure trace-to-log linking in Grafana
- Inject errors, find root cause via trace → log flow
- Write TraceQL queries for slow/error traces
📚 BÀI TIẾP THEO
Trong Bài 35: Grafana Dashboards & SLO, chúng ta sẽ build unified dashboards và implement SLO/SLI monitoring.