Chuyển đến nội dung chính

Lesson 16: Logging — Structured Logging, Loki & ELK Stack

Structured logging best practices, log levels strategy, Fluent Bit log collection, Loki vs Elasticsearch, LogQL, log correlation with traceId, log retention policies and cost optimization.

🏗️ Architecture — Lesson 16 Lesson 16: Logging — Structured Logging, Loki & ELK Stack

Cloud Native Microservices Architecture

Part 5: Observability — Three pillars

xdev.asia

Lesson 16: Logging — Structured Logging, Loki & ELK Stack

Introduction

In monolith, you can ssh Go to the server and tail -f log files. In microservices with dozens of services running on hundreds of dynamic pods, that is no longer possible. Centralized logging is a mandatory requirement.

But simply collecting logs is not enough — logs must be structured, queryable, and linked to traces for effective debugging.


1. Structured Logging

1.1 Why Structured Logging?

Unstructured log — difficult to parse automatically:

2026-03-31 10:15:30 INFO Order O-001 created for customer C-042, total 500000 VND, 3 items, took 45ms

Structured log (JSON) — machine-readable, queryable:

{
  "timestamp": "2026-03-31T10:15:30.123Z",
  "level": "INFO",
  "service": "order-service",
  "version": "1.2.3",
  "traceId": "abc123def456",
  "spanId": "span789abc",
  "message": "Order created successfully",
  "orderId": "O-001",
  "customerId": "C-042",
  "totalAmount": 500000,
  "currency": "VND",
  "itemCount": 3,
  "durationMs": 45
}

1.2 Log Levels Strategy

TRACE — Rất chi tiết, chỉ dùng khi debug cụ thể
        Không bao giờ enable ở production

DEBUG — Thông tin debug (function calls, variable values)
        Chỉ enable ở development, có thể bật tạm ở staging

INFO  — Sự kiện business quan trọng (order created, payment processed)
        Enable ở production — đây là log chính cần thu thập

WARN  — Tình huống không mong đợi nhưng system vẫn hoạt động
        (deprecated API call, slow query > 500ms, retry attempt)

ERROR — Lỗi cần xử lý nhưng service vẫn chạy
        (external service timeout, validation failure)

FATAL — Lỗi nghiêm trọng, service sắp shutdown
        (database connection lost, out of memory)

Principles:

  • INFO log is the "audit trail" — every important action needs an INFO log
  • ERROR log must include a stack trace and enough context to debug without needing additional information
  • Do not log sensitive data (password, credit card, PII)

1.3 Structured Logging Implementation

Java (Spring Boot with Logback + Logstash encoder):

<!-- logback-spring.xml -->
<configuration>
  <appender name="JSON" class="ch.qos.logback.core.ConsoleAppender">
    <encoder class="net.logstash.logback.encoder.LogstashEncoder">
      <customFields>{"service":"order-service","version":"${APP_VERSION}"}</customFields>
    </encoder>
  </appender>

  <root level="INFO">
    <appender-ref ref="JSON"/>
  </root>
</configuration>
// Sử dụng MDC (Mapped Diagnostic Context) cho correlation
import org.slf4j.MDC;

@Component
public class RequestLoggingFilter implements Filter {
    @Override
    public void doFilter(ServletRequest req, ServletResponse res, FilterChain chain) {
        MDC.put("traceId", extractOrGenerateTraceId(request));
        MDC.put("userId", extractUserId(request));
        try {
            chain.doFilter(req, res);
        } finally {
            MDC.clear();
        }
    }
}

// Trong service code
@Slf4j
public class OrderService {
    public Order createOrder(CreateOrderRequest req) {
        log.info("Creating order", // message
            kv("customerId", req.getCustomerId()),
            kv("itemCount", req.getItems().size()),
            kv("totalAmount", req.getTotal())
        );
        // ...
    }
}

Node.js (Pino):

import pino from 'pino';

const logger = pino({
  level: process.env.LOG_LEVEL || 'info',
  formatters: {
    level: (label) => ({ level: label }),
  },
  base: {
    service: 'order-service',
    version: process.env.APP_VERSION,
  },
});

// Usage
logger.info({ orderId, customerId, totalAmount }, 'Order created successfully');
logger.error({ err, orderId }, 'Failed to process order');

2. Log Collection Architecture

2.1 Pipeline Overview

┌──────────────────────────────────────────────────────────────┐
│                    Kubernetes Cluster                        │
│                                                              │
│  ┌──────────────┐  ┌──────────────┐  ┌──────────────┐      │
│  │ Order Service│  │Payment Svc   │  │ Inventory Svc│      │
│  │ (Pod)        │  │ (Pod)        │  │ (Pod)        │      │
│  │ stdout/stderr│  │ stdout/stderr│  │ stdout/stderr│      │
│  └──────┬───────┘  └──────┬───────┘  └──────┬───────┘      │
│         │                 │                  │              │
│         └─────────────────┴──────────────────┘              │
│                           │                                 │
│              ┌────────────▼──────────────┐                  │
│              │   Fluent Bit (DaemonSet)  │                  │
│              │  - Tail /var/log/pods/    │                  │
│              │  - Parse JSON             │                  │
│              │  - Enrich (node, pod)     │                  │
│              │  - Buffer & retry         │                  │
│              └────────────┬──────────────┘                  │
└───────────────────────────┼─────────────────────────────────┘
                            │
               ┌────────────▼────────────┐
               │    Log Aggregation      │
               │   Loki / Elasticsearch  │
               └────────────┬────────────┘
                            │
               ┌────────────▼────────────┐
               │      Visualization      │
               │   Grafana / Kibana      │
               └─────────────────────────┘

2.2 Fluent Bit Configuration

# fluent-bit ConfigMap
apiVersion: v1
kind: ConfigMap
metadata:
  name: fluent-bit-config
  namespace: monitoring
data:
  fluent-bit.conf: |
    [SERVICE]
        Flush         5
        Daemon        Off
        Log_Level     info

    [INPUT]
        Name              tail
        Tag               kube.*
        Path              /var/log/containers/*.log
        Parser            docker
        DB                /run/fluent-bit/flb_kube.db
        Mem_Buf_Limit     10MB
        Skip_Long_Lines   On
        Refresh_Interval  10

    [FILTER]
        Name                kubernetes
        Match               kube.*
        Kube_URL            https://kubernetes.default.svc:443
        Merge_Log           On
        K8S-Logging.Parser  On
        K8S-Logging.Exclude On

    [FILTER]
        Name    grep
        Match   kube.*
        # Loại bỏ health check logs
        Exclude log /health

    [OUTPUT]
        Name          loki
        Match         kube.*
        Host          loki.monitoring.svc
        Port          3100
        Labels        job=fluentbit, node=${NODE_NAME}
        Label_keys    $kubernetes['namespace_name'],$kubernetes['pod_name'],$kubernetes['container_name']
        Remove_keys   kubernetes,stream
        Auto_Kubernetes_Labels  On

3. Loki — Log Aggregation

3.1 Loki vs Elasticsearch

LokiElasticsearch
IndexingOnly index labels (metadata)Full-text index of the entire log
StorageMuch cheaper (S3/MinIO)Consumes storage and memory
QueryLogQL (simple, label-focused)Lucene/KQL (powerful full-text)
SetupSimpleComplex (clusters, shards)
Use casesCloud native, cost-sensitiveCompliance, full-text search
IntegrationGrafana nativeKibana

When to choose Loki: Most cases cloud native — low cost, native integration with Grafana, good enough for operational logs.

When to choose Elasticsearch: Need compliance/audit log search, full-text search in log content, already has Kibana ecosystem.

3.2 Loki Architecture

┌──────────────────────────────────────────────────┐
│                   Loki                           │
│                                                  │
│  ┌─────────────┐  ┌─────────────────────────┐   │
│  │  Distributor│──▶│   Ingester (in-memory)  │   │
│  │  (receive)  │  │                         │   │
│  └─────────────┘  └───────────┬─────────────┘   │
│                               │ flush            │
│  ┌─────────────┐  ┌───────────▼─────────────┐   │
│  │  Querier    │  │   Object Storage        │   │
│  │  (read)     │  │   (S3/MinIO/GCS)        │   │
│  └──────┬──────┘  └─────────────────────────┘   │
│         │                                        │
└─────────┼────────────────────────────────────────┘
          │ LogQL
          ▼
      Grafana

3.3 Install Loki and Grafana

helm repo add grafana https://grafana.github.io/helm-charts

# Cài Loki (simple scalable mode)
helm upgrade --install loki grafana/loki \
  --namespace monitoring \
  --set loki.storage.type=s3 \
  --set loki.storage.s3.bucket=loki-logs \
  --set loki.storage.s3.region=ap-southeast-1

# Cài Grafana Alloy (Fluent Bit alternative từ Grafana)
helm upgrade --install alloy grafana/alloy \
  --namespace monitoring

4. LogQL — Log Query Language

4.1 Stream Selector

Select log streams by labels:

# Tất cả log từ order-service
{service="order-service"}

# Log từ namespace services-prod
{namespace="services-prod"}

# Kết hợp nhiều labels
{namespace="services-prod", app="order-service", pod=~"order-service-.*"}

4.2 Filter Expressions

# Chứa chuỗi
{service="order-service"} |= "ERROR"

# Không chứa
{service="order-service"} != "health"

# Regex match
{service="order-service"} |~ "order.*created"

# Parse JSON và filter
{service="order-service"}
  | json
  | level = "ERROR"
  | durationMs > 1000

# Pipeline phức tạp
{namespace="services-prod"}
  | json
  | level = "ERROR"
  | line_format "{{.service}}: {{.message}} (trace: {{.traceId}})"

4.3 Metric Queries

# Đếm ERROR logs mỗi 5 phút
sum(rate({service="order-service"} |= "ERROR" [5m])) by (service)

# Log volume theo service
sum(bytes_rate({namespace="services-prod"}[5m])) by (service)

# Top 5 services nhiều lỗi nhất
topk(5,
  sum(count_over_time({namespace="services-prod"} |= "ERROR" [1h]))
  by (service)
)

5. Log Correlation with Traces

Goal: From a log entry, immediately jump to the corresponding trace, and vice versa.

5.1 Inject Trace Context into Log

// Spring Boot với Micrometer Tracing tự động inject
// Log sẽ có traceId và spanId từ MDC

// Kết quả log:
{
  "timestamp": "...",
  "level": "ERROR",
  "service": "order-service",
  "traceId": "abc123def456789",   ← đây
  "spanId": "span789",             ← đây
  "message": "Payment failed"
}

5.2 Grafana Derived Fields

Configure Grafana to create a link from the traceId in the log to Jaeger:

// Trong Loki datasource config (Grafana)
{
  "derivedFields": [
    {
      "matcherRegex": "traceId=(\\w+)",
      "name": "TraceID",
      "url": "http://jaeger:16686/trace/$${__value.raw}",
      "datasourceUid": "jaeger"
    }
  ]
}

Now when viewing logs in Grafana, the traceId will automatically become a clickable link to open the Jaeger trace.


6. Retention & Cost Optimization

6.1 Log Retention Policy

# Loki retention config
limits_config:
  # Giữ log 30 ngày cho prod
  retention_period: 720h

  # Per-stream override
  per_stream_rate_limit: 10MB
  per_stream_rate_limit_burst: 30MB

compactor:
  retention_enabled: true
  retention_delete_delay: 2h

Retention strategy by tier:

Hot  (0-7 ngày)   : SSD storage  — query nhanh
Warm (7-30 ngày)  : HDD storage  — query chậm hơn
Cold (30-365 ngày): S3 Glacier   — archive only

6.2 Reduce Log Volume

# Filtering trước khi ingest (Fluent Bit)
# Loại bỏ health check, static assets, debug logs từ noisy service
# Trong Fluent Bit:
[FILTER]
    Name    grep
    Match   kube.*
    Exclude log (GET /health|GET /metrics|\.css|\.js|\.png)

[FILTER]
    Name    grep
    Match   kube.monitoring.*
    Exclude log .*  # Bỏ hoàn toàn log từ monitoring namespace

Sampling for high-volume services:

// Chỉ log 10% DEBUG requests bình thường, 100% errors
if (random.nextDouble() < 0.1 || isError) {
    log.debug("Request processed", kv("endpoint", endpoint));
}

7. Best Practices

Logging Checklist:

□ Structured JSON logging
□ Consistent field names (snake_case) across services
□ traceId / spanId có mặt trong mọi log line
□ Correlation ID từ user request
□ Không log PII (email, phone, card numbers)
□ Error logs kèm stack trace
□ Business events quan trọng có INFO log
□ Health check endpoints bị filter khỏi logs
□ Log retention policy phù hợp với compliance
□ Alerting trên ERROR log spike

Field naming convention:

{
  "timestamp": "2026-03-31T10:00:00Z",   // ISO 8601
  "level": "INFO",                         // UPPERCASE
  "service": "order-service",              // kebab-case
  "version": "1.2.3",                      // semver
  "traceId": "abc123",                     // camelCase
  "spanId": "def456",
  "userId": "u-001",                       // camelCase
  "message": "...",                        // human-readable
  // Domain fields: snake_case
  "order_id": "O-001",
  "total_amount": 500000,
  "duration_ms": 45                        // unit suffix
}

Summary

ConceptPurpose
Structured LoggingMachine-readable, queryable log format
Log LevelsEnvironmentally appropriate verbosity control
Fluent BitDaemonSet collects logs from all pods
LokiLog aggregation, cost-effective for cloud native
LogQLQuery log using labels and filter expressions
Log CorrelationLink logs with distributed traces
Retention PolicyBalancing cost and compliance requirements

Next article: Distributed Tracing — OpenTelemetry & Jaeger