Chuyển đến nội dung chính

第 23 課:可觀察性 - 監控、記錄和追蹤

可觀察性的三大支柱:指標、日誌、追蹤。 Prometheus + Grafana 監控堆疊。集中式日誌記錄 (ELK/EFK)。分散式追蹤(Jaeger、OpenTelemetry)。警報策略。重新檢視 SLI/SLO/SLA。微服務的可觀察性。

🏗️ 建築 — 第 23 課 第 23 課:可觀察性 - 監控、 記錄和追蹤

系統架構:從零到英雄

第 6 部分:可靠性、安全性和可觀察性

亞洲開發網

簡介

監控顯示出了什麼問題。可觀察性告訴為什麼它是錯的。在分散式系統中,可觀察性至關重要——你無法修復看不到的東西。


1. 可觀察性的三大支柱

┌────────────────────────────────────────────────────┐
│                 OBSERVABILITY                       │
│                                                     │
│  ┌──────────┐    ┌──────────┐    ┌──────────┐      │
│  │ METRICS  │    │  LOGS    │    │ TRACES   │      │
│  │          │    │          │    │          │      │
│  │ What is  │    │ What     │    │ Where is │      │
│  │ happening│    │ happened │    │ the time │      │
│  │ now?     │    │ exactly? │    │ spent?   │      │
│  │          │    │          │    │          │      │
│  │Prometheus│    │ELK/Loki │    │Jaeger/   │      │
│  │Grafana   │    │Fluentd  │    │Zipkin    │      │
│  └──────────┘    └──────────┘    └──────────┘      │
│                                                     │
│  Numbers         Text Events     Request Flow       │
│  Time-series     Structured      Cross-service      │
│  Aggregatable    Searchable      Latency breakdown  │
└────────────────────────────────────────────────────┘

2. 指標(Prometheus + Grafana)

2.1 指標類型

Counter: Giá trị chỉ tăng
  http_requests_total{method="GET", status="200"} = 15234
  Use: Request count, error count, bytes transferred

Gauge: Giá trị lên xuống
  memory_usage_bytes = 1073741824
  Use: Temperature, queue size, active connections

Histogram: Phân phối giá trị
  http_request_duration_seconds_bucket{le="0.1"} = 5000
  http_request_duration_seconds_bucket{le="0.5"} = 8000
  http_request_duration_seconds_bucket{le="1.0"} = 9500
  Use: Latency percentiles, request sizes

Summary: Tương tự histogram, pre-calculated percentiles
  http_request_duration_seconds{quantile="0.99"} = 0.45

2.2 RED 方法(針對服務)

Rate:    Requests per second
Errors:  Errors per second
Duration: Latency distribution

Dashboard:
  ┌─────────────────────────────────────┐
  │ Service: Order API                   │
  │                                      │
  │ Rate:     523 req/s  [▓▓▓▓▓░░░░░]   │
  │ Errors:   0.3%       [▓░░░░░░░░░]   │
  │ Duration: p50=12ms   p99=145ms       │
  │           [graph ~~~~~~~~~~~~~~~~~~~~│
  └─────────────────────────────────────┘

2.3 USE 方法(針對資源)

Utilization: % resource busy
Saturation:  Work queued/waiting
Errors:      Error count

CPU:    Utilization 75%, Saturation (load avg) 2.3, Errors 0
Memory: Utilization 82%, Saturation (swap) 100MB, Errors 0
Disk:   Utilization 60%, Saturation (I/O wait) 5%, Errors 2
Network: Utilization 30%, Saturation (TCP retransmit) 0.1%

2.4 普羅米修斯架構

┌─────────────┐         ┌──────────────┐
│ App Server  │◄─scrape─│  Prometheus  │
│ /metrics    │         │  Server      │
└─────────────┘         │              │
                        │ TSDB storage │
┌─────────────┐         │ PromQL query │
│ Database    │◄─scrape─│ Alert rules  │
│ Exporter    │         └──────┬───────┘
└─────────────┘                │
                        ┌──────▼───────┐
┌─────────────┐         │ Alertmanager │──► PagerDuty
│ Node        │◄─scrape─│              │──► Slack
│ Exporter    │         └──────────────┘
└─────────────┘                │
                        ┌──────▼───────┐
                        │   Grafana    │
                        │ Dashboards   │
                        └──────────────┘

3. 日誌記錄

3.1 結構化日誌記錄

❌ Unstructured:
  "User 123 placed order 456 for $100.00"
  → Khó parse, search, aggregate

✅ Structured (JSON):
  {
    "timestamp": "2024-01-15T10:30:00Z",
    "level": "info",
    "service": "order-service",
    "trace_id": "abc123",
    "user_id": "123",
    "order_id": "456",
    "amount": 100.00,
    "message": "Order placed successfully"
  }
  → Dễ search, filter, aggregate
  → Correlate với traces (trace_id)

3.2 日誌等級

FATAL:   App sắp crash, cần intervention ngay
ERROR:   Operation failed, nhưng app vẫn chạy
WARN:    Sắp có vấn đề (disk 90%, high latency)
INFO:    Business events quan trọng
DEBUG:   Chi tiết cho troubleshooting
TRACE:   Rất chi tiết (function entry/exit)

Production: INFO + WARN + ERROR + FATAL
Debug mode: + DEBUG
Never in production: TRACE (quá nhiều data)

3.3 集中式日誌記錄(ELK)

┌─────────┐  ┌─────────┐  ┌─────────┐
│ App 1   │  │ App 2   │  │ App 3   │
│ stdout  │  │ stdout  │  │ stdout  │
└────┬────┘  └────┬────┘  └────┬────┘
     │            │            │
     ▼            ▼            ▼
┌─────────────────────────────────────┐
│ Fluentd / Filebeat / Vector        │ ← Collect & ship
└────────────────┬────────────────────┘
                 ▼
┌────────────────────────────────────┐
│ Elasticsearch / Loki              │ ← Store & index
└────────────────┬───────────────────┘
                 ▼
┌────────────────────────────────────┐
│ Kibana / Grafana                  │ ← Search & visualize
└────────────────────────────────────┘

Query example (Kibana):
  service:"order-service" AND level:"error" AND user_id:"123"
  → Tìm tất cả errors của user 123 trong order service

4. 分散式追蹤

4.1 為什麼我們需要追蹤?

Request: GET /api/orders/123

Monolith: 1 log file, dễ theo dõi

Microservices:
  API Gateway → Order Service → User Service
                              → Inventory Service
                              → Payment Service

  "Request chậm 5 giây. Service nào gây ra?"
  
  Không có tracing → Phải check logs từng service
  Có tracing → Thấy ngay bottleneck

4.2 追蹤結構

Trace ID: abc-123 (toàn bộ request flow)

  ┌────────────────────────────────────────────────┐
  │ Span: API Gateway (500ms total)                │
  │ ├── Span: Order Service (450ms)                │
  │ │   ├── Span: DB Query (50ms)                  │
  │ │   ├── Span: User Service (30ms)              │
  │ │   ├── Span: Inventory Service (350ms) ← SLOW!│
  │ │   │   └── Span: DB Query (340ms) ← ROOT CAUSE│
  │ │   └── Span: Payment Service (15ms)           │
  │ └── Span: Response Serialization (5ms)         │
  └────────────────────────────────────────────────┘

Context Propagation:
  Trace ID + Span ID passed via HTTP headers
  traceparent: 00-abc123-span456-01

4.3 開放遙測

OpenTelemetry (OTel): Vendor-neutral observability framework

┌─────────────────────────────────────┐
│ Application                         │
│ ┌─────────────────────────────────┐ │
│ │ OTel SDK                        │ │
│ │ Auto-instrumentation            │ │
│ │ (HTTP, DB, gRPC, messaging)     │ │
│ └──────────────┬──────────────────┘ │
└────────────────┼────────────────────┘
                 │ OTLP (protocol)
                 ▼
┌─────────────────────────────────────┐
│ OTel Collector                      │
│ Receive → Process → Export          │
└──────────┬──────────┬───────────────┘
           │          │
     ┌─────▼───┐  ┌───▼──────┐
     │ Jaeger  │  │Prometheus│
     │(traces) │  │(metrics) │
     └─────────┘  └──────────┘

Ưu điểm: Instrument once, export anywhere

5. 警報

5.1 警報設計

✅ Tốt:
  - Alert trên SYMPTOMS (user-facing impact)
  - "Error rate > 1% trong 5 phút"
  - "P99 latency > 2 giây"
  
❌ Xấu:
  - Alert trên CAUSES (noisy)
  - "CPU > 80%" (có thể bình thường)
  - "1 instance down" (auto-scaling xử lý)

Alert Severity:
  P1 (Critical): Revenue impact, data loss
    → Page on-call IMMEDIATELY
  P2 (High): Degraded performance, partial outage
    → Page during business hours
  P3 (Medium): Non-critical issue
    → Slack notification, fix next business day
  P4 (Low): Informational
    → Ticket, fix when convenient

5.2 待命最佳實踐

1. Runbooks cho mỗi alert
2. Escalation path rõ ràng
3. Post-incident review (blameless)
4. Alert fatigue prevention:
   - Mỗi alert phải actionable
   - Review alert hàng tháng (remove noisy ones)
   - Max 5-10 pages/week
5. Rotation: 1 tuần on-call, ít nhất 2 người trong pool

6. 可觀察性儀表板

┌─────────────────────────────────────────────────────┐
│ System Overview                                      │
│                                                       │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐  │
│ │ Requests/s   │ │ Error Rate   │ │ P99 Latency  │  │
│ │    1,234     │ │   0.12%      │ │   145ms      │  │
│ │ ▓▓▓▓▓▓▓░░░  │ │ ▓░░░░░░░░░  │ │ ▓▓▓░░░░░░░  │  │
│ └──────────────┘ └──────────────┘ └──────────────┘  │
│                                                       │
│ ┌──────────────────────────────────────────────────┐ │
│ │ Service Health Map                                │ │
│ │ API Gateway [✅] → Order [✅] → Payment [⚠️]     │ │
│ │                  → User [✅]  → Email [❌]        │ │
│ └──────────────────────────────────────────────────┘ │
│                                                       │
│ ┌──────────────────────────────────────────────────┐ │
│ │ Recent Alerts                                     │ │
│ │ 🔴 P1: Payment latency > 2s (10 min ago)         │ │
│ │ 🟡 P3: Email service connection timeout           │ │
│ └──────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────┘

總結

支柱工具目的
指標普羅米修斯 + Grafana現在發生了什麼
日誌ELK / 洛基發生了什麼(詳情)
痕跡耶格/節奏時間都花在哪裡了
警報警報管理器/PagerDuty何時採取行動

練習

  1. 監控設定: 微服務:API網關、使用者、訂單、付款、通知。對於每項服務,列出 5 個最重要的要監控的指標。使用紅色方法。

  2. 追蹤分析: 追蹤顯示:API網關(2s)→訂單(1.8s)→資料庫(50ms)→支付(1.7s)→外部API(1.5s)。識別瓶頸。建議3種優化方法。

  3. 警報設計: 設計電子商務警報系統。定義 3 個 P1、3 個 P2、3 個 P3 警報。對於每個警報:條件、閾值、升級、操作手冊摘要。