Chuyển đến nội dung chính

Bài 2: SRE Practices — SLO, SLI, SLA và Error Budgets

Định nghĩa SLIs, thiết lập SLOs với error budgets, SLA obligations, burn rate alerts và Google SRE principles.

🔒 DevSecOps — Bài 2 Bài 2: SRE Practices — SLO, SLI, SLA và Error Budgets

Performance Testing & Pentest: Quy trình Chuẩn Doanh nghiệp 2026

Phần 1: Nền tảng Performance Testing

xdev.asia

1. SLI, SLO, SLA — Phân biệt

┌─────────────────────────────────────────────────────────┐
│                SLI → SLO → SLA                         │
├──────────┬──────────────────────────────────────────────┤
│ SLI      │ Service Level INDICATOR                     │
│          │ = Metric đo lường (số liệu cụ thể)         │
│          │ Ví dụ: "Tỷ lệ requests có latency < 200ms" │
├──────────┼──────────────────────────────────────────────┤
│ SLO      │ Service Level OBJECTIVE                     │
│          │ = Mục tiêu nội bộ cho SLI                   │
│          │ Ví dụ: "99.9% requests < 200ms"             │
├──────────┼──────────────────────────────────────────────┤
│ SLA      │ Service Level AGREEMENT                     │
│          │ = Hợp đồng với khách hàng, có penalty       │
│          │ Ví dụ: "99.5% uptime, vi phạm = credit"     │
└──────────┴──────────────────────────────────────────────┘

Nguyên tắc: SLO phải strict hơn SLA
  SLA: 99.5% → SLO: 99.9% → "Buffer" 0.4%

2. Định nghĩa SLI

SLI tốt phải đáp ứng:
  ✅ Đo lường được (quantifiable)
  ✅ Phản ánh trải nghiệm người dùng
  ✅ Có threshold rõ ràng (good/bad event)

Các SLIs phổ biến:
┌────────────────┬─────────────────────────────────────────┐
│ Category       │ SLI Formula                             │
├────────────────┼─────────────────────────────────────────┤
│ Availability   │ Good requests / Total requests          │
│ Latency        │ Requests < threshold / Total requests   │
│ Throughput     │ Processed events / Expected events      │
│ Error Rate     │ Successful requests / Total requests    │
│ Freshness      │ Data updates < threshold / Total        │
│ Correctness    │ Correct responses / Total responses     │
└────────────────┴─────────────────────────────────────────┘
# Ví dụ SLI definitions cho API service
slis:
  availability:
    description: "Tỷ lệ requests trả về non-5xx status"
    formula: "count(http_status < 500) / count(total_requests)"
    measurement: prometheus
    query: |
      sum(rate(http_requests_total{status!~"5.."}[5m]))
      / sum(rate(http_requests_total[5m]))

  latency:
    description: "Tỷ lệ requests có latency < 300ms"
    formula: "count(latency < 300ms) / count(total_requests)"
    query: |
      sum(rate(http_request_duration_seconds_bucket{le="0.3"}[5m]))
      / sum(rate(http_request_duration_seconds_count[5m]))

3. Thiết lập SLO

# slo-definitions.yml
services:
  payment-api:
    slos:
      - name: "Availability"
        sli: availability
        target: 99.99%        # "Four nines"
        window: 30d           # rolling 30-day
        # Cho phép: 4.32 phút downtime / 30 ngày

      - name: "Latency (p95)"
        sli: latency_p95
        target: 99.5%
        threshold: 300ms
        window: 30d
        # Cho phép: 0.5% requests > 300ms

      - name: "Latency (p99)"
        sli: latency_p99
        target: 99.0%
        threshold: 1000ms
        window: 30d

  product-catalog:
    slos:
      - name: "Availability"
        sli: availability
        target: 99.9%         # "Three nines"
        window: 30d
        # Cho phép: 43.2 phút downtime / 30 ngày

Chọn SLO target phù hợp

Availability targets và downtime cho phép:
┌──────────┬───────────────┬──────────────┬───────────────┐
│ Target   │ / ngày        │ / tháng      │ / năm         │
├──────────┼───────────────┼──────────────┼───────────────┤
│ 99%      │ 14.4 phút     │ 7.3 giờ      │ 3.65 ngày     │
│ 99.9%    │ 1.44 phút     │ 43.2 phút    │ 8.76 giờ      │
│ 99.95%   │ 43.2 giây     │ 21.6 phút    │ 4.38 giờ      │
│ 99.99%   │ 8.64 giây     │ 4.32 phút    │ 52.6 phút     │
│ 99.999%  │ 0.86 giây     │ 25.9 giây    │ 5.26 phút     │
└──────────┴───────────────┴──────────────┴───────────────┘

4. Error Budgets

Error Budget = 1 - SLO target

Ví dụ: SLO = 99.9%  →  Error Budget = 0.1%

Trong 30 ngày (= 43,200 phút):
  Error Budget = 43,200 × 0.001 = 43.2 phút

Cách dùng Error Budget:
  ├── Còn budget → Deploy features mới, thử nghiệm
  ├── Budget cạn → Freeze deployments, focus stability
  └── Budget hết → Chỉ deploy hotfixes + reliability work
# Error budget policy
error_budget_policy:
  thresholds:
    - remaining: "> 50%"
      actions:
        - "Normal development velocity"
        - "Feature deployments allowed"
        - "Experiments allowed"

    - remaining: "25% - 50%"
      actions:
        - "Reduce deployment frequency"
        - "Mandatory canary deployments"
        - "Review recent incidents"

    - remaining: "< 25%"
      actions:
        - "Feature freeze"
        - "Focus on reliability improvements"
        - "Postmortem all incidents"

    - remaining: "0% (exhausted)"
      actions:
        - "Emergency mode"
        - "Only hotfixes allowed"
        - "Mandatory incident review"
        - "Escalate to VP Engineering"

5. Burn Rate Alerts

Burn Rate = Tốc độ tiêu thụ error budget

Burn Rate = 1  → Hết budget đúng cuối window
Burn Rate = 2  → Hết budget trong 15 ngày (thay vì 30)
Burn Rate = 10 → Hết budget trong 3 ngày
Burn Rate = 36 → Hết budget trong 20 giờ
# Multi-window burn rate alerts (Google SRE recommendation)
alerts:
  - name: "SLO_HighBurnRate_Page"
    severity: page          # PagerDuty, wake people up
    long_window: 1h
    short_window: 5m
    burn_rate: 14.4         # Hết budget trong ~2 ngày
    description: "Error budget burning 14.4x faster than normal"

  - name: "SLO_HighBurnRate_Ticket"
    severity: ticket        # Jira ticket, fix trong ngày
    long_window: 6h
    short_window: 30m
    burn_rate: 6            # Hết budget trong ~5 ngày

  - name: "SLO_MediumBurnRate"
    severity: warning
    long_window: 3d
    short_window: 6h
    burn_rate: 1            # Đang theo đúng tốc độ hết budget
# Prometheus alerting rule
groups:
  - name: slo-burn-rate
    rules:
      - alert: HighErrorBudgetBurn
        expr: |
          (
            1 - (sum(rate(http_requests_total{status!~"5.."}[1h]))
            / sum(rate(http_requests_total[1h])))
          ) > (14.4 * 0.001)  # 14.4 × error budget rate
        for: 2m
        labels:
          severity: page
        annotations:
          summary: "High error budget burn rate detected"
          description: "Service burning error budget 14.4x faster"

6. SLO Documents

SLO Document Template:
┌─────────────────────────────────────────────────────┐
│ Service: Payment API                                │
│ Owner: Platform Team                                │
│ Last Review: 2026-03-15                             │
├─────────────────────────────────────────────────────┤
│ SLI: Availability                                   │
│ Description: Non-5xx responses / total responses    │
│ SLO Target: 99.99% (rolling 30 days)               │
│ Error Budget: 4.32 minutes/month                   │
│ Data Source: Prometheus (http_requests_total)       │
│ Dashboard: https://grafana.../slo-payment           │
├─────────────────────────────────────────────────────┤
│ Escalation:                                         │
│   50% budget → Slow down deploys                   │
│   25% budget → Feature freeze                      │
│   0% budget → Emergency response                   │
├─────────────────────────────────────────────────────┤
│ Review Cadence: Monthly                             │
│ Next Review: 2026-04-15                             │
└─────────────────────────────────────────────────────┘

7. Tổng kết

  • SLI: Metric cụ thể đo lường quality (availability, latency, throughput)
  • SLO: Mục tiêu nội bộ cho SLI (99.9% availability)
  • SLA: Hợp đồng với khách hàng, có penalty
  • Error Budget: "Quota" cho phép fail — cân bằng velocity vs reliability
  • Burn Rate Alerts: Multi-window alerts phát hiện sớm khi budget cạn

Bài tiếp theo sẽ tìm hiểu Shift-left Performance Testing trong CI/CD.