1. SLI, SLO, SLA — Distinction
┌─────────────────────────────────────────────────────────┐
│ SLI → SLO → SLA │
├──────────┬──────────────────────────────────────────────┤
│ SLI │ Service Level INDICATOR │
│ │ = Metric đo lường (số liệu cụ thể) │
│ │ Ví dụ: "Tỷ lệ requests có latency < 200ms" │
├──────────┼──────────────────────────────────────────────┤
│ SLO │ Service Level OBJECTIVE │
│ │ = Mục tiêu nội bộ cho SLI │
│ │ Ví dụ: "99.9% requests < 200ms" │
├──────────┼──────────────────────────────────────────────┤
│ SLA │ Service Level AGREEMENT │
│ │ = Hợp đồng với khách hàng, có penalty │
│ │ Ví dụ: "99.5% uptime, vi phạm = credit" │
└──────────┴──────────────────────────────────────────────┘
Nguyên tắc: SLO phải strict hơn SLA
SLA: 99.5% → SLO: 99.9% → "Buffer" 0.4%
2. Definition of SLI
SLI tốt phải đáp ứng:
✅ Đo lường được (quantifiable)
✅ Phản ánh trải nghiệm người dùng
✅ Có threshold rõ ràng (good/bad event)
Các SLIs phổ biến:
┌────────────────┬─────────────────────────────────────────┐
│ Category │ SLI Formula │
├────────────────┼─────────────────────────────────────────┤
│ Availability │ Good requests / Total requests │
│ Latency │ Requests < threshold / Total requests │
│ Throughput │ Processed events / Expected events │
│ Error Rate │ Successful requests / Total requests │
│ Freshness │ Data updates < threshold / Total │
│ Correctness │ Correct responses / Total responses │
└────────────────┴─────────────────────────────────────────┘
# Ví dụ SLI definitions cho API service
slis:
availability:
description: "Tỷ lệ requests trả về non-5xx status"
formula: "count(http_status < 500) / count(total_requests)"
measurement: prometheus
query: |
sum(rate(http_requests_total{status!~"5.."}[5m]))
/ sum(rate(http_requests_total[5m]))
latency:
description: "Tỷ lệ requests có latency < 300ms"
formula: "count(latency < 300ms) / count(total_requests)"
query: |
sum(rate(http_request_duration_seconds_bucket{le="0.3"}[5m]))
/ sum(rate(http_request_duration_seconds_count[5m]))
3. Set SLO
# slo-definitions.yml
services:
payment-api:
slos:
- name: "Availability"
sli: availability
target: 99.99% # "Four nines"
window: 30d # rolling 30-day
# Cho phép: 4.32 phút downtime / 30 ngày
- name: "Latency (p95)"
sli: latency_p95
target: 99.5%
threshold: 300ms
window: 30d
# Cho phép: 0.5% requests > 300ms
- name: "Latency (p99)"
sli: latency_p99
target: 99.0%
threshold: 1000ms
window: 30d
product-catalog:
slos:
- name: "Availability"
sli: availability
target: 99.9% # "Three nines"
window: 30d
# Cho phép: 43.2 phút downtime / 30 ngày
Choose the appropriate SLO target__HTMLTAG_80___
Availability targets và downtime cho phép:
┌──────────┬───────────────┬──────────────┬───────────────┐
│ Target │ / ngày │ / tháng │ / năm │
├──────────┼───────────────┼──────────────┼───────────────┤
│ 99% │ 14.4 phút │ 7.3 giờ │ 3.65 ngày │
│ 99.9% │ 1.44 phút │ 43.2 phút │ 8.76 giờ │
│ 99.95% │ 43.2 giây │ 21.6 phút │ 4.38 giờ │
│ 99.99% │ 8.64 giây │ 4.32 phút │ 52.6 phút │
│ 99.999% │ 0.86 giây │ 25.9 giây │ 5.26 phút │
└──────────┴───────────────┴──────────────┴───────────────┘
Availability targets và downtime cho phép:
┌──────────┬───────────────┬──────────────┬───────────────┐
│ Target │ / ngày │ / tháng │ / năm │
├──────────┼───────────────┼──────────────┼───────────────┤
│ 99% │ 14.4 phút │ 7.3 giờ │ 3.65 ngày │
│ 99.9% │ 1.44 phút │ 43.2 phút │ 8.76 giờ │
│ 99.95% │ 43.2 giây │ 21.6 phút │ 4.38 giờ │
│ 99.99% │ 8.64 giây │ 4.32 phút │ 52.6 phút │
│ 99.999% │ 0.86 giây │ 25.9 giây │ 5.26 phút │
└──────────┴───────────────┴──────────────┴───────────────┘
4. Error Budgets
Error Budget = 1 - SLO target
Ví dụ: SLO = 99.9% → Error Budget = 0.1%
Trong 30 ngày (= 43,200 phút):
Error Budget = 43,200 × 0.001 = 43.2 phút
Cách dùng Error Budget:
├── Còn budget → Deploy features mới, thử nghiệm
├── Budget cạn → Freeze deployments, focus stability
└── Budget hết → Chỉ deploy hotfixes + reliability work
# Error budget policy
error_budget_policy:
thresholds:
- remaining: "> 50%"
actions:
- "Normal development velocity"
- "Feature deployments allowed"
- "Experiments allowed"
- remaining: "25% - 50%"
actions:
- "Reduce deployment frequency"
- "Mandatory canary deployments"
- "Review recent incidents"
- remaining: "< 25%"
actions:
- "Feature freeze"
- "Focus on reliability improvements"
- "Postmortem all incidents"
- remaining: "0% (exhausted)"
actions:
- "Emergency mode"
- "Only hotfixes allowed"
- "Mandatory incident review"
- "Escalate to VP Engineering"
5. Burn Rate Alerts
Burn Rate = Tốc độ tiêu thụ error budget
Burn Rate = 1 → Hết budget đúng cuối window
Burn Rate = 2 → Hết budget trong 15 ngày (thay vì 30)
Burn Rate = 10 → Hết budget trong 3 ngày
Burn Rate = 36 → Hết budget trong 20 giờ
# Multi-window burn rate alerts (Google SRE recommendation)
alerts:
- name: "SLO_HighBurnRate_Page"
severity: page # PagerDuty, wake people up
long_window: 1h
short_window: 5m
burn_rate: 14.4 # Hết budget trong ~2 ngày
description: "Error budget burning 14.4x faster than normal"
- name: "SLO_HighBurnRate_Ticket"
severity: ticket # Jira ticket, fix trong ngày
long_window: 6h
short_window: 30m
burn_rate: 6 # Hết budget trong ~5 ngày
- name: "SLO_MediumBurnRate"
severity: warning
long_window: 3d
short_window: 6h
burn_rate: 1 # Đang theo đúng tốc độ hết budget
# Prometheus alerting rule
groups:
- name: slo-burn-rate
rules:
- alert: HighErrorBudgetBurn
expr: |
(
1 - (sum(rate(http_requests_total{status!~"5.."}[1h]))
/ sum(rate(http_requests_total[1h])))
) > (14.4 * 0.001) # 14.4 × error budget rate
for: 2m
labels:
severity: page
annotations:
summary: "High error budget burn rate detected"
description: "Service burning error budget 14.4x faster"
6. SLO Documents
SLO Document Template:
┌─────────────────────────────────────────────────────┐
│ Service: Payment API │
│ Owner: Platform Team │
│ Last Review: 2026-03-15 │
├─────────────────────────────────────────────────────┤
│ SLI: Availability │
│ Description: Non-5xx responses / total responses │
│ SLO Target: 99.99% (rolling 30 days) │
│ Error Budget: 4.32 minutes/month │
│ Data Source: Prometheus (http_requests_total) │
│ Dashboard: https://grafana.../slo-payment │
├─────────────────────────────────────────────────────┤
│ Escalation: │
│ 50% budget → Slow down deploys │
│ 25% budget → Feature freeze │
│ 0% budget → Emergency response │
├─────────────────────────────────────────────────────┤
│ Review Cadence: Monthly │
│ Next Review: 2026-04-15 │
└─────────────────────────────────────────────────────┘
7. Summary
- SLI: Specific metrics that measure quality (availability, latency, throughput)
- SLO: Internal target for SLI (99.9% availability)
- SLA: Contract with customer, with penalty
- Error Budget: "Quota" allows failure — balancing velocity vs reliability
- Burn Rate Alerts: Multi-window alerts to detect early when budget is depleted
The next article will learn Shift-left Performance Testing in CI/CD.