1. OpenTelemetry Integration
Three Pillars of Observability:
Metrics ──→ Prometheus/Mimir ──→ Grafana Dashboards
Traces ──→ Jaeger/Tempo ──→ Trace Analysis
Logs ──→ Loki/ELK ──→ Log Correlation
OpenTelemetry = Unified SDK cho cả 3 pillars
├── Auto-instrumentation (zero-code)
├── Manual instrumentation (custom spans)
├── Context propagation (trace-id across services)
└── Exporters: OTLP, Jaeger, Prometheus, Zipkin
# OpenTelemetry Collector config
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
processors:
batch:
timeout: 5s
send_batch_size: 1000
exporters:
prometheusremotewrite:
endpoint: "http://mimir:9009/api/v1/push"
otlp/tempo:
endpoint: "tempo:4317"
tls:
insecure: true
loki:
endpoint: "http://loki:3100/loki/api/v1/push"
service:
pipelines:
metrics:
receivers: [otlp]
processors: [batch]
exporters: [prometheusremotewrite]
traces:
receivers: [otlp]
processors: [batch]
exporters: [otlp/tempo]
logs:
receivers: [otlp]
processors: [batch]
exporters: [loki]
2. Distributed Tracing During Load Tests
// k6 — Inject trace context vào load test requests
import http from 'k6/http';
import { uuidv4 } from 'https://jslib.k6.io/k6-utils/1.4.0/index.js';
export default function () {
const traceId = uuidv4().replace(/-/g, '');
const spanId = traceId.substring(0, 16);
const res = http.get(`${BASE_URL}/api/products`, {
headers: {
// W3C Trace Context propagation
'traceparent': `00-${traceId}-${spanId}-01`,
// Correlate k6 VU với traces
'x-k6-vu': `${__VU}`,
'x-k6-iter': `${__ITER}`,
},
tags: { name: 'ListProducts' },
});
// Log slow requests for trace lookup
if (res.timings.duration > 500) {
console.log(`Slow request trace: ${traceId} - ${res.timings.duration}ms`);
}
}
Trace Analysis During Load Test:
Request: GET /api/products (p99 = 1200ms)
Trace Timeline:
├── api-gateway [0-5ms] HTTP receive + route
├── auth-service [5-15ms] JWT validation
├── product-service [15-1100ms] ← BOTTLENECK
│ ├── cache-check [15-20ms] Redis GET (cache miss)
│ ├── db-query [20-950ms] ← Slow SQL query
│ │ └── pg: SELECT * FROM products JOIN categories...
│ └── serialize [950-1100ms] JSON serialization
├── response [1100-1200ms]
Root cause: db-query 930ms → Missing index on products.category_id
3. Grafana Dashboards cho Load Testing
k6 → Prometheus → Grafana Pipeline:
Dashboard panels:
┌─────────────────────────────────────────────────────┐
│ Row 1: Overview │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌─────────┐│
│ │ RPS │ │ p95 ms │ │ Error % │ │ VUs ││
│ │ 2,847 │ │ 245 │ │ 0.03% │ │ 500 ││
│ └──────────┘ └──────────┘ └──────────┘ └─────────┘│
├─────────────────────────────────────────────────────┤
│ Row 2: Latency Distribution │
│ ┌─────────────────────────────────────────────────┐│
│ │ Heatmap: Response time over time ││
│ │ p50 ████░░░░░░ 120ms ││
│ │ p95 ████████░░ 245ms ││
│ │ p99 ██████████ 890ms ││
│ └─────────────────────────────────────────────────┘│
├─────────────────────────────────────────────────────┤
│ Row 3: Per-endpoint breakdown │
│ ┌─────────────────────────────────────────────────┐│
│ │ Table: endpoint | RPS | p95 | p99 | errors ││
│ │ GET /products | 1200| 150 | 400 | 0.01% ││
│ │ GET /products/:id| 800| 100 | 250 | 0.02% ││
│ │ POST /orders | 200 | 450 | 1200| 0.1% ││
│ └─────────────────────────────────────────────────┘│
├─────────────────────────────────────────────────────┤
│ Row 4: Infrastructure │
│ ┌───────────┐ ┌───────────┐ ┌───────────┐ │
│ │ CPU % │ │ Memory % │ │ DB conns │ │
│ │ ▄▄▆█ 72% │ │ ▄▄▄▅ 58% │ │ ▄▄██ 45/50│ │
│ └───────────┘ └───────────┘ └───────────┘ │
└─────────────────────────────────────────────────────┘
4. eBPF-based Continuous Profiling
# Grafana Pyroscope — Continuous profiling
docker run -d --name pyroscope grafana/pyroscope
# Instrument application
# Node.js
npm install @pyroscope/nodejs
// Pyroscope integration
import Pyroscope from '@pyroscope/nodejs';
Pyroscope.init({
serverAddress: 'http://pyroscope:4040',
appName: 'api-server',
tags: { env: 'production', region: 'ap-southeast-1' },
});
Pyroscope.start();
// Flame graphs show:
// - CPU time per function
// - Memory allocation hotspots
// - Lock contention
// - Off-CPU time (I/O waits)
Profiling during load test reveals:
├── 35% CPU: JSON.parse/stringify → Use streaming parser
├── 25% CPU: bcrypt.hash → Move to worker thread
├── 20% CPU: RegExp validation → Compile once, reuse
├── 10% CPU: Logging → Buffer writes, async flush
└── 10% CPU: Framework overhead → Normal
5. AI-assisted Performance Testing
AI Applications in Performance Testing (2026):
1. Auto-generate test scripts from OpenAPI spec
OpenAPI spec → LLM → k6/Gatling scripts
→ Include realistic scenarios, data correlation
2. Intelligent workload modeling
Production traffic logs → ML → Synthetic workload model
→ Replicate actual user behavior distribution
3. Automated bottleneck root-cause analysis
Grafana metrics + traces → AI → "DB connection pool
saturated at 50 connections causing p99 spike"
4. Predictive capacity planning
Historical metrics → Time-series forecasting →
"At current growth rate, need 3 more pods by Q3 2026"
5. Anomaly detection
Baseline metrics → ML model → Auto-alert when
performance deviates from learned patterns
# Grafana ML — Anomaly detection trên metrics
# Configure trong Grafana UI:
# 1. Select metric: http_request_duration_seconds
# 2. Train model on last 30 days of data
# 3. Set sensitivity: medium
# 4. Alert when prediction band exceeded for > 5 minutes
6. Tổng kết
- OpenTelemetry: Unified observability cho metrics, traces, logs
- Distributed Tracing: Correlate k6 requests với backend traces
- Grafana Dashboards: Real-time visualization khi chạy load test
- eBPF/Pyroscope: Continuous profiling → flame graphs, hotspot analysis
- AI/ML: Auto-generate scripts, anomaly detection, predictive capacity
Bài tiếp theo sẽ chuyển sang Penetration Testing — Quy trình và Khung pháp lý.