Chuyển đến nội dung chính

Lesson 10: Observability-driven Testing and AI-assisted Performance

OpenTelemetry, distributed tracing, Grafana dashboards, eBPF profiling, AI-generated test scripts, predictive planning capacity.

🔒 DevSecOps — Lesson 10 Lesson 10: Observability-driven Testing and AI-assisted Performance

Performance Testing & Pentest: Enterprise Standard Process 2026

Part 2: Advanced Performance Testing

xdev.asia

1. OpenTelemetry Integration

Three Pillars of Observability:

  Metrics ──→ Prometheus/Mimir ──→ Grafana Dashboards
  Traces  ──→ Jaeger/Tempo     ──→ Trace Analysis
  Logs    ──→ Loki/ELK         ──→ Log Correlation

OpenTelemetry = Unified SDK cho cả 3 pillars
  ├── Auto-instrumentation (zero-code)
  ├── Manual instrumentation (custom spans)
  ├── Context propagation (trace-id across services)
  └── Exporters: OTLP, Jaeger, Prometheus, Zipkin
# OpenTelemetry Collector config
receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317
      http:
        endpoint: 0.0.0.0:4318

processors:
  batch:
    timeout: 5s
    send_batch_size: 1000

exporters:
  prometheusremotewrite:
    endpoint: "http://mimir:9009/api/v1/push"
  otlp/tempo:
    endpoint: "tempo:4317"
    tls:
      insecure: true
  loki:
    endpoint: "http://loki:3100/loki/api/v1/push"

service:
  pipelines:
    metrics:
      receivers: [otlp]
      processors: [batch]
      exporters: [prometheusremotewrite]
    traces:
      receivers: [otlp]
      processors: [batch]
      exporters: [otlp/tempo]
    logs:
      receivers: [otlp]
      processors: [batch]
      exporters: [loki]

2. Distributed Tracing During Load Tests

// k6 — Inject trace context vào load test requests
import http from 'k6/http';
import { uuidv4 } from 'https://jslib.k6.io/k6-utils/1.4.0/index.js';

export default function () {
  const traceId = uuidv4().replace(/-/g, '');
  const spanId = traceId.substring(0, 16);

  const res = http.get(`${BASE_URL}/api/products`, {
    headers: {
      // W3C Trace Context propagation
      'traceparent': `00-${traceId}-${spanId}-01`,
      // Correlate k6 VU với traces
      'x-k6-vu': `${__VU}`,
      'x-k6-iter': `${__ITER}`,
    },
    tags: { name: 'ListProducts' },
  });

  // Log slow requests for trace lookup
  if (res.timings.duration > 500) {
    console.log(`Slow request trace: ${traceId} - ${res.timings.duration}ms`);
  }
}
Trace Analysis During Load Test:

Request: GET /api/products (p99 = 1200ms)

Trace Timeline:
├── api-gateway     [0-5ms]      HTTP receive + route
├── auth-service    [5-15ms]     JWT validation
├── product-service [15-1100ms]  ← BOTTLENECK
│   ├── cache-check [15-20ms]   Redis GET (cache miss)
│   ├── db-query    [20-950ms]  ← Slow SQL query
│   │   └── pg: SELECT * FROM products JOIN categories...
│   └── serialize   [950-1100ms] JSON serialization
├── response        [1100-1200ms]

Root cause: db-query 930ms → Missing index on products.category_id

3. Grafana Dashboards for Load Testing

k6 → Prometheus → Grafana Pipeline:

Dashboard panels:
┌─────────────────────────────────────────────────────┐
│ Row 1: Overview                                     │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌─────────┐│
│ │ RPS      │ │ p95 ms   │ │ Error %  │ │ VUs     ││
│ │ 2,847    │ │ 245      │ │ 0.03%    │ │ 500     ││
│ └──────────┘ └──────────┘ └──────────┘ └─────────┘│
├─────────────────────────────────────────────────────┤
│ Row 2: Latency Distribution                         │
│ ┌─────────────────────────────────────────────────┐│
│ │ Heatmap: Response time over time                ││
│ │ p50 ████░░░░░░  120ms                           ││
│ │ p95 ████████░░  245ms                           ││
│ │ p99 ██████████  890ms                           ││
│ └─────────────────────────────────────────────────┘│
├─────────────────────────────────────────────────────┤
│ Row 3: Per-endpoint breakdown                       │
│ ┌─────────────────────────────────────────────────┐│
│ │ Table: endpoint | RPS | p95 | p99 | errors      ││
│ │ GET /products   | 1200| 150 | 400 | 0.01%       ││
│ │ GET /products/:id| 800| 100 | 250 | 0.02%       ││
│ │ POST /orders    | 200 | 450 | 1200| 0.1%        ││
│ └─────────────────────────────────────────────────┘│
├─────────────────────────────────────────────────────┤
│ Row 4: Infrastructure                               │
│ ┌───────────┐ ┌───────────┐ ┌───────────┐         │
│ │ CPU %     │ │ Memory %  │ │ DB conns  │         │
│ │ ▄▄▆█ 72% │ │ ▄▄▄▅ 58% │ │ ▄▄██ 45/50│         │
│ └───────────┘ └───────────┘ └───────────┘         │
└─────────────────────────────────────────────────────┘

4. eBPF-based Continuous Profiling

# Grafana Pyroscope — Continuous profiling
docker run -d --name pyroscope grafana/pyroscope

# Instrument application
# Node.js
npm install @pyroscope/nodejs
// Pyroscope integration
import Pyroscope from '@pyroscope/nodejs';

Pyroscope.init({
  serverAddress: 'http://pyroscope:4040',
  appName: 'api-server',
  tags: { env: 'production', region: 'ap-southeast-1' },
});

Pyroscope.start();

// Flame graphs show:
// - CPU time per function
// - Memory allocation hotspots
// - Lock contention
// - Off-CPU time (I/O waits)
Profiling during load test reveals:
  ├── 35% CPU: JSON.parse/stringify → Use streaming parser
  ├── 25% CPU: bcrypt.hash → Move to worker thread
  ├── 20% CPU: RegExp validation → Compile once, reuse
  ├── 10% CPU: Logging → Buffer writes, async flush
  └── 10% CPU: Framework overhead → Normal

5. AI-assisted Performance Testing

AI Applications in Performance Testing (2026):

1. Auto-generate test scripts from OpenAPI spec
   OpenAPI spec → LLM → k6/Gatling scripts
   → Include realistic scenarios, data correlation

2. Intelligent workload modeling
   Production traffic logs → ML → Synthetic workload model
   → Replicate actual user behavior distribution

3. Automated bottleneck root-cause analysis
   Grafana metrics + traces → AI → "DB connection pool
   saturated at 50 connections causing p99 spike"

4. Predictive capacity planning
   Historical metrics → Time-series forecasting →
   "At current growth rate, need 3 more pods by Q3 2026"

5. Anomaly detection
   Baseline metrics → ML model → Auto-alert when
   performance deviates from learned patterns
# Grafana ML — Anomaly detection trên metrics
# Configure trong Grafana UI:
# 1. Select metric: http_request_duration_seconds
# 2. Train model on last 30 days of data
# 3. Set sensitivity: medium
# 4. Alert when prediction band exceeded for > 5 minutes

6. Summary

  • OpenTelemetry: Unified observability for metrics, traces, logs
  • Distributed Tracing: Correlate k6 requests with backend traces
  • Grafana Dashboards: Real-time visualization when running load test
  • eBPF/Pyroscope: Continuous profiling → flame graphs, hotspot analysis
  • AI/ML: Auto-generate scripts, anomaly detection, predictive capacity

The next article will move to Penetration Testing — Process and Legal Framework.