Chuyển đến nội dung chính

Bài 8: Cloud-native Performance Testing

Testing auto-scaling, serverless cold start, service mesh overhead, CDN performance, multi-region latency testing.

🔒 DevSecOps — Bài 8 Bài 8: Cloud-native Performance Testing

Performance Testing & Pentest: Quy trình Chuẩn Doanh nghiệp 2026

Phần 2: Performance Testing Nâng cao

xdev.asia

1. Kubernetes Auto-scaling Testing

# HPA configuration
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: api-hpa
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: api-server
  minReplicas: 3
  maxReplicas: 50
  metrics:
    - type: Resource
      resource:
        name: cpu
        target:
          type: Utilization
          averageUtilization: 70
    - type: Pods
      pods:
        metric:
          name: http_requests_per_second
        target:
          type: AverageValue
          averageValue: 500      # Scale khi > 500 RPS / pod
  behavior:
    scaleUp:
      stabilizationWindowSeconds: 30
      policies:
        - type: Percent
          value: 100             # Double pods mỗi lần scale
          periodSeconds: 60
    scaleDown:
      stabilizationWindowSeconds: 300
      policies:
        - type: Percent
          value: 10
          periodSeconds: 60
// k6 — Test HPA scale-out behavior
export const options = {
  scenarios: {
    hpa_test: {
      executor: 'ramping-arrival-rate',
      startRate: 50,
      timeUnit: '1s',
      stages: [
        { duration: '2m', target: 50 },    // Baseline (3 pods)
        { duration: '3m', target: 500 },   // Trigger scale-out
        { duration: '5m', target: 1000 },  // Max scale
        { duration: '5m', target: 1000 },  // Sustain peak
        { duration: '5m', target: 50 },    // Scale-down
        { duration: '5m', target: 50 },    // Verify scale-down
      ],
      preAllocatedVUs: 200,
      maxVUs: 1000,
    },
  },
};

// Monitor cùng lúc:
// kubectl get hpa -w
// kubectl get pods -w
// Grafana dashboard: pods count + latency + RPS

2. Serverless Cold Start Profiling

// Test Lambda cold start vs warm invocation
import http from 'k6/http';
import { Trend } from 'k6/metrics';

const coldStart = new Trend('cold_start_duration');
const warmStart = new Trend('warm_start_duration');

export const options = {
  scenarios: {
    // Scenario 1: Sequential calls (force cold starts)
    cold_start_test: {
      executor: 'per-vu-iterations',
      vus: 50,              // 50 concurrent → 50 new containers
      iterations: 1,
      maxDuration: '2m',
    },
    // Scenario 2: Rapid calls (keep warm)
    warm_test: {
      executor: 'constant-arrival-rate',
      rate: 100,
      timeUnit: '1s',
      duration: '5m',
      preAllocatedVUs: 20,
      startTime: '3m',
    },
  },
};

export default function () {
  const res = http.get(`${LAMBDA_URL}/api/process`);
  
  // Check X-Cold-Start header hoặc init duration
  const isCold = res.headers['X-Cold-Start'] === 'true';
  
  if (isCold) {
    coldStart.add(res.timings.duration);
  } else {
    warmStart.add(res.timings.duration);
  }
}
Cold Start Optimization Strategies:
  ├── Provisioned Concurrency (AWS Lambda)
  ├── Min instances (Cloud Functions)
  ├── Smaller deployment package
  ├── Lazy initialization
  ├── SnapStart (Java Lambda)
  └── Warm-up cron (scheduled ping)

3. Service Mesh Overhead Measurement

Test: So sánh latency có/không có sidecar proxy

Scenario A: Direct pod-to-pod (no mesh)
  Service A → Service B
  p95: 5ms

Scenario B: With Istio sidecar
  Service A → Envoy proxy → Envoy proxy → Service B
  p95: 12ms
  
  → Overhead: ~7ms (140% increase)
  → Accept nếu thêm mTLS, observability, traffic management

Metrics cần đo:
  ├── Sidecar CPU/memory overhead per pod
  ├── Added latency per hop
  ├── Connection pool exhaustion
  ├── mTLS handshake overhead
  └── Control plane (istiod) resource usage

4. CDN Performance Validation

// Test CDN cache hit ratio và origin offload
import http from 'k6/http';
import { Counter } from 'k6/metrics';

const cacheHit = new Counter('cdn_cache_hit');
const cacheMiss = new Counter('cdn_cache_miss');

export default function () {
  // Test từ nhiều regions (k6 Cloud)
  const res = http.get('https://cdn.example.com/static/bundle.js');
  
  const cacheStatus = res.headers['X-Cache'] || res.headers['Cf-Cache-Status'];
  
  if (cacheStatus === 'HIT') {
    cacheHit.add(1);
  } else {
    cacheMiss.add(1);
  }

  // Measure TTFB (Time to First Byte)
  // res.timings.waiting = TTFB
  check(res, {
    'TTFB < 50ms': (r) => r.timings.waiting < 50,
    'Total < 200ms': (r) => r.timings.duration < 200,
  });
}

5. Multi-region Latency Testing

# k6 Cloud — Distributed load zones
export const options = {
  ext: {
    loadimpact: {
      distribution: {
        'amazon:us:ashburn':     { loadZone: 'amazon:us:ashburn', percent: 30 },
        'amazon:eu:frankfurt':   { loadZone: 'amazon:eu:frankfurt', percent: 25 },
        'amazon:ap:singapore':   { loadZone: 'amazon:ap:singapore', percent: 25 },
        'amazon:ap:sydney':      { loadZone: 'amazon:ap:sydney', percent: 20 },
      },
    },
  },
};
Multi-region Testing Checklist:
  ☐ DNS resolution time per region
  ☐ TCP connect time (geographic distance)
  ☐ TLS handshake overhead
  ☐ TTFB per region
  ☐ Data transfer time (payload size × bandwidth)
  ☐ CDN cache hit ratio per POP
  ☐ Database read latency (cross-region reads)
  ☐ Write replication lag (multi-master)

6. Tổng kết

  • HPA Testing: Validate scale-out speed, warm-up time, scale-down behavior
  • Serverless: Cold start profiling, provisioned concurrency testing
  • Service Mesh: Measure sidecar overhead per hop
  • CDN: Cache hit ratio, TTFB, origin offload
  • Multi-region: Distributed load generation, per-region SLOs

Bài tiếp theo sẽ tìm hiểu Chaos Engineering.