1. Kubernetes Auto-scaling Testing
# HPA configuration
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: api-hpa
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: api-server
minReplicas: 3
maxReplicas: 50
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
- type: Pods
pods:
metric:
name: http_requests_per_second
target:
type: AverageValue
averageValue: 500 # Scale khi > 500 RPS / pod
behavior:
scaleUp:
stabilizationWindowSeconds: 30
policies:
- type: Percent
value: 100 # Double pods mỗi lần scale
periodSeconds: 60
scaleDown:
stabilizationWindowSeconds: 300
policies:
- type: Percent
value: 10
periodSeconds: 60
// k6 — Test HPA scale-out behavior
export const options = {
scenarios: {
hpa_test: {
executor: 'ramping-arrival-rate',
startRate: 50,
timeUnit: '1s',
stages: [
{ duration: '2m', target: 50 }, // Baseline (3 pods)
{ duration: '3m', target: 500 }, // Trigger scale-out
{ duration: '5m', target: 1000 }, // Max scale
{ duration: '5m', target: 1000 }, // Sustain peak
{ duration: '5m', target: 50 }, // Scale-down
{ duration: '5m', target: 50 }, // Verify scale-down
],
preAllocatedVUs: 200,
maxVUs: 1000,
},
},
};
// Monitor cùng lúc:
// kubectl get hpa -w
// kubectl get pods -w
// Grafana dashboard: pods count + latency + RPS
2. Serverless Cold Start Profiling
// Test Lambda cold start vs warm invocation
import http from 'k6/http';
import { Trend } from 'k6/metrics';
const coldStart = new Trend('cold_start_duration');
const warmStart = new Trend('warm_start_duration');
export const options = {
scenarios: {
// Scenario 1: Sequential calls (force cold starts)
cold_start_test: {
executor: 'per-vu-iterations',
vus: 50, // 50 concurrent → 50 new containers
iterations: 1,
maxDuration: '2m',
},
// Scenario 2: Rapid calls (keep warm)
warm_test: {
executor: 'constant-arrival-rate',
rate: 100,
timeUnit: '1s',
duration: '5m',
preAllocatedVUs: 20,
startTime: '3m',
},
},
};
export default function () {
const res = http.get(`${LAMBDA_URL}/api/process`);
// Check X-Cold-Start header hoặc init duration
const isCold = res.headers['X-Cold-Start'] === 'true';
if (isCold) {
coldStart.add(res.timings.duration);
} else {
warmStart.add(res.timings.duration);
}
}
Cold Start Optimization Strategies:
├── Provisioned Concurrency (AWS Lambda)
├── Min instances (Cloud Functions)
├── Smaller deployment package
├── Lazy initialization
├── SnapStart (Java Lambda)
└── Warm-up cron (scheduled ping)
3. Service Mesh Overhead Measurement
Test: So sánh latency có/không có sidecar proxy
Scenario A: Direct pod-to-pod (no mesh)
Service A → Service B
p95: 5ms
Scenario B: With Istio sidecar
Service A → Envoy proxy → Envoy proxy → Service B
p95: 12ms
→ Overhead: ~7ms (140% increase)
→ Accept nếu thêm mTLS, observability, traffic management
Metrics cần đo:
├── Sidecar CPU/memory overhead per pod
├── Added latency per hop
├── Connection pool exhaustion
├── mTLS handshake overhead
└── Control plane (istiod) resource usage
4. CDN Performance Validation
// Test CDN cache hit ratio và origin offload
import http from 'k6/http';
import { Counter } from 'k6/metrics';
const cacheHit = new Counter('cdn_cache_hit');
const cacheMiss = new Counter('cdn_cache_miss');
export default function () {
// Test từ nhiều regions (k6 Cloud)
const res = http.get('https://cdn.example.com/static/bundle.js');
const cacheStatus = res.headers['X-Cache'] || res.headers['Cf-Cache-Status'];
if (cacheStatus === 'HIT') {
cacheHit.add(1);
} else {
cacheMiss.add(1);
}
// Measure TTFB (Time to First Byte)
// res.timings.waiting = TTFB
check(res, {
'TTFB < 50ms': (r) => r.timings.waiting < 50,
'Total < 200ms': (r) => r.timings.duration < 200,
});
}
5. Multi-region Latency Testing
# k6 Cloud — Distributed load zones
export const options = {
ext: {
loadimpact: {
distribution: {
'amazon:us:ashburn': { loadZone: 'amazon:us:ashburn', percent: 30 },
'amazon:eu:frankfurt': { loadZone: 'amazon:eu:frankfurt', percent: 25 },
'amazon:ap:singapore': { loadZone: 'amazon:ap:singapore', percent: 25 },
'amazon:ap:sydney': { loadZone: 'amazon:ap:sydney', percent: 20 },
},
},
},
};
Multi-region Testing Checklist:
☐ DNS resolution time per region
☐ TCP connect time (geographic distance)
☐ TLS handshake overhead
☐ TTFB per region
☐ Data transfer time (payload size × bandwidth)
☐ CDN cache hit ratio per POP
☐ Database read latency (cross-region reads)
☐ Write replication lag (multi-master)
6. Tổng kết
- HPA Testing: Validate scale-out speed, warm-up time, scale-down behavior
- Serverless: Cold start profiling, provisioned concurrency testing
- Service Mesh: Measure sidecar overhead per hop
- CDN: Cache hit ratio, TTFB, origin offload
- Multi-region: Distributed load generation, per-region SLOs
Bài tiếp theo sẽ tìm hiểu Chaos Engineering.