Load testing with k6/Locust, performance profiling, bottleneck identification, K8s cluster tuning, application optimization, and benchmark reporting.
🔒 DevSecOps — Lesson 48
LESSON 48: PERFORMANCE TESTING & OPTIMIZATION
Deploy Microservices On-Premises with Kubernetes HA
Part 12: Production Operations & Capstone Project
xdev.asia
🎯 LESSON OBJECTIVE__HTMLTAG_66___
✅ Load testing với k6 (Grafana k6)
✅ Performance profile types (load, stress, spike, soak)
✅ Bottleneck identification methodology
✅ K8s cluster tuning (kernel, containerd, kubelet)
✅ Application-level optimization
PART 1: LOAD TESTING WITH K6
# Install k6:
# Option 1: Binary
curl -L https://github.com/grafana/k6/releases/download/v0.49.0/k6-v0.49.0-linux-amd64.tar.gz | tar xz
mv k6-v0.49.0-linux-amd64/k6 /usr/local/bin/
# Option 2: Run in K8s:
kubectl run k6-test --image=grafana/k6 --rm -it -- run - < test.js
// load-test.js:
import http from 'k6/http';
import { check, sleep } from 'k6';
import { Rate, Trend } from 'k6/metrics';
const errorRate = new Rate('errors');
const orderLatency = new Trend('order_latency');
export const options = {
stages: [
{ duration: '2m', target: 50 }, // Ramp up
{ duration: '5m', target: 200 }, // Sustained load
{ duration: '2m', target: 500 }, // Peak
{ duration: '3m', target: 200 }, // Scale down
{ duration: '1m', target: 0 }, // Cool down
],
thresholds: {
http_req_duration: ['p(95)<500', 'p(99)<1000'],
errors: ['rate<0.01'],
http_req_failed: ['rate<0.01'],
},
};
export default function () {
// Health check:
const healthRes = http.get('http://api-gateway/health');
check(healthRes, { 'health ok': (r) => r.status === 200 });
// Create order:
const orderPayload = JSON.stringify({
user_id: `user-${__VU}`,
items: [
{ product_id: 'PROD-001', quantity: 1 },
{ product_id: 'PROD-002', quantity: 2 },
],
});
const orderRes = http.post('http://api-gateway/api/v1/orders', orderPayload, {
headers: { 'Content-Type': 'application/json' },
});
check(orderRes, {
'order created': (r) => r.status === 201,
});
errorRate.add(orderRes.status !== 201);
orderLatency.add(orderRes.timings.duration);
sleep(1);
}
# Run test:
k6 run load-test.js
# Run with Prometheus output:
k6 run --out experimental-prometheus-rw load-test.js
# Run with HTML report:
k6 run --out json=results.json load-test.js
PART 2: PERFORMANCE TEST TYPES
Type_ _Goal Pattern_ Duration
Load Test Verify SLO under expected load Gradual ramp to target_ 15-30 min
Stress Test Find breaking point_ Increase until failure_ 30+ min
Spike Test Test sudden traffic burst Instant jump to peak_ 10-15 min
Soak Test Find memory leaks, degradation Sustained moderate load_ 4-24 hours
Breakpoint Test_ Find max throughput_ Step increase until errors Variable
// Stress test (find breaking point):
export const options = {
stages: [
{ duration: '2m', target: 100 },
{ duration: '2m', target: 300 },
{ duration: '2m', target: 500 },
{ duration: '2m', target: 800 },
{ duration: '2m', target: 1000 },
{ duration: '5m', target: 1500 }, // Push beyond expected
{ duration: '2m', target: 0 },
],
};
// Spike test:
export const options = {
stages: [
{ duration: '1m', target: 50 },
{ duration: '10s', target: 1000 }, // Instant spike
{ duration: '3m', target: 1000 },
{ duration: '10s', target: 50 }, // Instant drop
{ duration: '2m', target: 50 },
],
};
PART 3: BOTTLENECK IDENTIFICATION
Bottleneck Analysis Flow:
1. Run load test → observe metrics
2. Check RED metrics per service
3. Find slowest span in trace
Bottleneck Layers:
┌── Application (slow queries, N+1, no caching)
├── Runtime (GC pauses, thread pool exhaustion)
├── Container (CPU throttling, OOM)
├── Kubernetes (scheduling, DNS, service routing)
├── Network (bandwidth, latency, packet loss)
├── Storage (IOPS, throughput, latency)
└── Hardware (CPU, RAM, NIC saturation)
Diagnosis Commands:
# CPU throttling:
cat /sys/fs/cgroup/cpu/cpu.stat | grep throttled
# DNS resolution time:
kubectl exec test-pod -- time nslookup order-service
# Network latency:
kubectl exec test-pod -- curl -w "connect:%{time_connect} ttfb:%{time_starttransfer} total:%{time_total}\n" -s -o /dev/null http://order-service:8080/health
PART 4: CLUSTER & APPLICATION TUNING
# Kernel tuning for high-traffic nodes:
cat >> /etc/sysctl.d/99-k8s-perf.conf << 'EOF'
# Network:
net.core.somaxconn = 65535
net.ipv4.tcp_max_syn_backlog = 65535
net.core.netdev_max_backlog = 65535
net.ipv4.ip_local_port_range = 1024 65535
net.ipv4.tcp_tw_reuse = 1
net.ipv4.tcp_fin_timeout = 15
# File descriptors:
fs.file-max = 2097152
fs.inotify.max_user_watches = 524288
# Memory:
vm.max_map_count = 262144
EOF
sysctl --system
# Container resource optimization:
# Before (over-provisioned):
resources:
requests:
cpu: 1000m
memory: 1Gi
limits:
cpu: 2000m
memory: 2Gi
# After (right-sized based on VPA recommendations):
resources:
requests:
cpu: 250m
memory: 256Mi
limits:
cpu: 500m
memory: 512Mi
# Application tuning:
# - Connection pooling (database, HTTP clients)
# - Add caching layer (Redis)
# - Optimize queries (indexes, query plans)
# - Reduce payload size (compression, pagination)
# - Async processing (message queue for heavy tasks)
PART 5: BENCHMARK REPORTING
Metric_ Target Actual Status
Max throughput_ 1000 req/s 1250 req/s ✅ PASS
P95 latency < 500ms 320ms ✅ PASS_
P99 latency < 1000ms 780ms ✅ PASS_
Error rate < 0.1% 0.02% ✅ PASS_
CPU at peak < 80% 65% ✅ PASS_
Memory at peak < 80% 72% ✅ PASS_
💡 KEY TAKEAWAYS
k6 : Modern load testing tool, scriptable, Grafana integration
Test types : Load, Stress, Spike, Soak — each serves different purpose
Bottleneck analysis : Layer by layer, trace-guided
Right-sizing : VPA recommendations → reduce waste
Kernel tuning : Increase connection limits, file descriptors
Report : Document baseline, compare after optimization
🎯 EXERCISE
Exercise 1: Load Test
Write k6 script for API endpoints__HTMLTAG_276___
Run load test, stress test, spike test__HTMLTAG_278___
Monitor Grafana dashboards during test__HTMLTAG_280___
Exercise 2: Optimization
Identify top 3 bottlenecks from test results
Apply optimizations (caching, right-sizing, tuning)
Re-test and compare improvement
📚 NEXT POST
In Lesson 49: Troubleshooting Guide , we will learn systematic troubleshooting for K8s production.