Chuyển đến nội dung chính

BÀI 48: PERFORMANCE TESTING & OPTIMIZATION

Load testing với k6/Locust, performance profiling, bottleneck identification, K8s cluster tuning, application optimization, và benchmark reporting.

🔒 DevSecOps — Bài 48 BÀI 48: PERFORMANCE TESTING & OPTIMIZATION

Deploy Microservices On-Premises với Kubernetes HA

Phần 12: Production Operations & Capstone Project

xdev.asia

🎯 MỤC TIÊU BÀI HỌC

  • ✅ Load testing với k6 (Grafana k6)
  • ✅ Performance profile types (load, stress, spike, soak)
  • ✅ Bottleneck identification methodology
  • ✅ K8s cluster tuning (kernel, containerd, kubelet)
  • ✅ Application-level optimization

PHẦN 1: LOAD TESTING VỚI K6

# Install k6:
# Option 1: Binary
curl -L https://github.com/grafana/k6/releases/download/v0.49.0/k6-v0.49.0-linux-amd64.tar.gz | tar xz
mv k6-v0.49.0-linux-amd64/k6 /usr/local/bin/

# Option 2: Run in K8s:
kubectl run k6-test --image=grafana/k6 --rm -it -- run - < test.js
// load-test.js:
import http from 'k6/http';
import { check, sleep } from 'k6';
import { Rate, Trend } from 'k6/metrics';

const errorRate = new Rate('errors');
const orderLatency = new Trend('order_latency');

export const options = {
  stages: [
    { duration: '2m', target: 50 },    // Ramp up
    { duration: '5m', target: 200 },   // Sustained load
    { duration: '2m', target: 500 },   // Peak
    { duration: '3m', target: 200 },   // Scale down
    { duration: '1m', target: 0 },     // Cool down
  ],
  thresholds: {
    http_req_duration: ['p(95)<500', 'p(99)<1000'],
    errors: ['rate<0.01'],
    http_req_failed: ['rate<0.01'],
  },
};

export default function () {
  // Health check:
  const healthRes = http.get('http://api-gateway/health');
  check(healthRes, { 'health ok': (r) => r.status === 200 });

  // Create order:
  const orderPayload = JSON.stringify({
    user_id: `user-${__VU}`,
    items: [
      { product_id: 'PROD-001', quantity: 1 },
      { product_id: 'PROD-002', quantity: 2 },
    ],
  });

  const orderRes = http.post('http://api-gateway/api/v1/orders', orderPayload, {
    headers: { 'Content-Type': 'application/json' },
  });

  check(orderRes, {
    'order created': (r) => r.status === 201,
  });

  errorRate.add(orderRes.status !== 201);
  orderLatency.add(orderRes.timings.duration);

  sleep(1);
}
# Run test:
k6 run load-test.js

# Run with Prometheus output:
k6 run --out experimental-prometheus-rw load-test.js

# Run with HTML report:
k6 run --out json=results.json load-test.js

PHẦN 2: PERFORMANCE TEST TYPES

TypeGoalPatternDuration
Load TestVerify SLO under expected loadGradual ramp to target15-30 min
Stress TestFind breaking pointIncrease until failure30+ min
Spike TestTest sudden traffic burstInstant jump to peak10-15 min
Soak TestFind memory leaks, degradationSustained moderate load4-24 hours
Breakpoint TestFind max throughputStep increase until errorsVariable
// Stress test (find breaking point):
export const options = {
  stages: [
    { duration: '2m', target: 100 },
    { duration: '2m', target: 300 },
    { duration: '2m', target: 500 },
    { duration: '2m', target: 800 },
    { duration: '2m', target: 1000 },
    { duration: '5m', target: 1500 },  // Push beyond expected
    { duration: '2m', target: 0 },
  ],
};

// Spike test:
export const options = {
  stages: [
    { duration: '1m', target: 50 },
    { duration: '10s', target: 1000 },   // Instant spike
    { duration: '3m', target: 1000 },
    { duration: '10s', target: 50 },     // Instant drop
    { duration: '2m', target: 50 },
  ],
};

PHẦN 3: BOTTLENECK IDENTIFICATION


Bottleneck Analysis Flow:

1. Run load test → observe metrics
2. Check RED metrics per service
3. Find slowest span in trace

Bottleneck Layers:
┌── Application  (slow queries, N+1, no caching)
├── Runtime      (GC pauses, thread pool exhaustion)
├── Container    (CPU throttling, OOM)
├── Kubernetes   (scheduling, DNS, service routing)
├── Network      (bandwidth, latency, packet loss)
├── Storage      (IOPS, throughput, latency)
└── Hardware     (CPU, RAM, NIC saturation)

Diagnosis Commands:
# CPU throttling:
cat /sys/fs/cgroup/cpu/cpu.stat | grep throttled

# DNS resolution time:
kubectl exec test-pod -- time nslookup order-service

# Network latency:
kubectl exec test-pod -- curl -w "connect:%{time_connect} ttfb:%{time_starttransfer} total:%{time_total}\n" -s -o /dev/null http://order-service:8080/health

PHẦN 4: CLUSTER & APPLICATION TUNING

# Kernel tuning for high-traffic nodes:
cat >> /etc/sysctl.d/99-k8s-perf.conf << 'EOF'
# Network:
net.core.somaxconn = 65535
net.ipv4.tcp_max_syn_backlog = 65535
net.core.netdev_max_backlog = 65535
net.ipv4.ip_local_port_range = 1024 65535
net.ipv4.tcp_tw_reuse = 1
net.ipv4.tcp_fin_timeout = 15

# File descriptors:
fs.file-max = 2097152
fs.inotify.max_user_watches = 524288

# Memory:
vm.max_map_count = 262144
EOF

sysctl --system
# Container resource optimization:
# Before (over-provisioned):
resources:
  requests:
    cpu: 1000m
    memory: 1Gi
  limits:
    cpu: 2000m
    memory: 2Gi

# After (right-sized based on VPA recommendations):
resources:
  requests:
    cpu: 250m
    memory: 256Mi
  limits:
    cpu: 500m
    memory: 512Mi

# Application tuning:
# - Connection pooling (database, HTTP clients)
# - Add caching layer (Redis)
# - Optimize queries (indexes, query plans)
# - Reduce payload size (compression, pagination)
# - Async processing (message queue for heavy tasks)

PHẦN 5: BENCHMARK REPORTING

MetricTargetActualStatus
Max throughput1000 req/s1250 req/s✅ PASS
P95 latency< 500ms320ms✅ PASS
P99 latency< 1000ms780ms✅ PASS
Error rate< 0.1%0.02%✅ PASS
CPU at peak< 80%65%✅ PASS
Memory at peak< 80%72%✅ PASS

💡 KEY TAKEAWAYS

  1. k6: Modern load testing tool, scriptable, Grafana integration
  2. Test types: Load, Stress, Spike, Soak — each serves different purpose
  3. Bottleneck analysis: Layer by layer, trace-guided
  4. Right-sizing: VPA recommendations → reduce waste
  5. Kernel tuning: Increase connection limits, file descriptors
  6. Report: Document baseline, compare after optimization

🎯 BÀI TẬP

Bài tập 1: Load Test

  • Write k6 script for API endpoints
  • Run load test, stress test, spike test
  • Monitor Grafana dashboards during test

Bài tập 2: Optimization

  • Identify top 3 bottlenecks from test results
  • Apply optimizations (caching, right-sizing, tuning)
  • Re-test and compare improvement

📚 BÀI TIẾP THEO

Trong Bài 49: Troubleshooting Guide, chúng ta sẽ học systematic troubleshooting cho K8s production.