
Giới thiệu
Bài trước đã xử lý cascade failure từ downstream service. Bài này bổ sung ba pattern quan trọng khác:
- Bulkhead — cô lập resources để một loại request không ảnh hưởng loại khác
- Rate Limiting — kiểm soát tốc độ request để bảo vệ service
- Health Check — giúp Kubernetes biết khi nào service sẵn sàng hay cần restart
1. Bulkhead Pattern
1.1 Ý tưởng từ tàu thủy
Tàu thủy được chia thành nhiều ngăn kín nước (bulkhead). Nếu một ngăn rò rỉ, các ngăn khác vẫn nguyên vẹn — tàu không chìm.
Trong microservices:
Không có Bulkhead:
┌────────────────────────────────────────────┐
│ Order Service │
│ Thread Pool: 50 threads (shared) │
│ │
│ Payment calls: 45 threads (blocked/slow) │
│ Inventory calls: 5 threads remaining │
│ │
│ → Inventory calls cũng bị ảnh hưởng! │
└────────────────────────────────────────────┘
Có Bulkhead:
┌────────────────────────────────────────────┐
│ Order Service │
│ │
│ ┌──────────────────┐ ┌──────────────────┐│
│ │ Payment Pool │ │ Inventory Pool ││
│ │ max: 20 threads │ │ max: 20 threads ││
│ │ │ │ ││
│ │ 19 blocked/slow │ │ 0 blocked ││
│ └──────────────────┘ └──────────────────┘│
│ ↑ Hoàn toàn độc lập! │
└────────────────────────────────────────────┘
1.2 Thread Pool Bulkhead với Resilience4j
// Cấu hình
BulkheadConfig bulkheadConfig = BulkheadConfig.custom()
.maxConcurrentCalls(20) // Tối đa 20 calls đồng thời
.maxWaitDuration(Duration.ofMillis(100)) // Chờ tối đa 100ms nếu pool đầy
.build();
// Hoặc ThreadPoolBulkhead (asynchronous, bounded queue)
ThreadPoolBulkheadConfig tpConfig = ThreadPoolBulkheadConfig.custom()
.maxThreadPoolSize(20)
.coreThreadPoolSize(10)
.queueCapacity(100)
.keepAliveDuration(Duration.ofMillis(20))
.build();
# application.yml
resilience4j:
bulkhead:
instances:
payment-service:
max-concurrent-calls: 20
max-wait-duration: 100ms
inventory-service:
max-concurrent-calls: 30
max-wait-duration: 50ms
thread-pool-bulkhead:
instances:
notification-service:
max-thread-pool-size: 10
core-thread-pool-size: 5
queue-capacity: 200
@CircuitBreaker(name = "payment-service", fallbackMethod = "paymentFallback")
@Bulkhead(name = "payment-service", type = Bulkhead.Type.SEMAPHORE)
public PaymentResponse charge(ChargeRequest request) {
return paymentClient.charge(request);
}
// Fallback khi bulkhead đầy
public PaymentResponse paymentFallback(ChargeRequest req, BulkheadFullException e) {
log.warn("Bulkhead full for payment-service, rejecting request");
throw new ServiceUnavailableException("Payment service busy, please retry");
}
1.3 Connection Pool Bulkhead
Database connection pool là bulkhead tự nhiên — nhưng cần cấu hình per-service:
# Với HikariCP — connection pool cho Order Service
spring:
datasource:
hikari:
maximum-pool-size: 20 # Max 20 connections
minimum-idle: 5 # Luôn giữ 5 connections sẵn
connection-timeout: 2000 # 2s timeout nếu không có connection
idle-timeout: 600000 # Đóng connection không dùng sau 10 phút
max-lifetime: 1800000 # Connection tồn tại tối đa 30 phút
pool-name: "OrderServicePool"
2. Rate Limiting
2.1 Tại sao cần Rate Limiting?
Trường hợp cần:
- Ngăn một client/user gửi quá nhiều request (abuse)
- Bảo vệ downstream service khỏi overload
- Đảm bảo fair usage giữa các clients
- Monetization (free tier vs paid tier)
Ví dụ:
- API Gateway: 1000 req/min per API key
- Login endpoint: 5 attempts/minute per IP (brute force prevention)
- Webhook: 100 req/sec per tenant
2.2 Token Bucket Algorithm
┌────────────────────────────────────────────────────┐
│ Token Bucket │
│ │
│ Capacity: 100 tokens ████████████████████ │
│ Refill rate: 10 tokens/second │
│ │
│ Request đến → lấy 1 token │
│ Nếu có token: ✅ Allow │
│ Nếu trống: ❌ Reject (429 Too Many Requests) │
│ │
│ Bucket tự refill theo thời gian │
└────────────────────────────────────────────────────┘
Ưu điểm:
- Cho phép burst ngắn (dùng hết tokens tích lũy)
- Smooth trong dài hạn
2.3 Sliding Window Algorithm
Window: 1 minute, limit: 100 requests
Timeline: |──────────────────────────────────|
0s 20s 40s 60s
Requests: ||||||||||||||||||||||| |
(100 requests trong 20s đầu)
Request thứ 101 tại 21s:
- Fixed window: ALLOW (window reset về 0) ← Có thể bị abuse
- Sliding window: REJECT (70 requests trong 60s trước) ← Chính xác hơn
2.4 Rate Limiting tại API Gateway (Kong)
# Kong Rate Limiting plugin
apiVersion: configuration.konghq.com/v1
kind: KongPlugin
metadata:
name: rate-limit-per-api-key
plugin: rate-limiting
config:
minute: 1000 # 1000 req/min
hour: 50000 # 50000 req/hour
policy: redis # Lưu counter trong Redis (cluster-aware)
redis_host: redis.platform
redis_port: 6379
hide_client_headers: false # Expose X-RateLimit-* headers to client
# Gắn plugin vào một ingress route
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
annotations:
konghq.com/plugins: rate-limit-per-api-key
2.5 Rate Limiting tại Service Level (Resilience4j)
RateLimiterConfig config = RateLimiterConfig.custom()
.limitForPeriod(100) // 100 permits per period
.limitRefreshPeriod(Duration.ofSeconds(1)) // Refresh mỗi 1s
.timeoutDuration(Duration.ofMillis(0)) // Reject ngay nếu không có permit
.build();
RateLimiter rateLimiter = RateLimiterRegistry.of(config)
.rateLimiter("order-creation");
@RateLimiter(name = "order-creation", fallbackMethod = "rateLimitFallback")
public Order createOrder(CreateOrderRequest request) {
return orderService.create(request);
}
public Order rateLimitFallback(CreateOrderRequest req, RequestNotPermitted e) {
throw new TooManyRequestsException("Rate limit exceeded, please retry later");
}
2.6 Response Headers cho Rate Limiting
HTTP/1.1 200 OK
X-RateLimit-Limit: 1000
X-RateLimit-Remaining: 743
X-RateLimit-Reset: 1711879200
# Khi bị reject:
HTTP/1.1 429 Too Many Requests
X-RateLimit-Limit: 1000
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 1711879200
Retry-After: 47
Content-Type: application/json
{"error": "rate_limit_exceeded", "retry_after": 47}
3. Health Check Pattern
3.1 Ba loại Probe trong Kubernetes
spec:
containers:
- name: order-service
# Startup Probe: Service đã khởi động xong chưa?
# K8s chờ startupProbe pass trước khi dùng liveness/readiness
startupProbe:
httpGet:
path: /actuator/health/liveness
port: 8080
failureThreshold: 30 # 30 × 2s = 60s để start
periodSeconds: 2
# Liveness Probe: Service có đang chạy không?
# Fail → K8s restart container
livenessProbe:
httpGet:
path: /actuator/health/liveness
port: 8080
initialDelaySeconds: 0 # Sau khi startupProbe pass
periodSeconds: 10
failureThreshold: 3 # Restart sau 3 lần fail liên tiếp
timeoutSeconds: 3
# Readiness Probe: Service có sẵn sàng nhận traffic không?
# Fail → K8s remove khỏi Service endpoints (không gửi traffic)
readinessProbe:
httpGet:
path: /actuator/health/readiness
port: 8080
initialDelaySeconds: 0
periodSeconds: 5
failureThreshold: 3
successThreshold: 1
timeoutSeconds: 3
3.2 Liveness vs Readiness — Sự khác biệt quan trọng
LIVENESS: "Tôi có đang sống không?"
→ Fail → RESTART container
→ Check: không bị deadlock, heap không bị full, process không bị hang
→ Ví dụ fail: OutOfMemoryError, deadlock trong thread pool
READINESS: "Tôi có sẵn sàng nhận request không?"
→ Fail → REMOVE from load balancer (K8s Service endpoints)
→ Container không bị restart, vẫn chạy
→ Check: DB connection có sẵn, required caches đã warm, dependencies sẵn sàng
→ Ví dụ fail: DB connection pool full, warming up cache, throttled by upstream
3.3 Spring Boot Actuator Health
// Tự động check:
// - Database connectivity
// - Disk space
// - Redis connectivity
// - Kafka connectivity (nếu dùng)
// Custom health indicator
@Component
public class OrderProcessorHealthIndicator implements HealthIndicator {
private final OrderQueue queue;
@Override
public Health health() {
int queueSize = queue.size();
if (queueSize > 10000) {
return Health.down()
.withDetail("queue_size", queueSize)
.withDetail("reason", "Queue backlog too large")
.build();
}
return Health.up()
.withDetail("queue_size", queueSize)
.build();
}
}
# application.yml — cấu hình health groups
management:
endpoint:
health:
show-details: always
group:
liveness:
include: livenessState
# Chỉ check internal state, không check external dependencies
readiness:
include: readinessState, db, redis, kafka
# Check đủ dependencies trước khi nhận traffic
health:
db:
enabled: true
redis:
enabled: true
Response /actuator/health/readiness:
{
"status": "UP",
"components": {
"db": { "status": "UP", "details": { "database": "PostgreSQL", "result": 1 } },
"redis": { "status": "UP" },
"kafka": { "status": "UP" },
"readinessState": { "status": "UP" }
}
}
3.4 Graceful Shutdown
# application.yml
server:
shutdown: graceful # Chờ requests đang xử lý hoàn thành
spring:
lifecycle:
timeout-per-shutdown-phase: 30s # Tối đa 30s chờ
# Kubernetes terminationGracePeriodSeconds phải > timeout-per-shutdown-phase
spec:
terminationGracePeriodSeconds: 60
@Bean
public GracefulShutdown gracefulShutdown() {
return new GracefulShutdown();
}
// Sequence khi K8s gửi SIGTERM:
// 1. Readiness probe → FAIL (K8s ngừng gửi traffic)
// 2. Spring starts graceful shutdown
// 3. Hoàn thành requests đang xử lý
// 4. Close connections (DB, Kafka, Redis)
// 5. Application exit(0)
4. Combine All Patterns
Một service production-ready kết hợp tất cả:
@Service
public class PaymentServiceClient {
@Retry(name = "payment")
@CircuitBreaker(name = "payment", fallbackMethod = "fallback")
@Bulkhead(name = "payment", type = Type.SEMAPHORE)
@RateLimiter(name = "payment-outbound")
public PaymentResponse charge(ChargeRequest request) {
return paymentClient.post("/charge", request, PaymentResponse.class);
}
public PaymentResponse fallback(ChargeRequest req, Throwable t) {
if (t instanceof BulkheadFullException) {
return PaymentResponse.queued(req.getOrderId());
}
if (t instanceof CallNotPermittedException) {
return PaymentResponse.pending(req.getOrderId());
}
return PaymentResponse.failed(req.getOrderId(), t.getMessage());
}
}
Thứ tự thực thi (quan trọng!)
Request đến
↓
RateLimiter → Có đủ permit không?
↓
Bulkhead → Có slot không?
↓
CircuitBreaker → Circuit có CLOSED không?
↓
TimeLimiter → Đặt timeout
↓
Retry → Thực thi và retry nếu fail
↓
Actual call đến Payment Service
Khi wrap bằng annotations, thứ tự là reverse: Retry → CircuitBreaker → Bulkhead → RateLimiter.
5. Monitoring tổng hợp
# Bulkhead đang gần đầy
resilience4j_bulkhead_available_concurrent_calls{name="payment-service"} < 3
# Rate limiter đang throttle nhiều requests
rate(resilience4j_ratelimiter_waiting_threads{name="payment-outbound"}[5m]) > 0
# Restart count cao trong 1 giờ
increase(kube_pod_container_status_restarts_total{
namespace="services-prod"
}[1h]) > 3
Tóm tắt
| Pattern | Cơ chế | Bảo vệ khỏi |
|---|---|---|
| Bulkhead | Thread pool / semaphore isolation | Resource starvation chéo |
| Token Bucket | Token accumulation | Bursty traffic |
| Sliding Window | Rolling count | Sustained high traffic |
| Liveness Probe | Internal state check | Stuck/deadlock process |
| Readiness Probe | External dependency check | Traffic khi chưa sẵn sàng |
| Graceful Shutdown | Drain in-flight requests | Request loss khi deploy |
Bài tiếp theo: Chaos Engineering — Kiểm chứng độ tin cậy hệ thống