
Introduction
The previous article dealt with cascade failures from downstream services. This article adds three other important patterns:
- Bulkhead — isolates resources so that one type of request does not affect another
- Rate Limiting — control request rate to protect service
- Health Check — helps Kubernetes know when the service is ready or needs to be restarted
1. Bulkhead Pattern
1.1 Ideas from ships
Ships are divided into several watertight compartments (bulkheads). If one compartment leaks, the other compartments remain intact — the ship does not sink.
In microservices:
Không có Bulkhead:
┌────────────────────────────────────────────┐
│ Order Service │
│ Thread Pool: 50 threads (shared) │
│ │
│ Payment calls: 45 threads (blocked/slow) │
│ Inventory calls: 5 threads remaining │
│ │
│ → Inventory calls cũng bị ảnh hưởng! │
└────────────────────────────────────────────┘
Có Bulkhead:
┌────────────────────────────────────────────┐
│ Order Service │
│ │
│ ┌──────────────────┐ ┌──────────────────┐│
│ │ Payment Pool │ │ Inventory Pool ││
│ │ max: 20 threads │ │ max: 20 threads ││
│ │ │ │ ││
│ │ 19 blocked/slow │ │ 0 blocked ││
│ └──────────────────┘ └──────────────────┘│
│ ↑ Hoàn toàn độc lập! │
└────────────────────────────────────────────┘
1.2 Thread Pool Bulkhead with Resilience4j
// Cấu hình
BulkheadConfig bulkheadConfig = BulkheadConfig.custom()
.maxConcurrentCalls(20) // Tối đa 20 calls đồng thời
.maxWaitDuration(Duration.ofMillis(100)) // Chờ tối đa 100ms nếu pool đầy
.build();
// Hoặc ThreadPoolBulkhead (asynchronous, bounded queue)
ThreadPoolBulkheadConfig tpConfig = ThreadPoolBulkheadConfig.custom()
.maxThreadPoolSize(20)
.coreThreadPoolSize(10)
.queueCapacity(100)
.keepAliveDuration(Duration.ofMillis(20))
.build();
# application.yml
resilience4j:
bulkhead:
instances:
payment-service:
max-concurrent-calls: 20
max-wait-duration: 100ms
inventory-service:
max-concurrent-calls: 30
max-wait-duration: 50ms
thread-pool-bulkhead:
instances:
notification-service:
max-thread-pool-size: 10
core-thread-pool-size: 5
queue-capacity: 200
@CircuitBreaker(name = "payment-service", fallbackMethod = "paymentFallback")
@Bulkhead(name = "payment-service", type = Bulkhead.Type.SEMAPHORE)
public PaymentResponse charge(ChargeRequest request) {
return paymentClient.charge(request);
}
// Fallback khi bulkhead đầy
public PaymentResponse paymentFallback(ChargeRequest req, BulkheadFullException e) {
log.warn("Bulkhead full for payment-service, rejecting request");
throw new ServiceUnavailableException("Payment service busy, please retry");
}
1.3 Connection Pool Bulkhead
Database connection pool is naturally bulkhead — but requires per-service configuration:
# Với HikariCP — connection pool cho Order Service
spring:
datasource:
hikari:
maximum-pool-size: 20 # Max 20 connections
minimum-idle: 5 # Luôn giữ 5 connections sẵn
connection-timeout: 2000 # 2s timeout nếu không có connection
idle-timeout: 600000 # Đóng connection không dùng sau 10 phút
max-lifetime: 1800000 # Connection tồn tại tối đa 30 phút
pool-name: "OrderServicePool"
2. Rate Limiting
2.1 Why is Rate Limiting needed?
Trường hợp cần:
- Ngăn một client/user gửi quá nhiều request (abuse)
- Bảo vệ downstream service khỏi overload
- Đảm bảo fair usage giữa các clients
- Monetization (free tier vs paid tier)
Ví dụ:
- API Gateway: 1000 req/min per API key
- Login endpoint: 5 attempts/minute per IP (brute force prevention)
- Webhook: 100 req/sec per tenant
2.2 Token Bucket Algorithm
┌────────────────────────────────────────────────────┐
│ Token Bucket │
│ │
│ Capacity: 100 tokens ████████████████████ │
│ Refill rate: 10 tokens/second │
│ │
│ Request đến → lấy 1 token │
│ Nếu có token: ✅ Allow │
│ Nếu trống: ❌ Reject (429 Too Many Requests) │
│ │
│ Bucket tự refill theo thời gian │
└────────────────────────────────────────────────────┘
Ưu điểm:
- Cho phép burst ngắn (dùng hết tokens tích lũy)
- Smooth trong dài hạn
2.3 Sliding Window Algorithm
Window: 1 minute, limit: 100 requests
Timeline: |──────────────────────────────────|
0s 20s 40s 60s
Requests: ||||||||||||||||||||||| |
(100 requests trong 20s đầu)
Request thứ 101 tại 21s:
- Fixed window: ALLOW (window reset về 0) ← Có thể bị abuse
- Sliding window: REJECT (70 requests trong 60s trước) ← Chính xác hơn
2.4 Rate Limiting at API Gateway (Kong)
# Kong Rate Limiting plugin
apiVersion: configuration.konghq.com/v1
kind: KongPlugin
metadata:
name: rate-limit-per-api-key
plugin: rate-limiting
config:
minute: 1000 # 1000 req/min
hour: 50000 # 50000 req/hour
policy: redis # Lưu counter trong Redis (cluster-aware)
redis_host: redis.platform
redis_port: 6379
hide_client_headers: false # Expose X-RateLimit-* headers to client
# Gắn plugin vào một ingress route
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
annotations:
konghq.com/plugins: rate-limit-per-api-key
2.5 Rate Limiting at Service Level (Resilience4j)
RateLimiterConfig config = RateLimiterConfig.custom()
.limitForPeriod(100) // 100 permits per period
.limitRefreshPeriod(Duration.ofSeconds(1)) // Refresh mỗi 1s
.timeoutDuration(Duration.ofMillis(0)) // Reject ngay nếu không có permit
.build();
RateLimiter rateLimiter = RateLimiterRegistry.of(config)
.rateLimiter("order-creation");
@RateLimiter(name = "order-creation", fallbackMethod = "rateLimitFallback")
public Order createOrder(CreateOrderRequest request) {
return orderService.create(request);
}
public Order rateLimitFallback(CreateOrderRequest req, RequestNotPermitted e) {
throw new TooManyRequestsException("Rate limit exceeded, please retry later");
}
2.6 Response Headers for Rate Limiting
HTTP/1.1 200 OK
X-RateLimit-Limit: 1000
X-RateLimit-Remaining: 743
X-RateLimit-Reset: 1711879200
# Khi bị reject:
HTTP/1.1 429 Too Many Requests
X-RateLimit-Limit: 1000
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 1711879200
Retry-After: 47
Content-Type: application/json
{"error": "rate_limit_exceeded", "retry_after": 47}
3. Health Check Pattern
3.1 Three types of Probes in Kubernetes
spec:
containers:
- name: order-service
# Startup Probe: Service đã khởi động xong chưa?
# K8s chờ startupProbe pass trước khi dùng liveness/readiness
startupProbe:
httpGet:
path: /actuator/health/liveness
port: 8080
failureThreshold: 30 # 30 × 2s = 60s để start
periodSeconds: 2
# Liveness Probe: Service có đang chạy không?
# Fail → K8s restart container
livenessProbe:
httpGet:
path: /actuator/health/liveness
port: 8080
initialDelaySeconds: 0 # Sau khi startupProbe pass
periodSeconds: 10
failureThreshold: 3 # Restart sau 3 lần fail liên tiếp
timeoutSeconds: 3
# Readiness Probe: Service có sẵn sàng nhận traffic không?
# Fail → K8s remove khỏi Service endpoints (không gửi traffic)
readinessProbe:
httpGet:
path: /actuator/health/readiness
port: 8080
initialDelaySeconds: 0
periodSeconds: 5
failureThreshold: 3
successThreshold: 1
timeoutSeconds: 3
3.2 Liveness vs Readiness — The important difference
LIVENESS: "Tôi có đang sống không?"
→ Fail → RESTART container
→ Check: không bị deadlock, heap không bị full, process không bị hang
→ Ví dụ fail: OutOfMemoryError, deadlock trong thread pool
READINESS: "Tôi có sẵn sàng nhận request không?"
→ Fail → REMOVE from load balancer (K8s Service endpoints)
→ Container không bị restart, vẫn chạy
→ Check: DB connection có sẵn, required caches đã warm, dependencies sẵn sàng
→ Ví dụ fail: DB connection pool full, warming up cache, throttled by upstream
3.3 Spring Boot Actuator Health
// Tự động check:
// - Database connectivity
// - Disk space
// - Redis connectivity
// - Kafka connectivity (nếu dùng)
// Custom health indicator
@Component
public class OrderProcessorHealthIndicator implements HealthIndicator {
private final OrderQueue queue;
@Override
public Health health() {
int queueSize = queue.size();
if (queueSize > 10000) {
return Health.down()
.withDetail("queue_size", queueSize)
.withDetail("reason", "Queue backlog too large")
.build();
}
return Health.up()
.withDetail("queue_size", queueSize)
.build();
}
}
# application.yml — cấu hình health groups
management:
endpoint:
health:
show-details: always
group:
liveness:
include: livenessState
# Chỉ check internal state, không check external dependencies
readiness:
include: readinessState, db, redis, kafka
# Check đủ dependencies trước khi nhận traffic
health:
db:
enabled: true
redis:
enabled: true
Response /actuator/health/readiness:
{
"status": "UP",
"components": {
"db": { "status": "UP", "details": { "database": "PostgreSQL", "result": 1 } },
"redis": { "status": "UP" },
"kafka": { "status": "UP" },
"readinessState": { "status": "UP" }
}
}
3.4 Graceful Shutdown
# application.yml
server:
shutdown: graceful # Chờ requests đang xử lý hoàn thành
spring:
lifecycle:
timeout-per-shutdown-phase: 30s # Tối đa 30s chờ
# Kubernetes terminationGracePeriodSeconds phải > timeout-per-shutdown-phase
spec:
terminationGracePeriodSeconds: 60
@Bean
public GracefulShutdown gracefulShutdown() {
return new GracefulShutdown();
}
// Sequence khi K8s gửi SIGTERM:
// 1. Readiness probe → FAIL (K8s ngừng gửi traffic)
// 2. Spring starts graceful shutdown
// 3. Hoàn thành requests đang xử lý
// 4. Close connections (DB, Kafka, Redis)
// 5. Application exit(0)
4. Combine All Patterns
A production-ready service combines it all:
@Service
public class PaymentServiceClient {
@Retry(name = "payment")
@CircuitBreaker(name = "payment", fallbackMethod = "fallback")
@Bulkhead(name = "payment", type = Type.SEMAPHORE)
@RateLimiter(name = "payment-outbound")
public PaymentResponse charge(ChargeRequest request) {
return paymentClient.post("/charge", request, PaymentResponse.class);
}
public PaymentResponse fallback(ChargeRequest req, Throwable t) {
if (t instanceof BulkheadFullException) {
return PaymentResponse.queued(req.getOrderId());
}
if (t instanceof CallNotPermittedException) {
return PaymentResponse.pending(req.getOrderId());
}
return PaymentResponse.failed(req.getOrderId(), t.getMessage());
}
}
Execution order (important!)
Request đến
↓
RateLimiter → Có đủ permit không?
↓
Bulkhead → Có slot không?
↓
CircuitBreaker → Circuit có CLOSED không?
↓
TimeLimiter → Đặt timeout
↓
Retry → Thực thi và retry nếu fail
↓
Actual call đến Payment Service
When wrapping with annotations, the order is reverse: Retry → CircuitBreaker → Bulkhead → RateLimiter.
5. General monitoring
# Bulkhead đang gần đầy
resilience4j_bulkhead_available_concurrent_calls{name="payment-service"} < 3
# Rate limiter đang throttle nhiều requests
rate(resilience4j_ratelimiter_waiting_threads{name="payment-outbound"}[5m]) > 0
# Restart count cao trong 1 giờ
increase(kube_pod_container_status_restarts_total{
namespace="services-prod"
}[1h]) > 3
Summary
| Pattern | Mechanism | Protect from |
|---|---|---|
| Bulkhead | Thread pool / semaphore isolation | Resource starvation cross |
| Token Bucket | Token accumulation | Bursty traffic |
| Sliding Window | Rolling count | Sustained high traffic |
| Liveness Probe | Internal state check | Stuck/deadlock process |
| Readiness Probe | External dependency check | Traffic when not ready |
| Graceful Shutdown | Drain in-flight requests | Request loss when deploying |
Next article: Chaos Engineering — Verifying system reliability