Chuyển đến nội dung chính

レッスン 19: バルクヘッド、レート制限、ヘルス チェック パターン

バルクヘッド パターン (スレッド プール分離)、レート制限アルゴリズム (トークン バケット、スライディング ウィンドウ)、ヘルス チェック パターン (活性、準備完了、起動プローブ)、グレースフル デグラデーション戦略。

🏗️ アーキテクチャ — レッスン 19 レッスン 19: バルクヘッド、レート制限、ヘルス チェックパターン

クラウドネイティブのマイクロサービスアーキテクチャ

パート 6: 回復力パターン

xdev.asia

レッスン 19: バルクヘッド、レート制限、ヘルス チェック パターン

はじめに

前回の記事では、ダウンストリーム サービスからのカスケード障害について説明しました。この記事では、他に 3 つの重要なパターンを追加します。

  • バルクヘッド — ある種類のリクエストが別の種類のリクエストに影響を与えないようにリソースを分離します。
  • レート制限 — サービスを保護するためにリクエスト レートを制御します
  • ヘルスチェック — Kubernetes がサービスの準備ができたとき、または再起動が必要なときを認識するのに役立ちます

1. 隔壁パターン

1.1 船舶からのアイデア

船はいくつかの水密区画(隔壁)に分かれています。 1 つのコンパートメントが漏れても、他のコンパートメントは無傷であり、船は沈みません。

マイクロサービスでは:

Không có Bulkhead:
┌────────────────────────────────────────────┐
│            Order Service                   │
│  Thread Pool: 50 threads (shared)          │
│                                            │
│  Payment calls: 45 threads (blocked/slow) │
│  Inventory calls: 5 threads remaining     │
│                                            │
│  → Inventory calls cũng bị ảnh hưởng!    │
└────────────────────────────────────────────┘

Có Bulkhead:
┌────────────────────────────────────────────┐
│            Order Service                   │
│                                            │
│  ┌──────────────────┐ ┌──────────────────┐│
│  │ Payment Pool     │ │ Inventory Pool   ││
│  │ max: 20 threads  │ │ max: 20 threads  ││
│  │                  │ │                  ││
│  │ 19 blocked/slow  │ │ 0 blocked        ││
│  └──────────────────┘ └──────────────────┘│
│                      ↑ Hoàn toàn độc lập! │
└────────────────────────────────────────────┘

Resilience4j を備えた 1.2 スレッド プール バルクヘッド

// Cấu hình
BulkheadConfig bulkheadConfig = BulkheadConfig.custom()
    .maxConcurrentCalls(20)           // Tối đa 20 calls đồng thời
    .maxWaitDuration(Duration.ofMillis(100))  // Chờ tối đa 100ms nếu pool đầy
    .build();

// Hoặc ThreadPoolBulkhead (asynchronous, bounded queue)
ThreadPoolBulkheadConfig tpConfig = ThreadPoolBulkheadConfig.custom()
    .maxThreadPoolSize(20)
    .coreThreadPoolSize(10)
    .queueCapacity(100)
    .keepAliveDuration(Duration.ofMillis(20))
    .build();
# application.yml
resilience4j:
  bulkhead:
    instances:
      payment-service:
        max-concurrent-calls: 20
        max-wait-duration: 100ms
      inventory-service:
        max-concurrent-calls: 30
        max-wait-duration: 50ms
  thread-pool-bulkhead:
    instances:
      notification-service:
        max-thread-pool-size: 10
        core-thread-pool-size: 5
        queue-capacity: 200
@CircuitBreaker(name = "payment-service", fallbackMethod = "paymentFallback")
@Bulkhead(name = "payment-service", type = Bulkhead.Type.SEMAPHORE)
public PaymentResponse charge(ChargeRequest request) {
    return paymentClient.charge(request);
}

// Fallback khi bulkhead đầy
public PaymentResponse paymentFallback(ChargeRequest req, BulkheadFullException e) {
    log.warn("Bulkhead full for payment-service, rejecting request");
    throw new ServiceUnavailableException("Payment service busy, please retry");
}

1.3 接続プールのバルクヘッド

データベース接続プールは当然バルクヘッドですが、サービスごとの構成が必要です。

# Với HikariCP — connection pool cho Order Service
spring:
  datasource:
    hikari:
      maximum-pool-size: 20      # Max 20 connections
      minimum-idle: 5            # Luôn giữ 5 connections sẵn
      connection-timeout: 2000   # 2s timeout nếu không có connection
      idle-timeout: 600000       # Đóng connection không dùng sau 10 phút
      max-lifetime: 1800000      # Connection tồn tại tối đa 30 phút
      pool-name: "OrderServicePool"

2. レート制限

2.1 レート制限はなぜ必要ですか?

Trường hợp cần:
- Ngăn một client/user gửi quá nhiều request (abuse)
- Bảo vệ downstream service khỏi overload
- Đảm bảo fair usage giữa các clients
- Monetization (free tier vs paid tier)

Ví dụ:
- API Gateway: 1000 req/min per API key
- Login endpoint: 5 attempts/minute per IP (brute force prevention)
- Webhook: 100 req/sec per tenant

2.2 トークンバケットアルゴリズム

┌────────────────────────────────────────────────────┐
│                  Token Bucket                      │
│                                                    │
│  Capacity: 100 tokens  ████████████████████        │
│  Refill rate: 10 tokens/second                     │
│                                                    │
│  Request đến → lấy 1 token                         │
│  Nếu có token: ✅ Allow                            │
│  Nếu trống:    ❌ Reject (429 Too Many Requests)   │
│                                                    │
│  Bucket tự refill theo thời gian                   │
└────────────────────────────────────────────────────┘

Ưu điểm:
- Cho phép burst ngắn (dùng hết tokens tích lũy)
- Smooth trong dài hạn

2.3 スライディング ウィンドウ アルゴリズム

Window: 1 minute, limit: 100 requests

Timeline: |──────────────────────────────────|
          0s        20s        40s       60s

Requests: |||||||||||||||||||||||            |
          (100 requests trong 20s đầu)

Request thứ 101 tại 21s:
- Fixed window: ALLOW (window reset về 0)  ← Có thể bị abuse
- Sliding window: REJECT (70 requests trong 60s trước) ← Chính xác hơn

2.4 API ゲートウェイ (Kong) でのレート制限

# Kong Rate Limiting plugin
apiVersion: configuration.konghq.com/v1
kind: KongPlugin
metadata:
  name: rate-limit-per-api-key
plugin: rate-limiting
config:
  minute: 1000          # 1000 req/min
  hour: 50000           # 50000 req/hour
  policy: redis          # Lưu counter trong Redis (cluster-aware)
  redis_host: redis.platform
  redis_port: 6379
  hide_client_headers: false  # Expose X-RateLimit-* headers to client
# Gắn plugin vào một ingress route
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  annotations:
    konghq.com/plugins: rate-limit-per-api-key

2.5 サービスレベルでのレート制限 (Resilience4j)

RateLimiterConfig config = RateLimiterConfig.custom()
    .limitForPeriod(100)                    // 100 permits per period
    .limitRefreshPeriod(Duration.ofSeconds(1))  // Refresh mỗi 1s
    .timeoutDuration(Duration.ofMillis(0))   // Reject ngay nếu không có permit
    .build();

RateLimiter rateLimiter = RateLimiterRegistry.of(config)
    .rateLimiter("order-creation");
@RateLimiter(name = "order-creation", fallbackMethod = "rateLimitFallback")
public Order createOrder(CreateOrderRequest request) {
    return orderService.create(request);
}

public Order rateLimitFallback(CreateOrderRequest req, RequestNotPermitted e) {
    throw new TooManyRequestsException("Rate limit exceeded, please retry later");
}

2.6 レート制限のための応答ヘッダー

HTTP/1.1 200 OK
X-RateLimit-Limit: 1000
X-RateLimit-Remaining: 743
X-RateLimit-Reset: 1711879200

# Khi bị reject:
HTTP/1.1 429 Too Many Requests
X-RateLimit-Limit: 1000
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 1711879200
Retry-After: 47
Content-Type: application/json

{"error": "rate_limit_exceeded", "retry_after": 47}

3. ヘルスチェックパターン

3.1 Kubernetes の 3 種類のプローブ

spec:
  containers:
    - name: order-service
      # Startup Probe: Service đã khởi động xong chưa?
      # K8s chờ startupProbe pass trước khi dùng liveness/readiness
      startupProbe:
        httpGet:
          path: /actuator/health/liveness
          port: 8080
        failureThreshold: 30    # 30 × 2s = 60s để start
        periodSeconds: 2

      # Liveness Probe: Service có đang chạy không?
      # Fail → K8s restart container
      livenessProbe:
        httpGet:
          path: /actuator/health/liveness
          port: 8080
        initialDelaySeconds: 0  # Sau khi startupProbe pass
        periodSeconds: 10
        failureThreshold: 3     # Restart sau 3 lần fail liên tiếp
        timeoutSeconds: 3

      # Readiness Probe: Service có sẵn sàng nhận traffic không?
      # Fail → K8s remove khỏi Service endpoints (không gửi traffic)
      readinessProbe:
        httpGet:
          path: /actuator/health/readiness
          port: 8080
        initialDelaySeconds: 0
        periodSeconds: 5
        failureThreshold: 3
        successThreshold: 1
        timeoutSeconds: 3

3.2 稼働性と準備性 — 重要な違い

LIVENESS: "Tôi có đang sống không?"
→ Fail → RESTART container
→ Check: không bị deadlock, heap không bị full, process không bị hang
→ Ví dụ fail: OutOfMemoryError, deadlock trong thread pool

READINESS: "Tôi có sẵn sàng nhận request không?"
→ Fail → REMOVE from load balancer (K8s Service endpoints)
→ Container không bị restart, vẫn chạy
→ Check: DB connection có sẵn, required caches đã warm, dependencies sẵn sàng
→ Ví dụ fail: DB connection pool full, warming up cache, throttled by upstream

3.3 Spring Boot アクチュエータの状態

// Tự động check:
// - Database connectivity
// - Disk space
// - Redis connectivity
// - Kafka connectivity (nếu dùng)

// Custom health indicator
@Component
public class OrderProcessorHealthIndicator implements HealthIndicator {

    private final OrderQueue queue;

    @Override
    public Health health() {
        int queueSize = queue.size();
        if (queueSize > 10000) {
            return Health.down()
                .withDetail("queue_size", queueSize)
                .withDetail("reason", "Queue backlog too large")
                .build();
        }

        return Health.up()
            .withDetail("queue_size", queueSize)
            .build();
    }
}
# application.yml — cấu hình health groups
management:
  endpoint:
    health:
      show-details: always
      group:
        liveness:
          include: livenessState
          # Chỉ check internal state, không check external dependencies
        readiness:
          include: readinessState, db, redis, kafka
          # Check đủ dependencies trước khi nhận traffic
  health:
    db:
      enabled: true
    redis:
      enabled: true

応答 /actuator/health/readiness:

{
  "status": "UP",
  "components": {
    "db": { "status": "UP", "details": { "database": "PostgreSQL", "result": 1 } },
    "redis": { "status": "UP" },
    "kafka": { "status": "UP" },
    "readinessState": { "status": "UP" }
  }
}

3.4 正常なシャットダウン

# application.yml
server:
  shutdown: graceful  # Chờ requests đang xử lý hoàn thành

spring:
  lifecycle:
    timeout-per-shutdown-phase: 30s  # Tối đa 30s chờ

# Kubernetes terminationGracePeriodSeconds phải > timeout-per-shutdown-phase
spec:
  terminationGracePeriodSeconds: 60
@Bean
public GracefulShutdown gracefulShutdown() {
    return new GracefulShutdown();
}

// Sequence khi K8s gửi SIGTERM:
// 1. Readiness probe → FAIL (K8s ngừng gửi traffic)
// 2. Spring starts graceful shutdown
// 3. Hoàn thành requests đang xử lý
// 4. Close connections (DB, Kafka, Redis)
// 5. Application exit(0)

4. すべてのパターンを結合する

本番環境に対応したサービスは、すべてを組み合わせたものです。

@Service
public class PaymentServiceClient {

    @Retry(name = "payment")
    @CircuitBreaker(name = "payment", fallbackMethod = "fallback")
    @Bulkhead(name = "payment", type = Type.SEMAPHORE)
    @RateLimiter(name = "payment-outbound")
    public PaymentResponse charge(ChargeRequest request) {
        return paymentClient.post("/charge", request, PaymentResponse.class);
    }

    public PaymentResponse fallback(ChargeRequest req, Throwable t) {
        if (t instanceof BulkheadFullException) {
            return PaymentResponse.queued(req.getOrderId());
        }
        if (t instanceof CallNotPermittedException) {
            return PaymentResponse.pending(req.getOrderId());
        }
        return PaymentResponse.failed(req.getOrderId(), t.getMessage());
    }
}

実行順序 (重要!)

Request đến
    ↓
RateLimiter → Có đủ permit không?
    ↓
Bulkhead → Có slot không?
    ↓
CircuitBreaker → Circuit có CLOSED không?
    ↓
TimeLimiter → Đặt timeout
    ↓
Retry → Thực thi và retry nếu fail
    ↓
Actual call đến Payment Service

アノテーションでラップする場合、順序は逆になります: Retry → CircuitBreaker → Bulkhead → RateLimiter。


5. 一般的なモニタリング

# Bulkhead đang gần đầy
resilience4j_bulkhead_available_concurrent_calls{name="payment-service"} < 3

# Rate limiter đang throttle nhiều requests
rate(resilience4j_ratelimiter_waiting_threads{name="payment-outbound"}[5m]) > 0

# Restart count cao trong 1 giờ
increase(kube_pod_container_status_restarts_total{
  namespace="services-prod"
}[1h]) > 3

概要

パターンメカニズムから守る
隔壁スレッド プール/セマフォの分離資源枯渇クロス
トークンバケットトークンの蓄積バーストトラフィック
スライディングウィンドウローリングカウント継続的な高トラフィック
活性プローブ内部状態のチェックスタック/デッドロックプロセス
レディネスプローブ外部依存関係のチェック準備ができていないときのトラフィック
正常なシャットダウン飛行中のリクエストを排出するデプロイ時のリクエストの損失

次の記事: カオス エンジニアリング — システムの信頼性の検証