Introduction
In microservices, when a service slows down or goes down, it can drag down the entire system (cascading failure). SmallRye Fault Tolerance (MicroProfile Fault Tolerance) provides patterns for protection: Retry, Timeout, Circuit Breaker, Bulkhead, and Fallback.
Dependency
<dependency>
<groupId>io.quarkus</groupId>
<artifactId>quarkus-smallrye-fault-tolerance</artifactId>
</dependency>
@Retry — Automatically retry
import org.eclipse.microprofile.faulttolerance.Retry;
@ApplicationScoped
public class OrderService {
@Inject @RestClient
ProductServiceClient productClient;
@Retry(maxRetries = 3,
delay = 500, // 500ms giữa các lần retry
jitter = 200, // ±200ms random delay
retryOn = {IOException.class,
WebApplicationException.class},
abortOn = {ResourceNotFoundException.class})
public ProductInfo getProduct(Long productId) {
return productClient.getById(productId);
}
}
Exponential Backoff
@Retry(maxRetries = 4,
delay = 1000,
maxDuration = 30000)
@ExponentialBackoff(factor = 2, maxDelay = 10000)
// Retry: 1s → 2s → 4s → 8s (max 10s)
public ProductInfo getProductWithBackoff(Long id) {
return productClient.getById(id);
}
@Timeout — Time limit
import org.eclipse.microprofile.faulttolerance.Timeout;
@Timeout(value = 5, unit = ChronoUnit.SECONDS)
public ProductInfo getProduct(Long id) {
// Nếu > 5s → throw TimeoutException
return productClient.getById(id);
}
Combine Retry + Timeout
@Retry(maxRetries = 3, delay = 1000)
@Timeout(5000) // Mỗi lần try tối đa 5s
public ProductInfo getProductReliable(Long id) {
return productClient.getById(id);
}
// Worst case: 3 retries × 5s timeout + 3 × 1s delay = 18s max
@CircuitBreaker — Circuit breaker
Circuit Breaker prevents cascading failures by "interrupting" calls to the failed service:
┌─────────┐ failures > threshold ┌──────┐
│ CLOSED │ ──────────────────────────> │ OPEN │
│(normal) │ │(fail)│
└─────────┘ └──┬───┘
^ │
│ ┌────────────┐ │
│ │ HALF-OPEN │ <─────────────┘
└────────│ (testing) │ after delay
success └────────────┘
import org.eclipse.microprofile.faulttolerance.CircuitBreaker;
@CircuitBreaker(
requestVolumeThreshold = 20, // Window: 20 requests
failureRatio = 0.5, // Mở khi 50% fail
delay = 10, // Đợi 10s trước khi thử lại
delayUnit = ChronoUnit.SECONDS,
successThreshold = 3) // Đóng lại sau 3 success
@Fallback(fallbackMethod = "getProductFallback")
public ProductInfo getProduct(Long id) {
return productClient.getById(id);
}
// Fallback khi circuit OPEN hoặc call fail
public ProductInfo getProductFallback(Long id) {
// Trả về cached data hoặc default
return cachedProducts.getOrDefault(id,
new ProductInfo(id, "Product Unavailable",
BigDecimal.ZERO, "VND"));
}
@Bulkhead — Limit concurrent calls
Prevent the service from being overloaded by too many concurrent requests:
import org.eclipse.microprofile.faulttolerance.Bulkhead;
@Bulkhead(value = 10, // Max 10 concurrent calls
waitingTaskQueue = 5) // Max 5 queued
@Timeout(5000)
public ProductInfo getProduct(Long id) {
return productClient.getById(id);
}
// Nếu > 15 (10 running + 5 queued) → BulkheadException
@Fallback — Replace value
import org.eclipse.microprofile.faulttolerance.Fallback;
// Method-level fallback
@Fallback(fallbackMethod = "getProductFallback")
@Retry(maxRetries = 2)
@Timeout(3000)
public ProductInfo getProduct(Long id) {
return productClient.getById(id);
}
private ProductInfo getProductFallback(Long id) {
Log.warnf("Fallback for product %d", id);
// Trả về từ local cache
return productCache.get(id);
}
// Handler class fallback
@Fallback(value = ProductFallbackHandler.class)
public ProductInfo getProduct2(Long id) {
return productClient.getById(id);
}
public class ProductFallbackHandler
implements FallbackHandler<ProductInfo> {
@Override
public ProductInfo handle(ExecutionContext context) {
// Log error, return default
return new ProductInfo(
0L, "Unavailable", BigDecimal.ZERO, "VND");
}
}
Execution order when combining Annotations
When using multiple annotations at the same time, execution order:
Request → Bulkhead → CircuitBreaker → Retry → Timeout → Method → Fallback
↑
(on any failure)
Ví dụ thực tế với getProduct():
┌─────────────────────────────────────────────────────────┐
│ ① Bulkhead: Có slot trống? │
│ ├─ YES → tiếp tục │
│ └─ NO → BulkheadException → ⑥ Fallback │
│ │
│ ② CircuitBreaker: CLOSED? │
│ ├─ CLOSED → tiếp tục │
│ ├─ HALF-OPEN → cho 1 request thử │
│ └─ OPEN → CircuitBreakerOpenException → ⑥ Fallback │
│ │
│ ③ Retry: lần thử thứ mấy? (max 3) │
│ ├─ Lần 1 → gọi method │
│ └─ Fail → đợi delay → retry lần 2, 3... │
│ │
│ ④ Timeout: method chạy < 5s? │
│ ├─ YES → trả kết quả │
│ └─ NO → TimeoutException → ③ Retry thử lại │
│ │
│ ⑤ Method: productClient.getById(id) │
│ │
│ ⑥ Fallback: khi tất cả retries fail │
│ → getProductFallback(id) │
└─────────────────────────────────────────────────────────┘
Combine it all — Real-world Pattern
@ApplicationScoped
public class ResilientProductClient {
@Inject @RestClient
ProductServiceClient productClient;
@Inject
ProductCache productCache;
@Retry(maxRetries = 3, delay = 500, jitter = 200,
retryOn = IOException.class)
@CircuitBreaker(requestVolumeThreshold = 20,
failureRatio = 0.5,
delay = 10, delayUnit = ChronoUnit.SECONDS)
@Bulkhead(value = 20, waitingTaskQueue = 10)
@Timeout(5000)
@Fallback(fallbackMethod = "getProductFallback")
public ProductInfo getProduct(Long id) {
ProductInfo product = productClient.getById(id);
// Cập nhật cache khi thành công
productCache.put(id, product);
return product;
}
public ProductInfo getProductFallback(Long id) {
Log.warnf("Using cached product for id: %d", id);
ProductInfo cached = productCache.get(id);
if (cached != null) return cached;
throw new ServiceUnavailableException(
"Product Service unavailable "
+ "and no cached data for product " + id);
}
}
@RateLimit — Request rate limit (Quarkus 3.x)
import io.smallrye.faulttolerance.api.RateLimit;
import java.time.temporal.ChronoUnit;
// Giới hạn 100 requests / 1 phút
@RateLimit(value = 100,
window = 1, windowUnit = ChronoUnit.MINUTES)
@Fallback(fallbackMethod = "rateLimitedFallback")
public ProductInfo getProduct(Long id) {
return productClient.getById(id);
}
public ProductInfo rateLimitedFallback(Long id) {
throw new WebApplicationException(
"Rate limit exceeded. Try again later.", 429);
}
Cascading Failure — Real-life example
Suppose Payment Service is slow (DB overload):
┌──────────┐ ┌──────────┐ ┌──────────┐
│ Client │────>│ Order │────>│ Payment │ ← DB chậm (30s response)
│ Browser │ │ Service │ │ Service │
└──────────┘ └──────────┘ └──────────┘
│ │
│ Thread pool │ Tất cả threads blocked
│ cạn kiệt! │ chờ Payment response
│ ▼
│ Order Service
│ KHÔNG THỂ xử lý
│ request mới
└────> 503 Service Unavailable
No Fault Tolerance: Order Service dies with Payment Service (cascading failure).
Has Fault Tolerance:
@ApplicationScoped
public class ResilientPaymentClient {
@Inject @RestClient
PaymentServiceClient paymentClient;
@Timeout(3000) // Không chờ quá 3s
@CircuitBreaker(
requestVolumeThreshold = 10,
failureRatio = 0.5,
delay = 30,
delayUnit = ChronoUnit.SECONDS)
@Bulkhead(value = 5) // Max 5 concurrent calls
@Fallback(fallbackMethod = "paymentFallback")
public PaymentResult processPayment(
PaymentRequest request) {
return paymentClient.process(request);
}
public PaymentResult paymentFallback(
PaymentRequest request) {
Log.warnf("Payment Service unavailable, "
+ "queuing payment for order %s",
request.orderId());
// Gửi vào Kafka queue để retry sau
paymentQueue.send(request);
return new PaymentResult(
request.orderId(),
"PENDING",
"Payment queued for processing");
}
}
Results:
- Timeout 3s → do not block thread for 30s
- Circuit Breaker → after 5/10 failures, break immediately → fail fast
- Bulkhead 5 → only 5 threads call Payment, the rest serve other requests
- Fallback → queue into Kafka, retry background → user is not blocked
Metrics Integration
SmallRye Fault Tolerance automatically exposes metrics when available quarkus-micrometer:
<dependency>
<groupId>io.quarkus</groupId>
<artifactId>quarkus-micrometer-registry-prometheus</artifactId>
</dependency>
Automated metrics are available at /q/metrics:
# Retry metrics
ft_retry_calls_total{method="getProduct",retried="true",retryResult="valueReturned"} 42
ft_retry_calls_total{method="getProduct",retried="true",retryResult="maxRetriesReached"} 3
ft_retry_retries_total{method="getProduct"} 86
# Circuit Breaker metrics
ft_circuitbreaker_state_total{method="getProduct",state="closed"} 95.0
ft_circuitbreaker_state_total{method="getProduct",state="open"} 3.0
ft_circuitbreaker_state_total{method="getProduct",state="halfOpen"} 2.0
ft_circuitbreaker_opened_total{method="getProduct"} 2
# Bulkhead metrics
ft_bulkhead_executionsRunning{method="getProduct"} 3
ft_bulkhead_executionsWaiting{method="getProduct"} 1
ft_bulkhead_runningDuration_seconds{method="getProduct",quantile="0.95"} 0.234
# Timeout metrics
ft_timeout_calls_total{method="getProduct",timedOut="true"} 5
ft_timeout_calls_total{method="getProduct",timedOut="false"} 195
Grafana Dashboard Query example
# Circuit Breaker open rate
rate(ft_circuitbreaker_opened_total[5m])
# Retry success rate
rate(ft_retry_calls_total{retryResult="valueReturned"}[5m])
/ rate(ft_retry_calls_total[5m])
# P95 bulkhead execution time
ft_bulkhead_runningDuration_seconds{quantile="0.95"}
Configuration via application.properties
# Override annotations qua config
com.xdev.ecommerce.order.ResilientProductClient/getProduct/Retry/maxRetries=5
com.xdev.ecommerce.order.ResilientProductClient/getProduct/Timeout/value=3000
com.xdev.ecommerce.order.ResilientProductClient/getProduct/CircuitBreaker/delay=15000
# Global defaults
Retry/maxRetries=3
Timeout/value=5000
CircuitBreaker/failureRatio=0.5
Testing Fault Tolerance
@QuarkusTest
class ResilientProductClientTest {
@InjectMock
@RestClient
ProductServiceClient productClient;
@Inject
ResilientProductClient resilientClient;
@Test
void testRetryOnFailure() {
// First two calls fail, third succeeds
when(productClient.getById(1L))
.thenThrow(new IOException("timeout"))
.thenThrow(new IOException("timeout"))
.thenReturn(new ProductInfo(1L, "Test",
BigDecimal.TEN, "VND"));
ProductInfo result = resilientClient.getProduct(1L);
assertEquals("Test", result.name());
// Verify 3 calls made
verify(productClient, times(3)).getById(1L);
}
@Test
void testFallbackWhenAllRetriesFail() {
when(productClient.getById(1L))
.thenThrow(new IOException("down"));
// Should return fallback (cached or throw)
// ...
}
}
Exercises
- Add
@Retry+@Timeoutfor Product Service REST Client calls - Implement
@CircuitBreakerwith fallback returns cached data - Add
@BulkheadLimit concurrent calls to Product Service - Create
ResilientProductClientwrapper with all patterns - Create
ResilientPaymentClientwith fallback queue into Kafka - Override config via
application.properties - Add Micrometer metrics — verify metrics at
/q/metrics - Test: mock Product Service down → verify retry → circuit open → fallback
- (Advanced) Create Grafana dashboard for Fault Tolerance metrics
Summary
@Retry— automatic retry, exponential backoff support@Timeout— limits processing time, prevents thread from being blocked@CircuitBreaker— circuit break when service fails too much (CLOSED → OPEN → HALF-OPEN)@Bulkhead— limit concurrent calls, avoid resource exhaustion@Fallback— returns cached/default data when all fail, or queue retry@RateLimit— limit request rate (requests/window)- Execution order: Bulkhead → CircuitBreaker → Retry → Timeout → Method → Fallback
- Cascading Failure Prevention — combines Timeout + CircuitBreaker + Bulkhead
- Metrics — automatically exposed via Micrometer/Prometheus
- Config override passed
application.properties— no need to change code
Next article: OpenTelemetry — Distributed Tracing & Metrics.