Why this matters
- When the payment service takes 30 seconds to respond, every checkout thread blocks — thread pools exhaust and the catalog API goes down too, even though catalog itself is healthy.
- Circuit breakers trip after repeated failures, returning a fallback immediately instead of waiting for timeouts on every request.
- Retries with exponential backoff recover from transient blips; rate limiters protect your own services from being overwhelmed by retries from callers.
ClosedNormal traffic
OpenFail fast
Half-openProbe request
Resilience4j patterns in Spring Boot
- Circuit Breaker — opens after failure threshold; short-circuits calls while upstream recovers.
- Retry — re-attempts failed calls with configurable backoff and max attempts.
- Rate Limiter — caps requests per time window to protect downstream services.
- Bulkhead — isolates thread pools so one slow dependency cannot starve others.
- Time Limiter — enforces a maximum duration on any call, including async ones.
Dependencies and configuration
<dependency>
<groupId>org.springframework.cloud</groupId>
<artifactId>spring-cloud-starter-circuitbreaker-resilience4j</artifactId>
</dependency>
resilience4j:
circuitbreaker:
instances:
paymentService:
sliding-window-size: 10
failure-rate-threshold: 50
wait-duration-in-open-state: 30s
permitted-number-of-calls-in-half-open-state: 3
retry:
instances:
paymentService:
max-attempts: 3
wait-duration: 500ms
exponential-backoff-multiplier: 2
timelimiter:
instances:
paymentService:
timeout-duration: 5s
Circuit breaker with fallback
@Service
public class CheckoutService {
private final CircuitBreakerFactory circuitBreakerFactory;
private final PaymentClient paymentClient;
private final OrderRepository orderRepository;
public CheckoutResult checkout(CheckoutRequest request) {
CircuitBreaker breaker = circuitBreakerFactory.create("paymentService");
PaymentResult payment = breaker.run(
() -> paymentClient.charge(request.toChargeRequest()),
throwable -> PaymentResult.pending(request.orderId())
);
Order order = orderRepository.save(buildOrder(request, payment));
return CheckoutResult.from(order, payment);
}
}
Declarative resilience with annotations
@Service
public class InventoryService {
private final InventoryClient inventoryClient;
@CircuitBreaker(name = "inventoryService", fallbackMethod = "stockUnavailableFallback")
@Retry(name = "inventoryService")
@TimeLimiter(name = "inventoryService")
public CompletableFuture<StockLevel> checkStockAsync(String isbn) {
return CompletableFuture.supplyAsync(() -> inventoryClient.checkStock(isbn));
}
private CompletableFuture<StockLevel> stockUnavailableFallback(String isbn, Throwable t) {
log.warn("Inventory check failed for {}, returning unknown stock", isbn, t);
return CompletableFuture.completedFuture(StockLevel.unknown(isbn));
}
}
Bulkhead isolation
@Configuration
public class ResilienceConfig {
@Bean
public ThreadPoolBulkheadRegistry bulkheadRegistry() {
ThreadPoolBulkheadConfig config = ThreadPoolBulkheadConfig.custom()
.maxThreadPoolSize(10)
.coreThreadPoolSize(5)
.queueCapacity(25)
.build();
ThreadPoolBulkheadRegistry registry = ThreadPoolBulkheadRegistry.of(config);
registry.bulkhead("paymentService", config);
registry.bulkhead("inventoryService", config);
return registry;
}
}
@Bulkhead(name = "paymentService", type = Bulkhead.Type.THREADPOOL)
public PaymentResult processPayment(ChargeRequest request) {
return paymentClient.charge(request);
}
Quick recall
Everything you need if you only revisit this box.
- Resilience4j integrates with Spring Cloud Circuit Breaker — configure instances in
application.yml. - Circuit breakers fail fast when upstream is unhealthy; combine with fallback methods for degraded responses.
- Retry transient failures with exponential backoff; never retry non-idempotent operations blindly.
- Bulkheads isolate thread pools per dependency so one outage does not exhaust the shared pool.
- Time limiters cap how long any single call can block, even inside a retry loop.
Test yourself
Answer these before moving on — recall is what makes it stick.