PrepZone Logo
PrepZone

Circuit Breaker and Bulkhead

Fail fast and isolate thread pools so one slow dependency cannot exhaust ShardPay workers.

Why this matters

  • Without bulkheads, a hung fraud vendor blocked all ShardPay transfer threads — including transfers that could skip fraud in low-risk corridors.
  • Circuit breakers convert slow failures into fast failures, preserving capacity for healthy paths and graceful degradation.
  • Resilience4j and Istio implement these patterns in Java/Kubernetes stacks ShardPay runs in production.
  • Interviewers ask how you degrade when fraud is down without losing committed ledger state.
ClosedNormal traffic
OpenFail fast
Half-openProbe request
Open circuit fails fast; half-open probes recovery.

Circuit breaker states

A circuit breaker wraps outbound calls to a dependency. Closed: calls pass through; failures increment a sliding-window counter. Open: calls fail immediately without hitting the dependency — ShardPay returns pending_fraud from cache policy. Half-open: one probe request tests recovery; success closes the circuit, failure reopens.

ShardPay opens the fraud circuit when error rate exceeds 50% over a 30-second window with minimum 20 calls — avoiding flapping on single errors during low traffic.

Walkthrough: fraud vendor outage

At 14:00 fraud vendor begins returning 503 for all requests. By 14:00:20 ShardPay's circuit opens. Transfers from 14:00:20 to 14:05 complete ledger commit but skip sync fraud; async workers queue scoring for when half-open probes succeed. Ledger p99 stays under 200ms; manual review queue depth rises but the platform stays up. At 14:05 a probe succeeds; circuit closes and drains the backlog with rate limiting.

Key points

  • Closed state — normal operation; failures tracked in a rolling window. ShardPay logs circuit metrics to Grafana alongside fraud vendor status page correlation for post-incident review.
  • Open state — fail fast without calling the dependency; protects the vendor from retry amplification and protects ShardPay threads from blocking. Open fraud circuit routes transfers to pending_fraud with customer-visible "verification in progress" copy.
  • Half-open state — allows a single probe (or small probe percentage in Istio) to test recovery before restoring full traffic. ShardPay requires three consecutive probe successes before closing to avoid flapping on intermittent 200s.
  • Bulkhead — isolates resource pools per dependency so one slow callee cannot exhaust all threads. ShardPay's transfer executor uses 80 threads for ledger and a dedicated 20-thread pool for fraud RPCs — fraud saturation never blocks ledger-only maintenance jobs.
  • Graceful degradation — define product behavior when circuit is open: sync path vs async queue vs hard fail. ShardPay never rolls back committed ledger entries when fraud opens; compliance accepts async scoring within 15 minutes for low-value corridors.

Bulkheads and thread isolation

Bulkheads apply to thread pools, connection pools, and semaphores. ShardPay limits concurrent open TCP connections to the fraud vendor at 50 regardless of total gateway traffic. Semaphore-based bulkheads in Resilience4j reject excess calls immediately with BulkheadFullException mapped to 503 with Retry-After.

Kubernetes resource limits are a coarse bulkhead — ShardPay pairs pod CPU limits with application-level pools so one bad dependency cannot fill the pool and starve health checks.

Walkthrough: ledger stays warm during fraud storm

During fraud outage, 5000 concurrent transfers arrive. Fraud bulkhead (20 threads) saturates instantly; excess fraud tasks reject in under 1ms. Ledger bulkhead (80 threads) continues processing debits. Without bulkheads all 5000 threads would block on fraud socket read timeout (30s pre-fix) and ledger would stall — the incident that motivated bulkhead adoption.

Java
// ShardPay fraud client — circuit breaker + bulkhead (Resilience4j, Java 17)
CircuitBreaker fraudBreaker = CircuitBreaker.of("fraud-vendor",
    CircuitBreakerConfig.custom()
        .failureRateThreshold(50)
        .waitDurationInOpenState(Duration.ofSeconds(30))
        .slidingWindowSize(20)
        .build());

Bulkhead fraudBulkhead = Bulkhead.of("fraud-pool",
    BulkheadConfig.custom().maxConcurrentCalls(20).build());

public FraudDecision score(TransferRequest req) {
    Supplier<FraudDecision> decorated = Decorators.ofSupplier(() -> fraudGrpcClient.score(req))
        .withBulkhead(fraudBulkhead)
        .withCircuitBreaker(fraudBreaker)
        .withFallback(e -> FraudDecision.pendingManualReview(req.transferId()))
        .decorate();

    return decorated.get();
}

Service mesh vs library breakers

ShardPay uses Resilience4j in the transfer orchestrator for fine-grained fallbacks and Istio destination rules for cluster-wide outlier detection on fraud egress. Library breakers know business fallbacks (pending_fraud); mesh breakers eject unhealthy endpoints from load balancing across pods.

Both layers complement each other — neither alone covers pod-level vs vendor-level failures.

Quick recall

Everything you need if you only revisit this box.

  1. Circuit breaker fails fast when a dependency is unhealthy — define degradation before opening.
  2. Bulkhead isolates thread and connection pools so one bad dependency cannot stall the ledger.
  3. Half-open probes recovery with bounded traffic before full restore.

Test yourself

Answer these before moving on — recall is what makes it stick.