PrepZone Logo
PrepZone

Retries, Backoff, and Jitter

Retry storms take down healthy services — exponential backoff with full jitter.

Why this matters

  • A retry storm during a 30-second fraud outage took down ShardPay's healthy ledger shards by multiplying traffic 10×.
  • Payment APIs without idempotency keys double-charge when clients retry POST on timeout.
  • Not every error is retriable — retrying 400 InsufficientFunds hammering the ledger wastes capacity and logs noise.
  • Interviewers ask for a full retry policy: which status codes, backoff shape, max attempts, and idempotency.
Retry delays
Attempt 10ms
Attempt 2~100ms + jitter
Attempt 3~200ms + jitter
Attempt 4~400ms + jitter
Spread retries to prevent thundering herd on recovery.

Retriable vs non-retriable errors

ShardPay classifies errors before retrying. Retriable: HTTP 408, 429, 502, 503, 504; gRPC UNAVAILABLE, DEADLINE_EXCEEDED; JDBC transient connection errors. Non-retriable: 400, 401, 403, 404, 409, 422; gRPC INVALID_ARGUMENT, FAILED_PRECONDITION; business errors like insufficient funds.

Ambiguous cases matter: a client timeout after the server committed must retry with the same idempotency key, not as a new transfer. ShardPay documents POST /transfers as safe to retry only when Idempotency-Key header is present.

Walkthrough: client timeout after commit

Mobile app POSTs a transfer; gateway times out at 200ms while ledger already committed. User taps "retry." Without idempotency the second POST creates a duplicate transfer. With Idempotency-Key: uuid-v4 the second request hits ShardPay's dedup store and returns the original transferId and 201 body — no double debit.

Key points

  • Exponential backoff — delay doubles each attempt (100ms, 200ms, 400ms) to give failing dependencies time to recover. ShardPay's Java RetryTemplate starts at 100ms base for ledger RPCs with max 3 attempts inside the parent deadline.
  • Full jitter — each delay is random(0, calculated_backoff) so synchronized clients desynchronize. After ShardPay added full jitter to fraud client retries, recovery time from vendor blips dropped because request peaks flattened.
  • Retry budget — cap total retry wall time (e.g. 150ms of retries inside 200ms SLA) and max attempt count. ShardPay aborts retries when Deadline.afterNow() would exceed gateway budget and returns pending state instead.
  • Idempotent retry — executing the same logical operation twice produces one durable effect. ShardPay requires idempotency keys on all mutating payment APIs; retries without keys get 428 Precondition Required.
  • Retry amplification — each client retry × each middleware retry × each downstream retry multiplies load. ShardPay disables automatic HTTP client retries on the gateway→ledger path; only the orchestrator retries with explicit policy.

Backoff configuration in practice

ShardPay uses Resilience4j Retry with exponential backoff and random jitter for outbound gRPC. Idempotency store lookups happen before the first attempt so retries never re-execute ledger logic without the guard.

For Kafka consumers, ShardPay relies on broker redelivery (at-least-once) plus idempotent handlers rather than application-level retry loops inside the poll thread — infinite consumer retries block partition progress.

Walkthrough: regional failover retry

Fraud vendor us-east returns 503. ShardPay's client retries twice with jitter (120ms, 240ms) then fails over to us-west read replica with a fresh attempt budget. Total elapsed 380ms — outside sync path SLA so orchestrator marks transfer pending_fraud and async worker completes scoring. Retries stayed bounded; no storm hit the failed region after failover.

Java
// ShardPay ledger client — bounded retry with jitter (Java 17)
public LedgerResult debitWithRetry(TransferRequest req, String idempotencyKey, Deadline deadline) {
    int attempt = 0;
    long backoffMs = 50;
    while (true) {
        attempt++;
        try {
            return ledgerStub
                .withDeadline(deadline)
                .debit(req.toProto(), metadataWithIdempotency(idempotencyKey));
        } catch (StatusRuntimeException e) {
            if (!isRetriable(e) || attempt >= 3 || deadline.timeRemaining().toMillis() < backoffMs) {
                throw e;
            }
            long sleep = ThreadLocalRandom.current().nextLong(0, backoffMs);
            backoffMs = Math.min(backoffMs * 2, 500);
            try {
                Thread.sleep(sleep);
            } catch (InterruptedException ie) {
                Thread.currentThread().interrupt();
                throw e;
            }
        }
    }
}

private static boolean isRetriable(StatusRuntimeException e) {
    return e.getStatus().getCode() == Status.Code.UNAVAILABLE
        || e.getStatus().getCode() == Status.Code.DEADLINE_EXCEEDED;
}

When not to retry

Do not retry on business rejection — insufficient funds, invalid account, fraud block. Do not retry unbounded on the consumer side of a queue. Do not nest retries: if service A retries service B three times and B retries database three times, worst case is 9× load.

ShardPay's rule: at most one retry layer owns policy per hop; inner layers fail fast to outer policy.

Quick recall

Everything you need if you only revisit this box.

  1. Backoff plus full jitter prevents synchronized retry storms.
  2. Only retry idempotent operations — idempotency keys on every mutating payment API.
  3. Cap retries with a budget inside the parent deadline; one retry owner per hop.

Test yourself

Answer these before moving on — recall is what makes it stick.