Why this matters
- A retry storm during a 30-second fraud outage took down ShardPay's healthy ledger shards by multiplying traffic 10×.
- Payment APIs without idempotency keys double-charge when clients retry POST on timeout.
- Not every error is retriable — retrying
400 InsufficientFundshammering the ledger wastes capacity and logs noise. - Interviewers ask for a full retry policy: which status codes, backoff shape, max attempts, and idempotency.
Retriable vs non-retriable errors
ShardPay classifies errors before retrying. Retriable: HTTP 408, 429, 502, 503, 504; gRPC UNAVAILABLE, DEADLINE_EXCEEDED; JDBC transient connection errors. Non-retriable: 400, 401, 403, 404, 409, 422; gRPC INVALID_ARGUMENT, FAILED_PRECONDITION; business errors like insufficient funds.
Ambiguous cases matter: a client timeout after the server committed must retry with the same idempotency key, not as a new transfer. ShardPay documents POST /transfers as safe to retry only when Idempotency-Key header is present.
Walkthrough: client timeout after commit
Mobile app POSTs a transfer; gateway times out at 200ms while ledger already committed. User taps "retry." Without idempotency the second POST creates a duplicate transfer. With Idempotency-Key: uuid-v4 the second request hits ShardPay's dedup store and returns the original transferId and 201 body — no double debit.
Key points
- Exponential backoff — delay doubles each attempt (100ms, 200ms, 400ms) to give failing dependencies time to recover. ShardPay's Java
RetryTemplatestarts at 100ms base for ledger RPCs with max 3 attempts inside the parent deadline. - Full jitter — each delay is
random(0, calculated_backoff)so synchronized clients desynchronize. After ShardPay added full jitter to fraud client retries, recovery time from vendor blips dropped because request peaks flattened. - Retry budget — cap total retry wall time (e.g. 150ms of retries inside 200ms SLA) and max attempt count. ShardPay aborts retries when
Deadline.afterNow()would exceed gateway budget and returns pending state instead. - Idempotent retry — executing the same logical operation twice produces one durable effect. ShardPay requires idempotency keys on all mutating payment APIs; retries without keys get
428 Precondition Required. - Retry amplification — each client retry × each middleware retry × each downstream retry multiplies load. ShardPay disables automatic HTTP client retries on the gateway→ledger path; only the orchestrator retries with explicit policy.
Backoff configuration in practice
ShardPay uses Resilience4j Retry with exponential backoff and random jitter for outbound gRPC. Idempotency store lookups happen before the first attempt so retries never re-execute ledger logic without the guard.
For Kafka consumers, ShardPay relies on broker redelivery (at-least-once) plus idempotent handlers rather than application-level retry loops inside the poll thread — infinite consumer retries block partition progress.
Walkthrough: regional failover retry
Fraud vendor us-east returns 503. ShardPay's client retries twice with jitter (120ms, 240ms) then fails over to us-west read replica with a fresh attempt budget. Total elapsed 380ms — outside sync path SLA so orchestrator marks transfer pending_fraud and async worker completes scoring. Retries stayed bounded; no storm hit the failed region after failover.
// ShardPay ledger client — bounded retry with jitter (Java 17)
public LedgerResult debitWithRetry(TransferRequest req, String idempotencyKey, Deadline deadline) {
int attempt = 0;
long backoffMs = 50;
while (true) {
attempt++;
try {
return ledgerStub
.withDeadline(deadline)
.debit(req.toProto(), metadataWithIdempotency(idempotencyKey));
} catch (StatusRuntimeException e) {
if (!isRetriable(e) || attempt >= 3 || deadline.timeRemaining().toMillis() < backoffMs) {
throw e;
}
long sleep = ThreadLocalRandom.current().nextLong(0, backoffMs);
backoffMs = Math.min(backoffMs * 2, 500);
try {
Thread.sleep(sleep);
} catch (InterruptedException ie) {
Thread.currentThread().interrupt();
throw e;
}
}
}
}
private static boolean isRetriable(StatusRuntimeException e) {
return e.getStatus().getCode() == Status.Code.UNAVAILABLE
|| e.getStatus().getCode() == Status.Code.DEADLINE_EXCEEDED;
}
When not to retry
Do not retry on business rejection — insufficient funds, invalid account, fraud block. Do not retry unbounded on the consumer side of a queue. Do not nest retries: if service A retries service B three times and B retries database three times, worst case is 9× load.
ShardPay's rule: at most one retry layer owns policy per hop; inner layers fail fast to outer policy.
Quick recall
Everything you need if you only revisit this box.
- Backoff plus full jitter prevents synchronized retry storms.
- Only retry idempotent operations — idempotency keys on every mutating payment API.
- Cap retries with a budget inside the parent deadline; one retry owner per hop.
Test yourself
Answer these before moving on — recall is what makes it stick.