PrepZone Logo
PrepZone

Timeout Hierarchies and Latency Budgets

Per-hop timeouts must sum under the end-to-end SLA — ShardPay's 200ms transfer budget.

Why this matters

  • Missing timeouts caused thread pool exhaustion at ShardPay when the fraud vendor hung for 90 seconds during a regional network glitch.
  • End-to-end latency budgets force explicit trade-offs instead of every dependency getting "30 seconds just in case."
  • Interviewers hand you a 200ms SLA and ask you to slice budget across five microservices.
  • Propagating deadlines prevents downstream work from running after the client already disconnected.
Gateway20ms
Auth15ms
Ledger80ms
Fraud40ms
Response200ms total
Each hop consumes part of the SLA. Sum of hop timeouts must stay under the total budget.

End-to-end budget anatomy

ShardPay's synchronous transfer path targets p99 ≤ 200ms from API gateway to response. That 200ms is not per service — it is the entire user-visible window. The gateway reserves 20ms for TLS and routing, 30ms for auth token validation, 100ms for ledger debit/credit on the shard, 50ms for fraud scoring, and 20ms slack for serialization and network jitter.

If fraud consistently consumes 50ms, ledger cannot borrow its unused slack without violating p99. Budgets are negotiated at design time and enforced in code, not renegotiated per request.

Walkthrough: budget exhausted mid-chain

A transfer hits the gateway at T+0ms. Auth completes at T+25ms (5ms over, but within hard cap). Ledger starts with 175ms remaining on the propagated deadline. Shard 7 is slow; ledger returns at T+120ms with 80ms left. Fraud starts with min(local_cap=50ms, remaining=80ms) = 50ms. Fraud's vendor times out at T+170ms. ShardPay returns 202 Accepted with status=pending_fraud instead of waiting another 30ms and breaching the 200ms SLA.

Key points

  • End-to-end budget — the total SLA clock started when the client sent the request. ShardPay's gateway starts a 200ms Deadline on every POST /transfers and refuses to enqueue work when fewer than 10ms remain — fail fast rather than guaranteed timeout.
  • Per-hop timeout — each dependency gets a hard cap smaller than the remaining budget. Ledger's 100ms cap prevents a slow shard from consuming fraud's allocation; ops tune caps from distributed trace percentiles quarterly.
  • Context deadline propagation — gRPC and W3C traceparent baggage carry deadline_ms so each hop knows remaining time. ShardPay's Java services read Deadline.afterNow() and pass shortened stubs to downstream clients automatically.
  • Graceful partial response — when budget is exhausted before the full pipeline completes, return a definitive async state (202, transferId, poll URL) instead of 504 Gateway Timeout. ShardPay reduced duplicate client retries by 40% after replacing bare 504s with structured pending responses.
  • Budget vs timeout — budget is the product SLA; timeout is the enforcement mechanism per hop. A hop may timeout at 50ms even if 80ms remains globally, to preserve slack for mandatory downstream steps.

Propagating deadlines in Java

ShardPay uses gRPC server interceptors to read incoming deadlines and client interceptors to attach shortened deadlines on outbound calls. A service must never ignore the propagated deadline and apply its own static 30s HTTP timeout — that defeats the budget.

For JDBC calls to colocated Postgres on the ledger shard, ShardPay sets statement_timeout from remaining budget minus 5ms guard. ORM-level defaults of "no timeout" are banned in production configs.

Walkthrough: cascading cancel

The mobile client disconnects at T+80ms (user closed app). The gateway cancels the request context. ShardPay's ledger service registers Context.current().addListener(...) and aborts the in-flight SQL when the listener fires — avoiding wasted shard capacity on a client that will never read the response.

Java
// ShardPay transfer orchestrator — budget-aware downstream calls (Java 17)
public TransferResponse executeTransfer(TransferRequest req, Deadline parentDeadline) {
    Deadline authDeadline = parentDeadline.withTimeout(Duration.ofMillis(30));
    AuthResult auth = authClient.verify(req.token(), authDeadline);

    Duration ledgerBudget = parentDeadline.timeRemaining().minusMillis(20);
    if (ledgerBudget.isNegative() || ledgerBudget.toMillis() < 10) {
        return TransferResponse.pending(req.clientRequestId(), "budget_exhausted");
    }
    LedgerResult ledger = ledgerClient.debitAndCredit(req, parentDeadline.withTimeout(ledgerBudget));

    Duration fraudBudget = Duration.ofMillis(Math.min(50, parentDeadline.timeRemaining().toMillis()));
    try {
        fraudClient.score(req, parentDeadline.withTimeout(fraudBudget));
        return TransferResponse.completed(ledger.transferId());
    } catch (DeadlineExceededException e) {
        return TransferResponse.pending(ledger.transferId(), "pending_fraud");
    }
}

Measuring and tuning budgets

ShardPay exports a "budget waterfall" metric per trace: milliseconds spent in auth, ledger, fraud, and unaccounted. Weekly SLO reviews compare allocated caps to p95 actuals. If ledger p95 is 60ms but cap is 100ms, slack moves to fraud or back to the user-facing headroom.

Cold-start and GC pauses eat budget too — ShardPay subtracts 10ms slack at the gateway during pod rollout windows. Budget math is living configuration, not a one-time architecture slide.

Quick recall

Everything you need if you only revisit this box.

  1. Child timeouts must sum inside the end-to-end SLA — not inherit generic 30s defaults.
  2. Propagate deadlines on every RPC; cancel downstream work when the client disconnects.
  3. Return structured pending responses when budget runs out after partial commit — not bare 504s.

Test yourself

Answer these before moving on — recall is what makes it stick.