PrepZone Logo
PrepZone

The Distributed Systems Interview Playbook

Framework for trade-off questions, failure scenarios, and architect-level system reasoning.

Why this matters

  • This article consolidates the Distributed Systems track into an interview-ready framework.
  • Interviewers at senior/staff level rarely ask trivia — they present ambiguous problems and evaluate structured thinking.
  • ShardPay examples from modules 1–12 provide concrete anchors for every abstract trade-off.
  • Revisit this article after completing all modules for spaced recall before interview loops.
ClientMobile / web
API gateway
Ledger AShard 1
Ledger BShard 2
Ledger CShard 3
Independent nodes communicate over a network. Any node can fail without warning.

The four-step framework

Every distributed systems design answer at ShardPay interview bar follows four steps. Skip any step and the answer feels junior.

Step 1: Clarify requirements

Before drawing boxes, state assumptions:

  • Scale: 1M merchants, 3M transfers/day, peak 500/sec
  • Consistency: ledger CP (no double-spend); activity feed AP
  • Latency: p99 transfer < 200ms; cross-shard acceptable at 400ms
  • Durability: no acknowledged transfer lost; reconciliation within 24h
  • Availability: 99.95% transfer SLO; brief read staleness OK for dashboards

Ask the interviewer which dimension matters most if time is short. "If I can only optimize one property, should it be correctness or latency?"

Step 2: Draw components and data flow

Sketch: API gateway → fraud → orchestrator → shard A debit → shard B credit → outbox → settlement consumer. Label sync vs async, CP vs AP per component.

Name ShardPay's actual choices: orchestrated saga (not 2PC), consistent hash sharding, cell-based routing, Kafka at-least-once with idempotent consumers.

Step 3: Name failure modes

For each arrow in the diagram, ask "what if this fails?"

FailureImpactShardPay mitigation
Network partition between shardsSaga stuck in DEBITEDStep timeout + compensation
Coordinator crashSaga pause, not duplicateDurable saga log + leader election
Duplicate Kafka eventDouble credit riskIdempotency key on consumer
Hot shardp99 latency spikeConsistent hash + sub-shards
AZ outage1/3 replica lossQuorum 2/3 writes continue
Bad deployRegression in transfer pathCell-scoped canary + rollback
  • Network partition between shards

    ImpactSaga stuck in DEBITED
    ShardPay mitigationStep timeout + compensation
  • Coordinator crash

    ImpactSaga pause, not duplicate
    ShardPay mitigationDurable saga log + leader election
  • Duplicate Kafka event

    ImpactDouble credit risk
    ShardPay mitigationIdempotency key on consumer
  • Hot shard

    Impactp99 latency spike
    ShardPay mitigationConsistent hash + sub-shards
  • AZ outage

    Impact1/3 replica loss
    ShardPay mitigationQuorum 2/3 writes continue
  • Bad deploy

    ImpactRegression in transfer path
    ShardPay mitigationCell-scoped canary + rollback

Common failure modes and mitigations

Step 4: Propose mitigations and observability

Mitigations without observability are unverifiable. Every ShardPay answer ends with: "I'd know this works because..."

  • Transfer SLI at gateway with 99.95% SLO and error budget alerts
  • Distributed traces with transferId baggage across all saga steps
  • Saga age metrics for stuck DEBITED detection
  • Nightly reconciliation as last-line defense
  • Game days testing AZ and cell failure quarterly

Key points

  • Requirements first — CAP stance, latency SLA, read/write ratio, durability guarantees. ShardPay's ledger is CP; saying "we need ACID everywhere" without qualifying scope fails interviews.
  • Failure modes per component — partition, crash, slow node, duplicate message, clock skew. Walk the data flow arrow by arrow, not as a generic list.
  • Mitigation catalog — retries with backoff, idempotency keys, quorum reads, saga + compensation, circuit breaker + bulkhead, cell isolation. Match mitigation to specific failure, not spray-and-pray.
  • Observability as proof — SLI/SLO, golden signals, traces, reconciliation. "We'd run a game day killing one AZ" demonstrates operational maturity.

Common traps and ShardPay answers

Interview traps recur across companies. Prepare one-paragraph responses anchored in this track's examples.

"We'll use 2PC for cross-shard transfers"

Trap: atomicity without discussing blocking and partition intolerance.

ShardPay answer: "We retired XA in 2023 after an 11-minute lock outage. Cross-shard transfers use orchestrated sagas with compensating debits, participant idempotency keys, and nightly reconciliation. We accept brief intermediate inconsistency bounded by 30-second step timeouts."

"Exactly-once delivery"

Trap: claiming end-to-end exactly-once without idempotency at every side effect.

ShardPay answer: "Kafka gives at-least-once; we achieve effective exactly-once with idempotent producers, consumer dedup by eventId, and Idempotency-Key on transfer API. Exactly-once is a pipeline property, not a broker toggle."

"Microservices are always better"

Trap: ignoring operational cost and distributed debugging complexity.

ShardPay answer: "We split along team boundaries (Conway) into ledger, fraud, merchant, settlement stream teams. Each service has on-call, SLO, and deploy pipeline. The monolith was correct at 10k merchants; cells + services at 1M. Operational cost is 4 on-call rotations and cross-service tracing — worth it at our scale."

Java
// Interview snippet — tie idempotency + saga + observability together
@PostMapping("/transfers")
public ResponseEntity<TransferResponse> createTransfer(
        @RequestHeader("Idempotency-Key") String idempotencyKey,
        @RequestBody TransferRequest req) {

    Optional<TransferResponse> cached = idempotencyStore.get(idempotencyKey);
    if (cached.isPresent()) return ResponseEntity.ok(cached.get());

    Span span = tracer.spanBuilder("transfer.create")
        .setAttribute("merchant.id", req.merchantId())
        .startSpan();
    try (Scope s = span.makeCurrent()) {
        SagaInstance saga = orchestrator.start(req, idempotencyKey);
        TransferResponse resp = TransferResponse.from(saga);
        idempotencyStore.put(idempotencyKey, resp);
        sliRecorder.recordTransfer(resp, span);
        return ResponseEntity.accepted().body(resp);
    } finally {
        span.end();
    }
}

Module map for quick recall

Link interview topics to track modules for deeper review:

TopicModuleShardPay example
CAP / consistency04CP ledger, AP feeds
Replication / sharding05Quorum shards, consistent hash
Consensus / locks06etcd leader election for orchestrator
Time / ordering07Version vectors for profile conflicts
Reliability patterns08Circuit breaker on fraud, idempotency keys
Messaging09Outbox, at-least-once consumers
Distributed transactions10Saga orchestration, reconciliation
Observability11SLI/SLO, traces, golden signals
Expert playbook12Cells, failure domains, Conway
  • CAP / consistency

    Module04
    ShardPay exampleCP ledger, AP feeds
  • Replication / sharding

    Module05
    ShardPay exampleQuorum shards, consistent hash
  • Consensus / locks

    Module06
    ShardPay exampleetcd leader election for orchestrator
  • Time / ordering

    Module07
    ShardPay exampleVersion vectors for profile conflicts
  • Reliability patterns

    Module08
    ShardPay exampleCircuit breaker on fraud, idempotency keys
  • Messaging

    Module09
    ShardPay exampleOutbox, at-least-once consumers
  • Distributed transactions

    Module10
    ShardPay exampleSaga orchestration, reconciliation
  • Observability

    Module11
    ShardPay exampleSLI/SLO, traces, golden signals
  • Expert playbook

    Module12
    ShardPay exampleCells, failure domains, Conway

Track module map for interview prep

Key points

  • Structured answer beats clever detail — four steps every time beats jumping to Kafka tuning without requirements.
  • Trade-offs explicitly stated — "We chose saga over 2PC because availability and partition tolerance matter more than instant cross-shard atomicity for our SLO."
  • Production validation — mention how you'd verify: SLI, game day, reconciliation, canary deploy. Theory without ops is incomplete.
  • Spaced recall — revisit modules 04, 08, 10, 11 closest to interview date; they cover the highest-frequency probe topics.

Quick recall

Everything you need if you only revisit this box.

  1. Four steps: clarify requirements → draw flow → name failures → mitigate + observe.
  2. Avoid traps: 2PC for everything, exactly-once without idempotency, microservices without ops cost.
  3. Anchor every trade-off in a ShardPay production example from this track.

Test yourself

Answer these before moving on — recall is what makes it stick.