Why this matters
- This article consolidates the Distributed Systems track into an interview-ready framework.
- Interviewers at senior/staff level rarely ask trivia — they present ambiguous problems and evaluate structured thinking.
- ShardPay examples from modules 1–12 provide concrete anchors for every abstract trade-off.
- Revisit this article after completing all modules for spaced recall before interview loops.
The four-step framework
Every distributed systems design answer at ShardPay interview bar follows four steps. Skip any step and the answer feels junior.
Step 1: Clarify requirements
Before drawing boxes, state assumptions:
- Scale: 1M merchants, 3M transfers/day, peak 500/sec
- Consistency: ledger CP (no double-spend); activity feed AP
- Latency: p99 transfer < 200ms; cross-shard acceptable at 400ms
- Durability: no acknowledged transfer lost; reconciliation within 24h
- Availability: 99.95% transfer SLO; brief read staleness OK for dashboards
Ask the interviewer which dimension matters most if time is short. "If I can only optimize one property, should it be correctness or latency?"
Step 2: Draw components and data flow
Sketch: API gateway → fraud → orchestrator → shard A debit → shard B credit → outbox → settlement consumer. Label sync vs async, CP vs AP per component.
Name ShardPay's actual choices: orchestrated saga (not 2PC), consistent hash sharding, cell-based routing, Kafka at-least-once with idempotent consumers.
Step 3: Name failure modes
For each arrow in the diagram, ask "what if this fails?"
| Failure | Impact | ShardPay mitigation |
|---|---|---|
| Network partition between shards | Saga stuck in DEBITED | Step timeout + compensation |
| Coordinator crash | Saga pause, not duplicate | Durable saga log + leader election |
| Duplicate Kafka event | Double credit risk | Idempotency key on consumer |
| Hot shard | p99 latency spike | Consistent hash + sub-shards |
| AZ outage | 1/3 replica loss | Quorum 2/3 writes continue |
| Bad deploy | Regression in transfer path | Cell-scoped canary + rollback |
Network partition between shards
ImpactSaga stuck in DEBITEDShardPay mitigationStep timeout + compensationCoordinator crash
ImpactSaga pause, not duplicateShardPay mitigationDurable saga log + leader electionDuplicate Kafka event
ImpactDouble credit riskShardPay mitigationIdempotency key on consumerHot shard
Impactp99 latency spikeShardPay mitigationConsistent hash + sub-shardsAZ outage
Impact1/3 replica lossShardPay mitigationQuorum 2/3 writes continueBad deploy
ImpactRegression in transfer pathShardPay mitigationCell-scoped canary + rollback
Common failure modes and mitigations
Step 4: Propose mitigations and observability
Mitigations without observability are unverifiable. Every ShardPay answer ends with: "I'd know this works because..."
- Transfer SLI at gateway with 99.95% SLO and error budget alerts
- Distributed traces with transferId baggage across all saga steps
- Saga age metrics for stuck DEBITED detection
- Nightly reconciliation as last-line defense
- Game days testing AZ and cell failure quarterly
Key points
- Requirements first — CAP stance, latency SLA, read/write ratio, durability guarantees. ShardPay's ledger is CP; saying "we need ACID everywhere" without qualifying scope fails interviews.
- Failure modes per component — partition, crash, slow node, duplicate message, clock skew. Walk the data flow arrow by arrow, not as a generic list.
- Mitigation catalog — retries with backoff, idempotency keys, quorum reads, saga + compensation, circuit breaker + bulkhead, cell isolation. Match mitigation to specific failure, not spray-and-pray.
- Observability as proof — SLI/SLO, golden signals, traces, reconciliation. "We'd run a game day killing one AZ" demonstrates operational maturity.
Common traps and ShardPay answers
Interview traps recur across companies. Prepare one-paragraph responses anchored in this track's examples.
"We'll use 2PC for cross-shard transfers"
Trap: atomicity without discussing blocking and partition intolerance.
ShardPay answer: "We retired XA in 2023 after an 11-minute lock outage. Cross-shard transfers use orchestrated sagas with compensating debits, participant idempotency keys, and nightly reconciliation. We accept brief intermediate inconsistency bounded by 30-second step timeouts."
"Exactly-once delivery"
Trap: claiming end-to-end exactly-once without idempotency at every side effect.
ShardPay answer: "Kafka gives at-least-once; we achieve effective exactly-once with idempotent producers, consumer dedup by eventId, and Idempotency-Key on transfer API. Exactly-once is a pipeline property, not a broker toggle."
"Microservices are always better"
Trap: ignoring operational cost and distributed debugging complexity.
ShardPay answer: "We split along team boundaries (Conway) into ledger, fraud, merchant, settlement stream teams. Each service has on-call, SLO, and deploy pipeline. The monolith was correct at 10k merchants; cells + services at 1M. Operational cost is 4 on-call rotations and cross-service tracing — worth it at our scale."
// Interview snippet — tie idempotency + saga + observability together
@PostMapping("/transfers")
public ResponseEntity<TransferResponse> createTransfer(
@RequestHeader("Idempotency-Key") String idempotencyKey,
@RequestBody TransferRequest req) {
Optional<TransferResponse> cached = idempotencyStore.get(idempotencyKey);
if (cached.isPresent()) return ResponseEntity.ok(cached.get());
Span span = tracer.spanBuilder("transfer.create")
.setAttribute("merchant.id", req.merchantId())
.startSpan();
try (Scope s = span.makeCurrent()) {
SagaInstance saga = orchestrator.start(req, idempotencyKey);
TransferResponse resp = TransferResponse.from(saga);
idempotencyStore.put(idempotencyKey, resp);
sliRecorder.recordTransfer(resp, span);
return ResponseEntity.accepted().body(resp);
} finally {
span.end();
}
}
Module map for quick recall
Link interview topics to track modules for deeper review:
| Topic | Module | ShardPay example |
|---|---|---|
| CAP / consistency | 04 | CP ledger, AP feeds |
| Replication / sharding | 05 | Quorum shards, consistent hash |
| Consensus / locks | 06 | etcd leader election for orchestrator |
| Time / ordering | 07 | Version vectors for profile conflicts |
| Reliability patterns | 08 | Circuit breaker on fraud, idempotency keys |
| Messaging | 09 | Outbox, at-least-once consumers |
| Distributed transactions | 10 | Saga orchestration, reconciliation |
| Observability | 11 | SLI/SLO, traces, golden signals |
| Expert playbook | 12 | Cells, failure domains, Conway |
CAP / consistency
Module04ShardPay exampleCP ledger, AP feedsReplication / sharding
Module05ShardPay exampleQuorum shards, consistent hashConsensus / locks
Module06ShardPay exampleetcd leader election for orchestratorTime / ordering
Module07ShardPay exampleVersion vectors for profile conflictsReliability patterns
Module08ShardPay exampleCircuit breaker on fraud, idempotency keysMessaging
Module09ShardPay exampleOutbox, at-least-once consumersDistributed transactions
Module10ShardPay exampleSaga orchestration, reconciliationObservability
Module11ShardPay exampleSLI/SLO, traces, golden signalsExpert playbook
Module12ShardPay exampleCells, failure domains, Conway
Track module map for interview prep
Key points
- Structured answer beats clever detail — four steps every time beats jumping to Kafka tuning without requirements.
- Trade-offs explicitly stated — "We chose saga over 2PC because availability and partition tolerance matter more than instant cross-shard atomicity for our SLO."
- Production validation — mention how you'd verify: SLI, game day, reconciliation, canary deploy. Theory without ops is incomplete.
- Spaced recall — revisit modules 04, 08, 10, 11 closest to interview date; they cover the highest-frequency probe topics.
Quick recall
Everything you need if you only revisit this box.
- Four steps: clarify requirements → draw flow → name failures → mitigate + observe.
- Avoid traps: 2PC for everything, exactly-once without idempotency, microservices without ops cost.
- Anchor every trade-off in a ShardPay production example from this track.
Test yourself
Answer these before moving on — recall is what makes it stick.