Why this matters
- Dynamo-style quorums appear in Cassandra, Riak, DynamoDB, and custom ShardPay storage paths.
- Interviewers ask you to derive W, R, N for a given durability and latency SLA — show the math.
- SLA targets (RPO, RTO, read staleness) should drive quorum configuration, not framework defaults.
- Misconfigured quorums cause "sometimes stale, sometimes lost" behavior that is hardest to debug.
N=3, W=2, R=2 → strong read after quorum write
Quorum configuration
ShardPay runs N=5 replicas per ledger shard. W=3 on write ensures any commit is on a majority before ack. R=2 on read with W+R>5 guarantees the read set overlaps the write set — a reader always contacts at least one node that participated in the latest commit.
This configuration trades extra write latency (3 round-trips) for read speed (2 nodes) while preserving read-after-write correctness for quorum reads.
Core parameters
- N (replication factor) — Total copies of each key or shard range. ShardPay uses N=5 in production for ledger shards — survives loss of any two replicas without data loss if W uses a majority.
- W (write quorum) — Minimum replicas that must acknowledge a write before success. ShardPay sets W=3 so a committed transfer survives any two simultaneous replica failures.
- R (read quorum) — Minimum replicas consulted on read. ShardPay uses R=2 for standard balance reads when paired with W=3 — satisfies W+R>N overlap property.
- W + R > N — Guarantees read and write quorums intersect, so a read sees the latest committed write. ShardPay's fraud pipeline uses this overlap; analytics uses R=1 (no overlap guarantee) with explicit staleness acceptance.
- Sloppy quorum / hinted handoff — Writes accepted by fallback nodes when primary replicas are down, returned later. ShardPay disables sloppy quorum on the ledger — financial data does not use hinted handoff.
Deriving W and R from SLAs
Start from business requirements, not defaults.
| SLA | Requirement | Quorum implication |
|---|---|---|
| Durability | Zero committed transfer loss if ≤2 replicas die | W > N/2, typically W=3 for N=5 |
| Read freshness | Post-commit read sees new balance | W+R > N |
| Read latency p99 | < 15ms | Lower R (e.g., R=2 not R=5) |
| Write latency p99 | < 40ms | Lower W if durability allows (ShardPay does not lower W) |
Durability
RequirementZero committed transfer loss if ≤2 replicas dieQuorum implicationW > N/2, typically W=3 for N=5Read freshness
RequirementPost-commit read sees new balanceQuorum implicationW+R > NRead latency p99
Requirement< 15msQuorum implicationLower R (e.g., R=2 not R=5)Write latency p99
Requirement< 40msQuorum implicationLower W if durability allows (ShardPay does not lower W)
Mapping ShardPay SLAs to quorum settings
Walkthrough: Prove overlap for ShardPay N=5, W=3, R=2
- Write quorum touches any 3 of 5 nodes — call this set W-set.
- Read quorum touches any 2 of 5 nodes — call this set R-set.
- If W-set and R-set were disjoint, they would need 3+2=5 distinct nodes — the entire cluster.
- Any two sets of size 2 and 3 from 5 elements must intersect in at least one node.
- That intersecting node holds the latest committed value for the key (assuming last-write-wins with version vectors on conflicts — ShardPay uses leader serialization to avoid conflicts).
// ShardPay quorum write — wait for W acknowledgements
public CommitResult quorumWrite(Transfer transfer) {
List<Replica> targets = replicaSet.pickAny(W);
List<Ack> acks = parallelWrite(targets, transfer);
if (acks.size() < W) {
return CommitResult.failed("WRITE_QUORUM_NOT_REACHED");
}
return CommitResult.success(acks.highestVersion());
}
// Quorum read — merge versions from R replicas
public Balance quorumRead(String accountId) {
List<Balance> versions = parallelRead(replicaSet.pickAny(R), accountId);
return versions.maxBy(Balance::version);
}
SLA mapping and operational trade-offs
Tuning guide
- Higher W — More durability and write latency. ShardPay increased W from 2 to 3 after a rare dual-replica loss during async replication exposed a committed-but-lost transfer.
- Lower R — Faster reads, risk of stale data if W+R ≤ N. ShardPay's merchant dashboard uses R=1 with session routing; fraud uses R=2 with W=3.
- R=W=majority — Strongest common pattern; read and write both pay majority latency. ShardPay's
?fresh=truebalance path uses R=3, W=3 for maximum overlap without reading all five nodes. - Read repair — Background fix when R=1 reads detect version mismatch. ShardPay runs read repair on analytics replicas only — not on the critical path for transfers.
- Monitoring — Track
write_quorum_time_ms,read_quorum_miss_stale, andreplica_version_skew. Alert when any replica falls more than 500 versions behind the leader.
Example: Tightening R for settlement batch
ShardPay's end-of-day settlement job aggregates net positions per merchant. A stale read could misstate settlement by millions. The job switches to R=3, W=3 (full majority overlap) for the duration of the batch — p99 read latency rises from 8ms to 22ms, acceptable for a nightly window. Dashboard traffic during the same window stays on R=2.
Quick recall
Everything you need if you only revisit this box.
- W + R > N guarantees overlapping read/write sets for fresh quorum reads.
- Tune W and R per SLA — higher W for durability, lower R for read latency when overlap still holds.
- ShardPay ledger: N=5, W=3, R=2 for default path; settlement uses majority R.
Test yourself
Answer these before moving on — recall is what makes it stick.