Why this matters
- CAP alone misses the everyday latency vs consistency choices that dominate p99 dashboards.
- Async replication — the default in most databases — is a PACELC trade-off, not a partition scenario.
- Staff interviews expect PACELC after CAP: "What happens when the network is fine but you read from a replica?"
- ShardPay documents PACELC per endpoint so API consumers know whether a 5ms read might be stale.
Beyond the partition
Even with a healthy network, ShardPay's read replica returns balances that trail the leader by 50–200ms. A merchant refreshing their dashboard after a large payout might see the old balance for one poll cycle. That is the EL branch of PACELC: Else (no partition), choose Latency over Consistency.
The PA/PC branch still matters when AZs split, but most incident tickets at ShardPay come from replica lag under load — not from formal network partitions.
PACELC branches
- PA (Partition → Availability) — During a network split, favor serving requests even if replicas disagree. ShardPay's notification ingress in each region is PA: events buffer locally rather than block merchants when cross-region links flap.
- PC (Partition → Consistency) — During a split, favor consistent answers over always-on responses. ShardPay's ledger shard is PC: minority replicas stop accepting writes until quorum returns.
- EL (Else → Latency) — When the network is healthy, favor fast responses even if reads are slightly stale. ShardPay routes merchant history queries to async replicas with a 30s staleness budget — p99 drops from 40ms (leader) to 8ms (replica).
- EC (Else → Consistency) — When the network is healthy, pay latency for fresh reads. ShardPay's
GET /accounts/{id}/balanceafter a transfer usesreadConcern=majorityon the leader — every response reflects all committed debits.
Mapping ShardPay endpoints to PACELC
Production systems spend most of their life in the "Else" branch. Documenting which endpoints are EL vs EC prevents engineers from accidentally calling a replica for a post-transfer balance check.
| Endpoint | PACELC | Rationale |
|---|---|---|
| POST /transfers | PC / EC | Must commit to quorum; no stale write path |
| GET /balance (default) | EL | Replica OK for dashboard polling |
| GET /balance?fresh=true | EC | Leader read after user action |
| GET /transactions | EL | History tolerates seconds of lag |
| Regional notification ingress | PA / EL | Accept fast; merge later |
POST /transfers
PACELCPC / ECRationaleMust commit to quorum; no stale write pathGET /balance (default)
PACELCELRationaleReplica OK for dashboard pollingGET /balance?fresh=true
PACELCECRationaleLeader read after user actionGET /transactions
PACELCELRationaleHistory tolerates seconds of lagRegional notification ingress
PACELCPA / ELRationaleAccept fast; merge later
ShardPay PACELC decisions per endpoint
Walkthrough: Healthy network, stale read
- Merchant completes a $10,000 transfer; leader commits in 25ms.
- Dashboard JavaScript polls
GET /balanceagainst a read replica (EL path). - Replication lag is 120ms; replica still shows the pre-transfer balance.
- Merchant clicks "Refresh" — client adds
?fresh=true, routed to leader (EC path). - Balance updates; UI shows a one-time "syncing" badge if
replicationLagMs > 100.
// ShardPay read router — PACELC in code
public Balance getBalance(String accountId, boolean fresh) {
if (fresh) {
return leaderClient.read(accountId); // EC: consistency over latency
}
return replicaPool.readWithFallback(accountId); // EL: latency, bounded staleness
}
Tunable consistency without a partition
Modern databases expose read concerns, consistency levels, and session tokens. PACELC gives you vocabulary to explain why those knobs exist and which ShardPay flows set them.
Tunable knobs
- Read concern / consistency level — Per-query choice of how many replicas must respond. ShardPay uses
ONEfor analytics rollups andQUORUMfor fraud-score inputs that must not act on pre-fraud-transfer balances. - Session stickiness — Route a user's reads to a replica that has applied at least their last write (read-your-writes without hitting the leader every time). ShardPay encodes
sessionTokenin the transfer response; subsequent reads include it so the router picks an caught-up replica. - Monotonic read routing — Never move a session backward in time across replicas. ShardPay's router tracks the highest
commitIndexseen per session and skips replicas below that index. - Default path documentation — Most traffic should hit a documented default. ShardPay's OpenAPI spec tags each GET with
pacelc: ELorpacelc: ECso client SDKs generate the right default for balance vs history.
Example: Fraud check needs EC, dashboard needs EL
A fraud rule blocks transfers when daily outflow exceeds $50,000. The rule engine reads cumulative daily debit from the ledger with readConcern=majority (EC) — acting on a stale replica could miss a burst of transfers and approve fraud. The same merchant's dashboard chart aggregates hourly totals from a columnar replica (EL) where 60 seconds of lag is invisible on a bar chart.
Databases trackSee PACELC applied to database engines
Quick recall
Everything you need if you only revisit this box.
- PACELC covers normal operation: most trade-offs are latency vs consistency, not partition scenarios.
- Async replication is EL — document staleness budgets per endpoint.
- ShardPay ledger writes are PC/EC; dashboard reads are EL with an explicit fresh path.
Test yourself
Answer these before moving on — recall is what makes it stick.