PrepZone Logo
PrepZone

PACELC — Beyond CAP

Even without partitions, latency forces consistency vs response-time trade-offs.

Why this matters

  • CAP alone misses the everyday latency vs consistency choices that dominate p99 dashboards.
  • Async replication — the default in most databases — is a PACELC trade-off, not a partition scenario.
  • Staff interviews expect PACELC after CAP: "What happens when the network is fine but you read from a replica?"
  • ShardPay documents PACELC per endpoint so API consumers know whether a 5ms read might be stale.
During partition (PAC)
AvailabilityAP — activity feed
ConsistencyCP — ledger
Normal operation (ELC)
LatencyFollower reads
ConsistencyLeader reads
If Partition then choose A or C; Else (normal operation) choose Latency or Consistency.

Beyond the partition

Even with a healthy network, ShardPay's read replica returns balances that trail the leader by 50–200ms. A merchant refreshing their dashboard after a large payout might see the old balance for one poll cycle. That is the EL branch of PACELC: Else (no partition), choose Latency over Consistency.

The PA/PC branch still matters when AZs split, but most incident tickets at ShardPay come from replica lag under load — not from formal network partitions.

PACELC branches

  • PA (Partition → Availability) — During a network split, favor serving requests even if replicas disagree. ShardPay's notification ingress in each region is PA: events buffer locally rather than block merchants when cross-region links flap.
  • PC (Partition → Consistency) — During a split, favor consistent answers over always-on responses. ShardPay's ledger shard is PC: minority replicas stop accepting writes until quorum returns.
  • EL (Else → Latency) — When the network is healthy, favor fast responses even if reads are slightly stale. ShardPay routes merchant history queries to async replicas with a 30s staleness budget — p99 drops from 40ms (leader) to 8ms (replica).
  • EC (Else → Consistency) — When the network is healthy, pay latency for fresh reads. ShardPay's GET /accounts/{id}/balance after a transfer uses readConcern=majority on the leader — every response reflects all committed debits.

Mapping ShardPay endpoints to PACELC

Production systems spend most of their life in the "Else" branch. Documenting which endpoints are EL vs EC prevents engineers from accidentally calling a replica for a post-transfer balance check.

EndpointPACELCRationale
POST /transfersPC / ECMust commit to quorum; no stale write path
GET /balance (default)ELReplica OK for dashboard polling
GET /balance?fresh=trueECLeader read after user action
GET /transactionsELHistory tolerates seconds of lag
Regional notification ingressPA / ELAccept fast; merge later
  • POST /transfers

    PACELCPC / EC
    RationaleMust commit to quorum; no stale write path
  • GET /balance (default)

    PACELCEL
    RationaleReplica OK for dashboard polling
  • GET /balance?fresh=true

    PACELCEC
    RationaleLeader read after user action
  • GET /transactions

    PACELCEL
    RationaleHistory tolerates seconds of lag
  • Regional notification ingress

    PACELCPA / EL
    RationaleAccept fast; merge later

ShardPay PACELC decisions per endpoint

Walkthrough: Healthy network, stale read

  1. Merchant completes a $10,000 transfer; leader commits in 25ms.
  2. Dashboard JavaScript polls GET /balance against a read replica (EL path).
  3. Replication lag is 120ms; replica still shows the pre-transfer balance.
  4. Merchant clicks "Refresh" — client adds ?fresh=true, routed to leader (EC path).
  5. Balance updates; UI shows a one-time "syncing" badge if replicationLagMs > 100.
Java
// ShardPay read router — PACELC in code
public Balance getBalance(String accountId, boolean fresh) {
    if (fresh) {
        return leaderClient.read(accountId);           // EC: consistency over latency
    }
    return replicaPool.readWithFallback(accountId);  // EL: latency, bounded staleness
}

Tunable consistency without a partition

Modern databases expose read concerns, consistency levels, and session tokens. PACELC gives you vocabulary to explain why those knobs exist and which ShardPay flows set them.

Tunable knobs

  • Read concern / consistency level — Per-query choice of how many replicas must respond. ShardPay uses ONE for analytics rollups and QUORUM for fraud-score inputs that must not act on pre-fraud-transfer balances.
  • Session stickiness — Route a user's reads to a replica that has applied at least their last write (read-your-writes without hitting the leader every time). ShardPay encodes sessionToken in the transfer response; subsequent reads include it so the router picks an caught-up replica.
  • Monotonic read routing — Never move a session backward in time across replicas. ShardPay's router tracks the highest commitIndex seen per session and skips replicas below that index.
  • Default path documentation — Most traffic should hit a documented default. ShardPay's OpenAPI spec tags each GET with pacelc: EL or pacelc: EC so client SDKs generate the right default for balance vs history.

Example: Fraud check needs EC, dashboard needs EL

A fraud rule blocks transfers when daily outflow exceeds $50,000. The rule engine reads cumulative daily debit from the ledger with readConcern=majority (EC) — acting on a stale replica could miss a burst of transfers and approve fraud. The same merchant's dashboard chart aggregates hourly totals from a columnar replica (EL) where 60 seconds of lag is invisible on a bar chart.

Quick recall

Everything you need if you only revisit this box.

  1. PACELC covers normal operation: most trade-offs are latency vs consistency, not partition scenarios.
  2. Async replication is EL — document staleness budgets per endpoint.
  3. ShardPay ledger writes are PC/EC; dashboard reads are EL with an explicit fresh path.

Test yourself

Answer these before moving on — recall is what makes it stick.