Why this matters
- Read replicas are the first scaling lever after connection pooling; misrouting writes or freshness-sensitive reads causes subtle production bugs.
- Failover during Black Friday tests whether your replication topology promotes cleanly in under two minutes.
- Interviews contrast single-leader, multi-leader, and leaderless designs with concrete trade-offs.
Single-leader replication
The most common pattern: one primary accepts writes; replicas apply the WAL asynchronously.
Single-leader characteristics
- Write path — all
INSERT/UPDATE/DELETEgo to the leader. - Read path — replicas serve read traffic with bounded staleness.
- Failover — promote a replica when the leader fails; brief write outage.
- VaultCommerce use — Postgres primary + two read replicas for browse and search hydration.
-- Application routing (HikariCP data sources)
-- writeDataSource → primary.vaultcommerce.db
INSERT INTO orders (id, user_id, total) VALUES ($1, $2, $3);
-- readDataSource → replica.vaultcommerce.db
SELECT id, status, total FROM orders WHERE user_id = $1 ORDER BY created_at DESC;
After checkout, VaultCommerce reads the confirmation page from the primary so the customer never sees "order not found" due to replica lag.
Synchronous vs asynchronous replication
| Aspect | Async replication | Sync replication |
|---|---|---|
| Write latency | Low — leader commits locally | Higher — waits for replica ack |
| Data loss risk | Possible if leader dies before replicate | Minimal on failover |
| Read staleness | Milliseconds to seconds | Near-zero on synced replica |
| VaultCommerce | Browse, recommendations | Payment audit replica |
Write latency
Async replicationLow — leader commits locallySync replicationHigher — waits for replica ackData loss risk
Async replicationPossible if leader dies before replicateSync replicationMinimal on failoverRead staleness
Async replicationMilliseconds to secondsSync replicationNear-zero on synced replicaVaultCommerce
Async replicationBrowse, recommendationsSync replicationPayment audit replica
Semi-sync — wait for one replica before acknowledging — is a middle ground VaultCommerce uses for the payment ledger replica in the same region.
Multi-leader replication
Multiple nodes accept writes, usually partitioned by geography (eu-leader, us-leader). Writes within a region are fast; cross-region conflicts need resolution (last-write-wins, custom merge, or avoidance).
VaultCommerce does not multi-leader Postgres for orders — concurrent inventory decrements across leaders cause overselling. Multi-leader fits calendars, user preferences, or shopping lists where conflicts are rare and mergeable.
Leaderless replication
Dynamo-style quorum reads and writes: no fixed leader; clients contact N nodes and reconcile versions with vector clocks or timestamps. Tunable W and R values trade consistency for availability.
N=3 replicas, W=2, R=2
Write succeeds when 2 nodes ack
Read fetches 2 nodes, returns latest version
VaultCommerce's session store on DynamoDB uses leaderless replication under the hood — the application sees a key-value API, not leader election.
Monitoring replica health
Track pg_stat_replication lag, replication slot backlog, and application-level "read after write" errors. Alert when lag exceeds the product's staleness budget — VaultCommerce pages tolerate 2 seconds on browse, zero on post-payment.
Quick recall
Everything you need if you only revisit this box.
- Single-leader: one write node, replicas for read scale and failover.
- Async replication is fast but stale; sync replication is durable but slower.
- VaultCommerce routes post-checkout reads to primary, browse to replicas.
- Multi-leader suits conflict-tolerant data; avoid it for inventory ledgers.
- Leaderless stores use quorum reads/writes — DynamoDB and Cassandra under the hood.
Test yourself
Answer these before moving on — recall is what makes it stick.