Why this matters
- Global payment products need regional presence for latency (EU merchant hitting US-only API adds 120ms RTT) and compliance (GDPR data residency).
- Active-active avoids disaster failover RTO but introduces split-brain and conflict scenarios absent in single-writer designs.
- ShardPay uses active-passive for the CP ledger, active-active for merchant profile reads, and geo-routed API ingress — a deliberate per-data-type strategy.
- Interviewers probe whether you understand why financial ledgers rarely go active-active without strong conflict semantics.
Async replication between regions; conflict resolution per entity type.
Regional deployment strategies
Three patterns dominate: active-passive (one write region), active-active reads (writes in one region, reads everywhere), and full active-active (writes in multiple regions).
ShardPay's ledger is active-passive: primary writes in us-east-1, async replica in eu-west-1 for disaster recovery. Failover promotes EU replica to primary — RTO 15 minutes, RPO 5 seconds of async lag. This CP choice prevents split-brain balance corruption.
ShardPay's merchant API and profile cache are active-active: both US and EU regions accept writes behind geo-DNS. Profile updates use optimistic locking with version vectors; conflicts on non-financial fields merge with last-write-wins.
Key points
- Active-passive — one designated write region; standby for disaster failover. ShardPay's ledger primary in us-east-1 rejects EU-origin writes routed incorrectly — enforced at cell router.
- Active-active — multiple regions accept writes concurrently. ShardPay's merchant profile service runs active-active; the ledger intentionally does not.
- Conflict resolution — required when two regions write the same record. ShardPay uses version vectors for profiles, application-level merge for notification preferences, and rejects conflicts for ledger (single writer only).
- Data residency — legal requirement to keep EU citizen data in EU. ShardPay pins EU merchants to EU cells; cross-border transfer sagas carry compliance metadata for audit.
ShardPay multi-region architecture
ShardPay operates in us-east-1 (primary) and eu-west-1 (EU presence) with geo-DNS routing merchants to nearest healthy region.
Ledger: active-passive CP
Financial correctness trumps regional write latency. EU merchants hit EU API edge, but ledger writes proxy to US primary shard (unless merchant is in EU cell 4 with EU-local shards — a hybrid model). Async replication streams WAL to EU standby. Failover is manual with confirmation — automatic failover risks split brain.
Merchant API: active-active AP
Profile, API keys, webhook config accept writes in both regions. Version field increments on every update; stale writes return HTTP 409 Conflict. ShardPay's dashboard handles conflict with "your session is stale, refresh" UX.
Read path optimization
Merchant balance reads served from local read replica with max_replica_lag_ms=500 guard. If replica lag exceeds threshold, read falls back to primary — trading latency for freshness on balance display.
// ShardPay regional write with version conflict detection
public ProfileUpdateResult updateProfile(String merchantId,
ProfilePatch patch,
long clientVersion) {
Profile current = profileRepo.find(merchantId);
if (current.version() != clientVersion) {
return ProfileUpdateResult.conflict(current.version(), current.data());
}
Profile updated = current.apply(patch).incrementVersion();
profileRepo.save(updated);
replicationBus.publish(ProfileUpdatedEvent.from(updated)); // to other region
return ProfileUpdateResult.success(updated);
}
Conflict handling and failover
Active-active without conflict strategy is data loss waiting to happen. ShardPay documents resolution per entity type.
Walkthrough: simultaneous profile edits in US and EU
- Merchant admin A in New York sets webhook URL to
https://us.example.com/hook(version 5 → 6) in us-east-1. - Admin B in Berlin sets webhook URL to
https://eu.example.com/hook(version 5 → 6) in eu-west-1 — same base version. - US write succeeds first; version 6 replicates to EU.
- EU write detects version mismatch (expected 5, found 6) → returns 409 to admin B with current state.
- Admin B refreshes, sees US admin's URL, decides whether to overwrite (version 6 → 7).
For ledger, this scenario cannot happen — single write region prevents concurrent conflicting balance updates.
Failover drill: US region unavailable
- Geo-DNS health checks fail for us-east-1 after 3 minutes.
- Traffic routes to eu-west-1 API edge — already active-active, no change needed for API.
- Ledger failover: manual promotion of EU standby to primary (runbook step, ~15 min).
- EU cell 4 absorbs US merchant traffic temporarily; cross-region capacity pre-provisioned at 130%.
- RPO verified: max 5-second replication lag means at most 5 seconds of writes at risk — reconciled from WAL on promotion.
Key points
- RTO vs RPO — Recovery Time Objective (how fast) vs Recovery Point Objective (how much data loss). ShardPay's ledger failover RTO is 15 minutes, RPO is 5 seconds with async replication.
- Split brain prevention — two primaries accepting writes corrupts balances. ShardPay uses etcd fencing tokens and manual promotion checklist — never automatic for ledger.
- CRDTs and LWW — Conflict-free Replicated Data Types or Last-Write-Wins for AP data. ShardPay uses LWW with wall-clock timestamps for notification prefs (acceptable staleness); not for money.
- Geo-DNS and anycast — route users to nearest healthy region. ShardPay uses health-checked DNS with 60-second TTL; mobile SDKs cache region endpoint with fallback list.
Quick recall
Everything you need if you only revisit this box.
- Active-active needs explicit conflict resolution; ledgers usually stay active-passive (CP).
- ShardPay: active-passive ledger, active-active profiles, geo-routed API.
- Failover trades RTO for correctness — manual ledger promotion prevents split brain.
Test yourself
Answer these before moving on — recall is what makes it stick.