PrepZone Logo
PrepZone

Multi-Region Active-Active Design

Write in multiple regions without double-spending — conflict resolution and data residency.

Why this matters

  • Global payment products need regional presence for latency (EU merchant hitting US-only API adds 120ms RTT) and compliance (GDPR data residency).
  • Active-active avoids disaster failover RTO but introduces split-brain and conflict scenarios absent in single-writer designs.
  • ShardPay uses active-passive for the CP ledger, active-active for merchant profile reads, and geo-routed API ingress — a deliberate per-data-type strategy.
  • Interviewers probe whether you understand why financial ledgers rarely go active-active without strong conflict semantics.
US
US regionfull stack
EU
EU regionfull stack
APAC
APAC regionfull stack

Async replication between regions; conflict resolution per entity type.

Users route to the nearest healthy region via GeoDNS.

Regional deployment strategies

Three patterns dominate: active-passive (one write region), active-active reads (writes in one region, reads everywhere), and full active-active (writes in multiple regions).

ShardPay's ledger is active-passive: primary writes in us-east-1, async replica in eu-west-1 for disaster recovery. Failover promotes EU replica to primary — RTO 15 minutes, RPO 5 seconds of async lag. This CP choice prevents split-brain balance corruption.

ShardPay's merchant API and profile cache are active-active: both US and EU regions accept writes behind geo-DNS. Profile updates use optimistic locking with version vectors; conflicts on non-financial fields merge with last-write-wins.

Key points

  • Active-passive — one designated write region; standby for disaster failover. ShardPay's ledger primary in us-east-1 rejects EU-origin writes routed incorrectly — enforced at cell router.
  • Active-active — multiple regions accept writes concurrently. ShardPay's merchant profile service runs active-active; the ledger intentionally does not.
  • Conflict resolution — required when two regions write the same record. ShardPay uses version vectors for profiles, application-level merge for notification preferences, and rejects conflicts for ledger (single writer only).
  • Data residency — legal requirement to keep EU citizen data in EU. ShardPay pins EU merchants to EU cells; cross-border transfer sagas carry compliance metadata for audit.

ShardPay multi-region architecture

ShardPay operates in us-east-1 (primary) and eu-west-1 (EU presence) with geo-DNS routing merchants to nearest healthy region.

Ledger: active-passive CP

Financial correctness trumps regional write latency. EU merchants hit EU API edge, but ledger writes proxy to US primary shard (unless merchant is in EU cell 4 with EU-local shards — a hybrid model). Async replication streams WAL to EU standby. Failover is manual with confirmation — automatic failover risks split brain.

Merchant API: active-active AP

Profile, API keys, webhook config accept writes in both regions. Version field increments on every update; stale writes return HTTP 409 Conflict. ShardPay's dashboard handles conflict with "your session is stale, refresh" UX.

Read path optimization

Merchant balance reads served from local read replica with max_replica_lag_ms=500 guard. If replica lag exceeds threshold, read falls back to primary — trading latency for freshness on balance display.

Java
// ShardPay regional write with version conflict detection
public ProfileUpdateResult updateProfile(String merchantId,
                                          ProfilePatch patch,
                                          long clientVersion) {
    Profile current = profileRepo.find(merchantId);
    if (current.version() != clientVersion) {
        return ProfileUpdateResult.conflict(current.version(), current.data());
    }
    Profile updated = current.apply(patch).incrementVersion();
    profileRepo.save(updated);
    replicationBus.publish(ProfileUpdatedEvent.from(updated)); // to other region
    return ProfileUpdateResult.success(updated);
}

Conflict handling and failover

Active-active without conflict strategy is data loss waiting to happen. ShardPay documents resolution per entity type.

Walkthrough: simultaneous profile edits in US and EU

  1. Merchant admin A in New York sets webhook URL to https://us.example.com/hook (version 5 → 6) in us-east-1.
  2. Admin B in Berlin sets webhook URL to https://eu.example.com/hook (version 5 → 6) in eu-west-1 — same base version.
  3. US write succeeds first; version 6 replicates to EU.
  4. EU write detects version mismatch (expected 5, found 6) → returns 409 to admin B with current state.
  5. Admin B refreshes, sees US admin's URL, decides whether to overwrite (version 6 → 7).

For ledger, this scenario cannot happen — single write region prevents concurrent conflicting balance updates.

Failover drill: US region unavailable

  1. Geo-DNS health checks fail for us-east-1 after 3 minutes.
  2. Traffic routes to eu-west-1 API edge — already active-active, no change needed for API.
  3. Ledger failover: manual promotion of EU standby to primary (runbook step, ~15 min).
  4. EU cell 4 absorbs US merchant traffic temporarily; cross-region capacity pre-provisioned at 130%.
  5. RPO verified: max 5-second replication lag means at most 5 seconds of writes at risk — reconciled from WAL on promotion.

Key points

  • RTO vs RPO — Recovery Time Objective (how fast) vs Recovery Point Objective (how much data loss). ShardPay's ledger failover RTO is 15 minutes, RPO is 5 seconds with async replication.
  • Split brain prevention — two primaries accepting writes corrupts balances. ShardPay uses etcd fencing tokens and manual promotion checklist — never automatic for ledger.
  • CRDTs and LWW — Conflict-free Replicated Data Types or Last-Write-Wins for AP data. ShardPay uses LWW with wall-clock timestamps for notification prefs (acceptable staleness); not for money.
  • Geo-DNS and anycast — route users to nearest healthy region. ShardPay uses health-checked DNS with 60-second TTL; mobile SDKs cache region endpoint with fallback list.

Quick recall

Everything you need if you only revisit this box.

  1. Active-active needs explicit conflict resolution; ledgers usually stay active-passive (CP).
  2. ShardPay: active-passive ledger, active-active profiles, geo-routed API.
  3. Failover trades RTO for correctness — manual ledger promotion prevents split brain.

Test yourself

Answer these before moving on — recall is what makes it stick.