PrepZone Logo
PrepZone

Multi-Region Architecture

Active-active regions, data replication lag and routing users to the nearest healthy cluster.

Read these first

Why this matters

  • StreamHub serves 80M DAU across 120 countries — a single-region outage in us-east-1 would disconnect millions of viewers during peak events.
  • Latency to the nearest region matters: 200 ms round-trip to a distant data centre degrades live chat and video playback.
  • Data replication across regions introduces consistency trade-offs that must be designed per entity type.

Multi-region building blocks

  • GeoDNS / global load balancer — route users to nearest healthy region (Route 53, Cloudflare, GCP GLB).
  • Regional Kubernetes clusters — independent EKS/GKE clusters per region with identical service topology.
  • Cross-region replication — async database replication, S3 cross-region replication, Kafka MirrorMaker.
  • Conflict resolution — last-write-wins, version vectors, or region-scoped writes for contested data.
  • Failover automation — health checks trigger DNS failover when a region degrades.

Active-active topology

Multi-region active-active (AWS)

REGIONus-east-1ALB + EKS + RDS + S3
REGIONeu-west-1ALB + EKS + RDS + S3
REGIONap-southeast-1ALB + EKS + RDS + S3
NETWORKRoute 53latency routing
INTEGRATIONGlobal MSKcross-region sync
DATABASEDynamoDB Globalmulti-master
Route 53 latency routing → regional EKS stacks → async cross-region replication.
Java
# Route 53 latency-based routing
apiVersion: v1
kind: ConfigMap
metadata:
  name: dns-config
data:
  records: |
    streamhub.com:
      - region: us-east-1
        type: A
        routing_policy: latency
        health_check: /health
      - region: eu-west-1
        type: A
        routing_policy: latency
        health_check: /health
      - region: ap-south-1
        type: A
        routing_policy: latency
        health_check: /health

Data replication strategies

AspectData typeReplication approach
User profilesGlobal, low write rateAsync multi-region replication; LWW on conflict
Video metadataRegional write, global readWrite to home region; replicate to read replicas
Live chat messagesHigh write, regional scopeWrite to local region only; no cross-region sync
Payment ledgerStrong consistency requiredSingle-region writes; no multi-master
Static assets / CDNImmutable blobsS3 cross-region replication + CloudFront
  • User profiles

    Data typeGlobal, low write rate
    Replication approachAsync multi-region replication; LWW on conflict
  • Video metadata

    Data typeRegional write, global read
    Replication approachWrite to home region; replicate to read replicas
  • Live chat messages

    Data typeHigh write, regional scope
    Replication approachWrite to local region only; no cross-region sync
  • Payment ledger

    Data typeStrong consistency required
    Replication approachSingle-region writes; no multi-master
  • Static assets / CDN

    Data typeImmutable blobs
    Replication approachS3 cross-region replication + CloudFront

Not everything needs global consistency. Match replication strategy to the business requirement per entity.

Handling replication lag

Async replication means a write in us-east-1 may not appear in eu-west-1 for 200–500 ms.

Java
{
  "user_id": "usr_42",
  "display_name": "Alex",
  "home_region": "us-east-1",
  "updated_at": "2026-04-01T12:00:00Z",
  "version": 7
}

Strategies to manage lag:

Consistency patterns

  • Read-your-writes — route subsequent reads to the write region for 5 seconds after a mutation.
  • Sticky sessions — pin user to home region via cookie or JWT claim.
  • Conflict resolution — compare version numbers; higher version wins; log conflicts for audit.
  • Region-scoped data — live stream chat exists only in the region hosting the stream.

Failover procedure

When us-east-1 degrades, traffic shifts to eu-west-1 and ap-south-1.

Java
# Automated failover (Route 53 health check fails → DNS update)
# Manual escalation steps:
# 1. Confirm region health via Grafana dashboard
# 2. Disable us-east-1 in GeoDNS (weight = 0)
# 3. Scale eu-west-1 and ap-south-1 HPA max replicas
# 4. Promote eu-west-1 read replica to primary for global writes
# 5. Monitor error rates and replication catch-up

StreamHub regional deployment

RegionRoleCapacity
us-east-1Primary + global writes40% traffic
eu-west-1Active read/write30% traffic
ap-south-1Active read/write20% traffic
ap-northeast-1Active read10% traffic
  • us-east-1

    RolePrimary + global writes
    Capacity40% traffic
  • eu-west-1

    RoleActive read/write
    Capacity30% traffic
  • ap-south-1

    RoleActive read/write
    Capacity20% traffic
  • ap-northeast-1

    RoleActive read
    Capacity10% traffic

Capacity percentages shift during regional events (Cricket World Cup spikes ap-south-1 to 35%).

Cost considerations

Running three full stacks is expensive. Optimise with:

  • Regional data residency — keep EU user data in eu-west-1 for GDPR compliance.
  • Spot/preemptible nodes for batch workloads (transcoding) in non-primary regions.
  • CDN absorbs 90% of read traffic — origin servers only handle cache misses per region.

Quick recall

Everything you need if you only revisit this box.

  • Active-active: every region serves live traffic; GeoDNS routes to nearest healthy cluster.
  • Match replication strategy per data type — not everything needs global strong consistency.
  • Async replication lag (200–500 ms) requires read-your-writes or sticky sessions.
  • Failover drills quarterly; untested failover is untrusted failover.
  • CDN + regional clusters reduce latency; payment ledgers stay single-region.

Test yourself

Answer these before moving on — recall is what makes it stick.