Why this matters
- StreamHub serves 80M DAU across 120 countries — a single-region outage in us-east-1 would disconnect millions of viewers during peak events.
- Latency to the nearest region matters: 200 ms round-trip to a distant data centre degrades live chat and video playback.
- Data replication across regions introduces consistency trade-offs that must be designed per entity type.
Multi-region building blocks
- GeoDNS / global load balancer — route users to nearest healthy region (Route 53, Cloudflare, GCP GLB).
- Regional Kubernetes clusters — independent EKS/GKE clusters per region with identical service topology.
- Cross-region replication — async database replication, S3 cross-region replication, Kafka MirrorMaker.
- Conflict resolution — last-write-wins, version vectors, or region-scoped writes for contested data.
- Failover automation — health checks trigger DNS failover when a region degrades.
Active-active topology
Multi-region active-active (AWS)
# Route 53 latency-based routing
apiVersion: v1
kind: ConfigMap
metadata:
name: dns-config
data:
records: |
streamhub.com:
- region: us-east-1
type: A
routing_policy: latency
health_check: /health
- region: eu-west-1
type: A
routing_policy: latency
health_check: /health
- region: ap-south-1
type: A
routing_policy: latency
health_check: /health
Data replication strategies
| Aspect | Data type | Replication approach |
|---|---|---|
| User profiles | Global, low write rate | Async multi-region replication; LWW on conflict |
| Video metadata | Regional write, global read | Write to home region; replicate to read replicas |
| Live chat messages | High write, regional scope | Write to local region only; no cross-region sync |
| Payment ledger | Strong consistency required | Single-region writes; no multi-master |
| Static assets / CDN | Immutable blobs | S3 cross-region replication + CloudFront |
User profiles
Data typeGlobal, low write rateReplication approachAsync multi-region replication; LWW on conflictVideo metadata
Data typeRegional write, global readReplication approachWrite to home region; replicate to read replicasLive chat messages
Data typeHigh write, regional scopeReplication approachWrite to local region only; no cross-region syncPayment ledger
Data typeStrong consistency requiredReplication approachSingle-region writes; no multi-masterStatic assets / CDN
Data typeImmutable blobsReplication approachS3 cross-region replication + CloudFront
Not everything needs global consistency. Match replication strategy to the business requirement per entity.
Handling replication lag
Async replication means a write in us-east-1 may not appear in eu-west-1 for 200–500 ms.
{
"user_id": "usr_42",
"display_name": "Alex",
"home_region": "us-east-1",
"updated_at": "2026-04-01T12:00:00Z",
"version": 7
}
Strategies to manage lag:
Consistency patterns
- Read-your-writes — route subsequent reads to the write region for 5 seconds after a mutation.
- Sticky sessions — pin user to home region via cookie or JWT claim.
- Conflict resolution — compare version numbers; higher version wins; log conflicts for audit.
- Region-scoped data — live stream chat exists only in the region hosting the stream.
Failover procedure
When us-east-1 degrades, traffic shifts to eu-west-1 and ap-south-1.
# Automated failover (Route 53 health check fails → DNS update)
# Manual escalation steps:
# 1. Confirm region health via Grafana dashboard
# 2. Disable us-east-1 in GeoDNS (weight = 0)
# 3. Scale eu-west-1 and ap-south-1 HPA max replicas
# 4. Promote eu-west-1 read replica to primary for global writes
# 5. Monitor error rates and replication catch-up
StreamHub regional deployment
| Region | Role | Capacity |
|---|---|---|
| us-east-1 | Primary + global writes | 40% traffic |
| eu-west-1 | Active read/write | 30% traffic |
| ap-south-1 | Active read/write | 20% traffic |
| ap-northeast-1 | Active read | 10% traffic |
us-east-1
RolePrimary + global writesCapacity40% trafficeu-west-1
RoleActive read/writeCapacity30% trafficap-south-1
RoleActive read/writeCapacity20% trafficap-northeast-1
RoleActive readCapacity10% traffic
Capacity percentages shift during regional events (Cricket World Cup spikes ap-south-1 to 35%).
Cost considerations
Running three full stacks is expensive. Optimise with:
- Regional data residency — keep EU user data in eu-west-1 for GDPR compliance.
- Spot/preemptible nodes for batch workloads (transcoding) in non-primary regions.
- CDN absorbs 90% of read traffic — origin servers only handle cache misses per region.
Quick recall
Everything you need if you only revisit this box.
- Active-active: every region serves live traffic; GeoDNS routes to nearest healthy cluster.
- Match replication strategy per data type — not everything needs global strong consistency.
- Async replication lag (200–500 ms) requires read-your-writes or sticky sessions.
- Failover drills quarterly; untested failover is untrusted failover.
- CDN + regional clusters reduce latency; payment ledgers stay single-region.
Test yourself
Answer these before moving on — recall is what makes it stick.