Why this matters
- Engineers underestimate WAN latency in multi-region designs.
- Large windows and pipelining exist because of BDP.
- Interviewers ask why cross-region strong consistency is expensive.
- Physics sets hard limits that no amount of optimization can bypass.
BDP = bandwidth × RTT — bytes in flight before ACK returns
Physics of the network
Light travels ~200,000 km/s in fiber. A round trip between US coasts (~5,000 km) takes ~50ms minimum — before switches, queuing, TLS handshakes, or application processing. ShardPay cannot strong-consensus every transfer across regions without paying that tax on every write. This is not a software problem; it is physics.
Bandwidth is how wide the pipe is (bits per second). Latency (RTT) is how long a single bit takes to make a round trip. They are independent dimensions. A 10 Gbps link with 70ms RTT does not deliver 10 Gbps of throughput for small messages — each message waits 70ms for acknowledgment before the next can be sent (without pipelining).
Network fundamentals
- RTT (Round-Trip Time) — time for a packet to reach the destination and return. ShardPay measures ~1ms RTT within an AZ, ~2ms cross-AZ in the same region, and ~70ms cross-coast US. Every synchronous RPC pays this cost twice (request + response).
- BDP (Bandwidth-Delay Product) — bandwidth × RTT = bytes that must be in flight to saturate the pipe. A 10 Gbps link with 70ms RTT has BDP = 10 × 10⁹ × 0.07 / 8 ≈ 87 MB. Without 87 MB in flight, the pipe is underutilized.
- Pipelining — send multiple requests before the first response returns, keeping the pipe full. ShardPay's cross-shard replication pipeline sends up to 64 entries per RTT instead of one, multiplying effective throughput by ~64×.
- Batching — group multiple records into one message to amortize RTT overhead. ShardPay batches up to 50 transfer events per replication RPC, reducing per-transfer network cost from 70ms to ~1.4ms amortized.
Why cross-region consistency is expensive
ShardPay's EU expansion required a regional ledger cell. The alternative — strong-consensus every write across US and EU — would add 130ms RTT (US-East to EU-West) to every transfer. At 5,000 transfers/sec, that is 5,000 synchronous cross-Atlantic round trips per second, consuming bandwidth and making p99 latency unacceptable.
The solution: regional cells with async cross-region replication for disaster recovery, not per-write consensus. RPO for cross-region is 5 seconds (async replication lag), not zero. This is a deliberate trade-off: durability within a region (zero RPO) vs cross-region (seconds of RPO) to avoid paying RTT on every write.
Walkthrough: replication throughput math
ShardPay replicates ledger entries from US-East leader to US-West follower. Link: 10 Gbps, RTT: 70ms. Entry size: 2 KB.
Without pipelining: 1 entry per RTT = 2 KB / 70ms ≈ 28 KB/s ≈ 0.22 Mbps — using 0.002% of the 10 Gbps pipe.
With pipelining (64 entries in flight): 64 × 2 KB / 70ms ≈ 1.8 MB/s ≈ 14 Mbps — still only 0.14% utilization.
With batching (50 entries per message, 64 batches in flight): 64 × 50 × 2 KB / 70ms ≈ 91 MB/s ≈ 730 Mbps — approaching useful utilization.
// ShardPay replication client — pipeline + batch to overcome RTT
public class ReplicationClient {
private static final int PIPELINE_DEPTH = 64;
private static final int BATCH_SIZE = 50;
private final BlockingQueue<LedgerEntry> pending = new LinkedBlockingQueue<>();
private final Channel channel; // gRPC to follower region
public void replicate(LedgerEntry entry) {
pending.offer(entry);
}
@Scheduled(fixedDelay = 5) // flush every 5ms or when batch full
public void flushBatch() {
List<LedgerEntry> batch = new ArrayList<>(BATCH_SIZE);
pending.drainTo(batch, BATCH_SIZE);
if (!batch.isEmpty()) {
channel.replicateBatch(batch); // one RTT for up to 50 entries
}
}
}
Latency budgets across regions
ShardPay's 200ms transfer SLO cannot survive a 130ms network round trip for every hop. Regional cells keep the critical path within one region (~5ms RTT). Cross-region traffic is async: backup replication, analytics export, and disaster-recovery failover — never on the synchronous transfer path.
When designing multi-region systems, map every synchronous call on the critical path and sum the RTTs. If the sum exceeds the SLA, move calls off the critical path (async replication, regional cells, edge caching) or accept higher latency.
Multi-region design patterns
- Regional cells — self-contained deployments per geography; cross-region is async. ShardPay's EU cell handles EU merchants with under 10ms ledger RTT; US cell handles US merchants.
- Read replicas — serve reads locally, accept replication lag. ShardPay's EU dashboard reads from the EU replica (max 2s stale), not the US leader.
- Async replication — batch and pipeline cross-region writes; accept RPO > 0. ShardPay's cross-region RPO is 5 seconds, monitored and alerted.
- Edge caching — cache static and semi-static content at CDN PoPs. ShardPay serves merchant logos and receipt PDFs from edge; transfer APIs always hit origin.
Quick recall
Everything you need if you only revisit this box.
- RTT dominates WAN latency — physics, not bandwidth, sets the floor.
- BDP = bandwidth × RTT; pipelining and batching fill the pipe.
- Cross-region strong consistency costs one RTT per write — regional cells avoid this.
Test yourself
Answer these before moving on — recall is what makes it stick.