Why this matters
- Wrong LB algorithm creates hot instances and tail latency.
- Consistent hashing at the edge preserves cache locality.
- Every system design diagram starts with a load balancer.
- Layer 4 vs Layer 7 choice affects routing flexibility and cost.
Algorithms at a glance
A load balancer sits between clients and backend instances, distributing requests so no single node is overwhelmed. The algorithm determines which instance gets each request — and the wrong choice creates hot spots, breaks stateful routing, or wastes cache warmth.
ShardPay uses different algorithms per tier: least-connections for the stateless API (variable request cost), consistent hash on accountId for ledger shard routing (stateful partitioning), and geographic routing at the CDN edge (lowest latency for static assets).
Load balancing algorithms
- Round-robin — send each request to the next instance in rotation. Simple and fair for uniform request cost, but ignores actual load. ShardPay avoids round-robin for the API tier because transfer requests (80ms) and balance lookups (5ms) have very different costs.
- Least connections — route to the instance with the fewest active connections. Adapts to variable request duration. ShardPay's API load balancer uses least-connections, keeping p99 flat when transfer and read traffic mix.
- Consistent hash — map a key (accountId, sessionId) to a backend via a hash ring; same key always routes to the same backend until the ring changes. ShardPay routes ledger writes by
accountIdhash so each shard receives only its partition's traffic. - Weighted routing — send proportionally more traffic to larger instances. ShardPay uses weights during deploys: new pods start at weight 10% and ramp to 100% over 5 minutes to warm JVM caches before taking full load.
Layer 4 vs Layer 7 load balancing
Layer 4 (L4) balances based on IP and port — fast, low overhead, no protocol awareness. Layer 7 (L7) balances based on HTTP headers, URL paths, or gRPC metadata — slower but enables path-based routing, TLS termination, and request inspection.
ShardPay's architecture uses both: an L4 network load balancer for raw TCP throughput to ledger shards, and an L7 application load balancer at the API gateway for path-based routing (/v1/transfers → transfer service, /v1/balances → read service) and JWT validation.
Walkthrough: hot instance from round-robin
ShardPay's API used round-robin across 10 pods. During a merchant bulk-import event, one pod received a burst of 500 complex transfer requests while others handled lightweight balance reads. That pod's p99 hit 800ms; the fleet average was 120ms. Switching to least-connections redistributed the bulk-import requests across pods, bringing p99 to 160ms.
For ledger shards, round-robin would be catastrophic — writes to the wrong shard would fail or require expensive cross-shard forwarding. Consistent hash on accountId ensures each request reaches the shard that owns the data.
// ShardPay gateway — L7 routing with least-connections upstream
@Configuration
public class GatewayRouting {
@Bean
public RouteLocator routes(RouteLocatorBuilder builder) {
return builder.routes()
.route("transfers", r -> r.path("/v1/transfers/**")
.filters(f -> f.requestRateLimiter(c -> c.setRateLimiter(transferRateLimiter)))
.uri("lb://transfer-service")) // least-connections via Spring Cloud LoadBalancer
.route("balances", r -> r.path("/v1/balances/**")
.uri("lb://balance-service"))
.build();
}
}
// Shard-aware routing — consistent hash, not round-robin
public class ShardLoadBalancer {
private final ConsistentHashRing<LedgerShard> ring;
public LedgerShard selectShard(AccountId accountId) {
return ring.getNode(accountId.hashCode()); // same account → same shard
}
}
Health-aware routing and draining
A load balancer must not send traffic to unhealthy instances. ShardPay's LB polls /health every 5 seconds and removes instances that fail two consecutive checks. During deploys, instances enter draining mode: health check returns 503 (remove from rotation) but the instance finishes in-flight requests before shutting down.
Connection draining is critical for long-lived transfers. ShardPay's preStop hook signals draining, waits 15 seconds for active requests to complete, then terminates. Without draining, in-flight transfers fail mid-execution and clients retry — creating duplicate processing risk without idempotency keys.
Operational concerns
- Health-aware routing — only send traffic to instances passing health checks. ShardPay removes a ledger pod from rotation when replica lag exceeds 2 seconds, even if the process is alive.
- Connection draining — gracefully remove an instance from rotation while completing in-flight work. ShardPay's deploy pipeline sets draining 15 seconds before pod termination.
- Sticky sessions — route the same client IP or cookie to the same backend. ShardPay uses stickiness only for WebSocket merchant dashboards; HTTP APIs are stateless with no stickiness.
- SSL termination — L7 LB decrypts TLS and forwards plain HTTP to backends, centralizing certificate management. ShardPay terminates TLS at the gateway and uses mTLS for internal service-to-service traffic.
Load balancing at multiple layers
ShardPay load-balances at three layers: CDN edge (geographic, for static assets), API gateway (least-connections, for stateless compute), and shard router (consistent hash, for stateful ledger). Each layer uses the algorithm matched to its tier's state model.
Misplacing the algorithm is a common mistake: consistent hash at the API layer (unnecessary stickiness), round-robin at the ledger layer (wrong shard), or no health checks anywhere (traffic to dead pods).
Quick recall
Everything you need if you only revisit this box.
- Match algorithm to stateless vs stateful: least-connections for API, consistent hash for shards.
- Health-aware routing and connection draining prevent traffic to dead or terminating pods.
- L4 for throughput, L7 for path-based routing and TLS termination.
Test yourself
Answer these before moving on — recall is what makes it stick.