Why StreamHub needed a load balancer
At ~20K DAU, StreamHub's single API instance hit connection limits during peak evening viewing. Adding an Application Load Balancer (ALB) in front of four identical pods removed the bottleneck without code changes.
AWS ALB → EKS traffic path
StreamHub production architecture (AWS)
Load balancer responsibilities
- Distribute requests — round-robin, least connections, weighted, or hash-based.
- Health checks — remove unhealthy backends automatically.
- TLS termination — decrypt HTTPS at the edge, forward HTTP internally.
- Sticky sessions — route same client to same backend when needed.
- Path-based routing —
/api/*to API pool,/admin/*to admin pool.
Layer 4 vs Layer 7
| Aspect | L4 (Transport) | L7 (Application) |
|---|---|---|
| Routes on | IP + port | URL path, headers, cookies |
| Awareness | TCP/UDP only | HTTP headers, WebSocket upgrade |
| Examples | AWS NLB, HAProxy TCP mode | AWS ALB, NGINX, Envoy |
| StreamHub | Internal DB proxy (rare) | Public API entry point |
Routes on
L4 (Transport)IP + portL7 (Application)URL path, headers, cookiesAwareness
L4 (Transport)TCP/UDP onlyL7 (Application)HTTP headers, WebSocket upgradeExamples
L4 (Transport)AWS NLB, HAProxy TCP modeL7 (Application)AWS ALB, NGINX, EnvoyStreamHub
L4 (Transport)Internal DB proxy (rare)L7 (Application)Public API entry point
Most HTTP APIs use L7 balancers for path routing and header inspection.
Load balancing algorithms
| Algorithm | Behaviour | When to use |
|---|---|---|
| Round-robin | Cycle through backends in order | Equal-capacity stateless servers |
| Least connections | Send to fewest active connections | Long-lived requests, varying work |
| Weighted | Proportional traffic by weight | Canary deploys, mixed instance sizes |
| IP hash | Same client IP → same backend | Simple stickiness without cookies |
| Consistent hash | Hash on header/key | Cache-friendly routing (advanced) |
Round-robin
BehaviourCycle through backends in orderWhen to useEqual-capacity stateless serversLeast connections
BehaviourSend to fewest active connectionsWhen to useLong-lived requests, varying workWeighted
BehaviourProportional traffic by weightWhen to useCanary deploys, mixed instance sizesIP hash
BehaviourSame client IP → same backendWhen to useSimple stickiness without cookiesConsistent hash
BehaviourHash on header/keyWhen to useCache-friendly routing (advanced)
# AWS ALB target group — StreamHub API
target_group: streamhub-api-tg
protocol: HTTP
port: 8080
health_check:
path: /health
interval: 15
healthy_threshold: 2
unhealthy_threshold: 3
algorithm: least_outstanding_requests
Health checks and graceful shutdown
Backends that fail health checks are drained from the pool. Implement /health to verify dependencies:
GET /health
{
"status": "ok",
"checks": {
"database": "ok",
"redis": "ok",
"disk_free_pct": 42
}
}
On deploy, send SIGTERM, stop accepting new connections, finish in-flight requests, then exit — the LB stops routing within one health-check interval.
Sticky sessions
Stateless APIs avoid stickiness. Use it when:
- In-memory session without Redis (legacy).
- WebSocket connections tied to a specific server.
Prefer external session stores over stickiness — sticky routing breaks when instances scale down.
# ALB sticky session cookie (avoid if possible)
Set-Cookie: AWSALB=...; Path=/; HttpOnly
TLS termination
StreamHub terminates TLS at the ALB. Traffic inside the VPC travels HTTP to app pods — acceptable with network isolation; some teams re-encrypt with mTLS via service mesh.
# Client sees HTTPS; internal hop may be HTTP
curl -v https://api.streamhub.com/v1/feed
# * SSL connection using TLSv1.3
Scaling the load balancer itself
Managed LBs (ALB, Cloud LB) scale automatically. Self-hosted NGINX pairs use DNS round-robin or anycast VIPs. At extreme scale, DNS + multiple regional LBs + anycast IP distribute edge load globally.
Quick recall
Everything you need if you only revisit this box.
- Load balancers spread traffic, terminate TLS, and remove unhealthy backends.
- L7 balancers route by HTTP path/headers; L4 by IP and port.
- Round-robin suits equal stateless servers; least-connections for variable work.
- Health checks and graceful shutdown prevent dropped requests during deploys.
- Avoid sticky sessions when Redis or JWT handles session state externally.
- StreamHub added an ALB at ~20K DAU when one instance hit connection limits.
Test yourself
Answer these before moving on — recall is what makes it stick.