Stage 1: Single server
Every product starts here: one machine runs the web app, API, and database.
StreamHub at launch (2024)
- Stack — Node.js monolith + Postgres on one VPS.
- Users — ~500 DAU, friends and early creators.
- Bottleneck signal — none yet; simplicity is the feature.
- Ops — one deploy script, one backup cron.
# Deploy v0 — literally one box
ssh streamhub-prod
git pull && npm run build && pm2 restart api
pg_dump streamhub > /backups/$(date +%F).sql
This is correct engineering for the stage. Premature distribution adds failure modes without users to justify them.
Stage 2: Separate web and database
When CPU contention appears, split the database onto its own host (or managed RDS). The app tier can still be one instance.
StreamHub hit this at ~5K DAU when nightly analytics queries slowed live uploads.
Stage 3: Load balancer + multiple app servers
Horizontal scaling begins when one app process cannot handle concurrent connections.
AWS ALB → EKS traffic path
- Session state moves to Redis.
- File uploads go to S3 (not local disk).
- Health checks remove unhealthy instances automatically.
POST /api/v1/videos
Authorization: Bearer <token>
Content-Type: application/json
{ "title": "Sunset timelapse", "s3_key": "uploads/tmp/abc123" }
At 50K DAU, StreamHub ran four API instances behind an ALB in one region.
Stage 4: Database read replicas + cache
Read-heavy workloads (feeds, video metadata, creator profiles) overwhelm the primary.
StreamHub read path (Stage 4)
- Primary — writes only (uploads, likes, follows).
- Replicas — feed assembly, search suggestions, analytics.
- Redis — hot video metadata, session tokens, rate-limit counters.
- Cache-aside — app checks Redis, on miss reads replica, populates cache.
GET video:meta:v_9182
# miss → SELECT from replica → SETEX video:meta:v_9182 3600 <json>
Replica lag of 1–2 seconds was acceptable for StreamHub feeds — stated explicitly in NFRs.
Stage 5: CDN + object storage
Video files and static assets (thumbnails, JS bundles) must not traverse the origin for every view.
StreamHub serves .mp4 segments and poster images from a CDN. Origin stores only in S3; edge PoPs cache popular content. Global p99 startup latency dropped from 800 ms to under 120 ms.
# S3 → CDN origin pull
bucket: streamhub-media-prod
cdn_behavior:
path: /segments/*
ttl_default: 86400
compress: true
Stage 6: Async workers + message queue
Heavy work — transcoding, thumbnail generation, email digests — must not block API responses.
Amazon MSK event pipeline
// Message published after upload accepted
{
"event": "video.uploaded",
"video_id": "v_9182",
"s3_key": "uploads/tmp/abc123",
"profiles": ["360p", "720p", "1080p"]
}
Worker pools scale independently from API pods. A transcode backlog no longer stalls the upload endpoint.
Stage 7: Database sharding + multi-region
At millions of DAU, a single primary cannot absorb write volume. StreamHub shards by creator_id — all of a creator's videos and stats live on one shard, limiting cross-shard joins.
| Aspect | Before sharding | After sharding |
|---|---|---|
| Write ceiling | ~10K writes/sec on one primary | N shards × per-shard limit |
| Queries | Any join in one DB | Shard-key in every hot query |
| Ops | One backup pipeline | Per-shard migrations, rebalancing |
| StreamHub trigger | 500K DAU, upload spikes | 3M DAU, global creators |
Write ceiling
Before sharding~10K writes/sec on one primaryAfter shardingN shards × per-shard limitQueries
Before shardingAny join in one DBAfter shardingShard-key in every hot queryOps
Before shardingOne backup pipelineAfter shardingPer-shard migrations, rebalancingStreamHub trigger
Before sharding500K DAU, upload spikesAfter sharding3M DAU, global creators
Multi-region active-active came later for latency — users in APAC hit Singapore cells, not US-East.
Evolution principles
| Principle | Application |
|---|---|
| Solve today's bottleneck | Do not shard at 1K users |
| Measure before migrating | Profile, load-test, then add complexity |
| Keep rollback paths | Feature flags for cache, dual-write during shard migration |
| Document NFR changes | Feed staleness OK → replica lag OK |
Solve today's bottleneck
ApplicationDo not shard at 1K usersMeasure before migrating
ApplicationProfile, load-test, then add complexityKeep rollback paths
ApplicationFeature flags for cache, dual-write during shard migrationDocument NFR changes
ApplicationFeed staleness OK → replica lag OK
Quick recall
Everything you need if you only revisit this box.
- Products evolve: monolith → LB + app pool → replicas + cache → CDN → queue → shards.
- StreamHub started on one VPS; each stage fixed a measured bottleneck.
- Stateless APIs and external object storage enable horizontal app scaling.
- Read replicas and Redis address read-heavy feeds; async queues offload transcoding.
- CDN serves video segments globally; origin stays in object storage.
- Shard when write volume exceeds one primary — partition by a stable key like creator_id.
Test yourself
Answer these before moving on — recall is what makes it stick.