PrepZone Logo
PrepZone

Zero to Millions of Users

Follow StreamHub from one server to CDN, cache, queues and sharded databases as traffic grows.

Stage 1EC2 + Postgres on one box
Stage 2RDS + separate EC2 app
Stage 3ALB + ASG + ElastiCache Redis
Stage 4CloudFront + S3 media origin
Stage 5MSK Kafka + EKS worker fleet
Stage 6Sharded RDS + multi-region EKS
Each stage adds one AWS building block when the previous bottleneck appears.

Stage 1: Single server

Every product starts here: one machine runs the web app, API, and database.

StreamHub at launch (2024)

  • Stack — Node.js monolith + Postgres on one VPS.
  • Users — ~500 DAU, friends and early creators.
  • Bottleneck signal — none yet; simplicity is the feature.
  • Ops — one deploy script, one backup cron.
Java
# Deploy v0 — literally one box
ssh streamhub-prod
git pull && npm run build && pm2 restart api
pg_dump streamhub > /backups/$(date +%F).sql

This is correct engineering for the stage. Premature distribution adds failure modes without users to justify them.

Stage 2: Separate web and database

When CPU contention appears, split the database onto its own host (or managed RDS). The app tier can still be one instance.

StreamHub hit this at ~5K DAU when nightly analytics queries slowed live uploads.

Stage 3: Load balancer + multiple app servers

Horizontal scaling begins when one app process cannot handle concurrent connections.

AWS ALB → EKS traffic path

TLSCLIENT
UsersHTTPS requests
NETWORK
Route 53alias → ALB
NETWORK
AWS ALBtarget group
COMPUTE
EKS pod 1streamhub-api
COMPUTE
EKS pod 2streamhub-api
COMPUTE
EKS pod NHPA scaled
L7 ALB terminates TLS, health-checks targets, and fans out to stateless pods.
  • Session state moves to Redis.
  • File uploads go to S3 (not local disk).
  • Health checks remove unhealthy instances automatically.
Java
POST /api/v1/videos
Authorization: Bearer <token>
Content-Type: application/json

{ "title": "Sunset timelapse", "s3_key": "uploads/tmp/abc123" }

At 50K DAU, StreamHub ran four API instances behind an ALB in one region.

Stage 4: Database read replicas + cache

Read-heavy workloads (feeds, video metadata, creator profiles) overwhelm the primary.

StreamHub read path (Stage 4)

  • Primary — writes only (uploads, likes, follows).
  • Replicas — feed assembly, search suggestions, analytics.
  • Redis — hot video metadata, session tokens, rate-limit counters.
  • Cache-aside — app checks Redis, on miss reads replica, populates cache.
Java
GET video:meta:v_9182
# miss → SELECT from replica → SETEX video:meta:v_9182 3600 <json>

Replica lag of 1–2 seconds was acceptable for StreamHub feeds — stated explicitly in NFRs.

Stage 5: CDN + object storage

Video files and static assets (thumbnails, JS bundles) must not traverse the origin for every view.

Stage 1EC2 + Postgres on one box
Stage 2RDS + separate EC2 app
Stage 3ALB + ASG + ElastiCache Redis
Stage 4CloudFront + S3 media origin
Stage 5MSK Kafka + EKS worker fleet
Stage 6Sharded RDS + multi-region EKS
Each stage adds one AWS building block when the previous bottleneck appears.

StreamHub serves .mp4 segments and poster images from a CDN. Origin stores only in S3; edge PoPs cache popular content. Global p99 startup latency dropped from 800 ms to under 120 ms.

Java
# S3 → CDN origin pull
bucket: streamhub-media-prod
cdn_behavior:
  path: /segments/*
  ttl_default: 86400
  compress: true

Stage 6: Async workers + message queue

Heavy work — transcoding, thumbnail generation, email digests — must not block API responses.

Amazon MSK event pipeline

produceCOMPUTE
EKS APIorder placed
INTEGRATION
Amazon MSKorders.placed.v1
COMPUTE
InventoryEKS worker
COMPUTE
Email svcEKS worker
ANALYTICS
AnalyticsFlink / EMR
API publishes events; worker fleets scale independently on consumer lag.
Java
// Message published after upload accepted
{
  "event": "video.uploaded",
  "video_id": "v_9182",
  "s3_key": "uploads/tmp/abc123",
  "profiles": ["360p", "720p", "1080p"]
}

Worker pools scale independently from API pods. A transcode backlog no longer stalls the upload endpoint.

Stage 7: Database sharding + multi-region

At millions of DAU, a single primary cannot absorb write volume. StreamHub shards by creator_id — all of a creator's videos and stats live on one shard, limiting cross-shard joins.

AspectBefore shardingAfter sharding
Write ceiling~10K writes/sec on one primaryN shards × per-shard limit
QueriesAny join in one DBShard-key in every hot query
OpsOne backup pipelinePer-shard migrations, rebalancing
StreamHub trigger500K DAU, upload spikes3M DAU, global creators
  • Write ceiling

    Before sharding~10K writes/sec on one primary
    After shardingN shards × per-shard limit
  • Queries

    Before shardingAny join in one DB
    After shardingShard-key in every hot query
  • Ops

    Before shardingOne backup pipeline
    After shardingPer-shard migrations, rebalancing
  • StreamHub trigger

    Before sharding500K DAU, upload spikes
    After sharding3M DAU, global creators

Multi-region active-active came later for latency — users in APAC hit Singapore cells, not US-East.

Evolution principles

PrincipleApplication
Solve today's bottleneckDo not shard at 1K users
Measure before migratingProfile, load-test, then add complexity
Keep rollback pathsFeature flags for cache, dual-write during shard migration
Document NFR changesFeed staleness OK → replica lag OK
  • Solve today's bottleneck

    ApplicationDo not shard at 1K users
  • Measure before migrating

    ApplicationProfile, load-test, then add complexity
  • Keep rollback paths

    ApplicationFeature flags for cache, dual-write during shard migration
  • Document NFR changes

    ApplicationFeed staleness OK → replica lag OK

Quick recall

Everything you need if you only revisit this box.

  • Products evolve: monolith → LB + app pool → replicas + cache → CDN → queue → shards.
  • StreamHub started on one VPS; each stage fixed a measured bottleneck.
  • Stateless APIs and external object storage enable horizontal app scaling.
  • Read replicas and Redis address read-heavy feeds; async queues offload transcoding.
  • CDN serves video segments globally; origin stays in object storage.
  • Shard when write volume exceeds one primary — partition by a stable key like creator_id.

Test yourself

Answer these before moving on — recall is what makes it stick.