Vertical scaling (scale up)
Vertical scaling increases CPU, RAM, or disk on a single instance. It is simple: no code changes, no distributed-systems complexity.
Vertical scaling pros and cons
- Pros — zero architecture change, no network partition issues, easy ops.
- Cons — hard ceiling (largest instance size), single point of failure, expensive at the top end.
- Best for — early StreamHub MVP, managed databases before sharding, stateful workloads that resist distribution.
# StreamHub phase 1 — vertical growth
instance: t3.medium → t3.xlarge → r6g.2xlarge
database: db.t3.medium → db.r6g.4xlarge
# Worked until ~20K concurrent users
Horizontal scaling (scale out)
Horizontal scaling adds identical nodes behind a load balancer. Stateless app servers are the classic example.
AWS ALB → EKS traffic path
StreamHub's API tier went stateless early: session data moved to Redis, uploaded files to S3, so any app instance could handle any request.
# Health check endpoint every LB uses
curl -f http://app-3.internal:8080/health
# → 200 { "status": "ok", "db": "reachable", "redis": "reachable" }
Stateless vs stateful tiers
| Aspect | Stateless (easy to scale out) | Stateful (harder to scale) |
|---|---|---|
| Examples | API servers, workers | Primary DB, WebSocket sessions |
| Scaling | Add nodes behind LB | Replication, sharding, sticky sessions |
| Failure | Kill one node, traffic reroutes | Failover, leader election |
| StreamHub | Video API pods | Postgres primary, live chat rooms |
Examples
Stateless (easy to scale out)API servers, workersStateful (harder to scale)Primary DB, WebSocket sessionsScaling
Stateless (easy to scale out)Add nodes behind LBStateful (harder to scale)Replication, sharding, sticky sessionsFailure
Stateless (easy to scale out)Kill one node, traffic reroutesStateful (harder to scale)Failover, leader electionStreamHub
Stateless (easy to scale out)Video API podsStateful (harder to scale)Postgres primary, live chat rooms
Push state out of app servers into Redis, the database, or object storage. StreamHub stores upload progress in Redis so workers and APIs share status without in-memory coupling.
Database scaling path
Databases resist horizontal scaling more than app tiers. The usual progression:
Database scaling ladder
- Vertical scale — bigger instance, tuned queries, indexes.
- Read replicas — offload read-heavy workloads to followers.
- Caching — Redis for hot keys (covered in later modules).
- Sharding — partition data across multiple primaries.
- Specialised stores — search index, time-series DB for analytics.
-- StreamHub: route analytics queries to replica
-- Application config (conceptual)
SET transaction_read_only = on;
-- connection string points to replica.endpoint
SELECT date, view_count FROM daily_stats WHERE creator_id = $1;
Auto-scaling
Cloud auto-scaling groups adjust instance count from metrics:
# Example HPA-style rule for StreamHub API
scale_target: deployment/streamhub-api
metrics:
- type: cpu_utilization
target: 70%
- type: request_rate
target: 8000 per instance
min_replicas: 3
max_replicas: 50
cooldown_seconds: 120
Scale on sustained load, not spikes — flapping (rapid scale up/down) wastes money and causes cold starts.
When vertical still wins
| Component | Why vertical first |
|---|---|
| Postgres primary (pre-shard) | ACID transactions across rows need one leader |
| Redis single-shard hot cache | Simpler than cluster until memory exceeds one node |
| Early MVP | Team size 2 — ops cost of K8s cluster not justified |
| GPU inference node | Model weights do not split trivially |
Postgres primary (pre-shard)
Why vertical firstACID transactions across rows need one leaderRedis single-shard hot cache
Why vertical firstSimpler than cluster until memory exceeds one nodeEarly MVP
Why vertical firstTeam size 2 — ops cost of K8s cluster not justifiedGPU inference node
Why vertical firstModel weights do not split trivially
Cost and operational trade-offs
Horizontal scaling adds: load balancer fees, more deploy targets, log aggregation, and consistency challenges. Vertical scaling adds: bigger bills per box and downtime during resize.
StreamHub's inflection point came at roughly $8K/month on the largest single RDS instance — sharding planning started before the bill doubled again.
Quick recall
Everything you need if you only revisit this box.
- Scale up = bigger machine; scale out = more machines — most APIs scale out.
- Make app servers stateless; externalise sessions, files, and job state.
- Databases scale vertically first, then replicas, cache, then shards.
- Auto-scale on sustained metrics with cooldowns to avoid flapping.
- Vertical scaling has a ceiling and single-point-of-failure risk.
- StreamHub moved API tier horizontal early; database followed the replica ladder.
Test yourself
Answer these before moving on — recall is what makes it stick.