PrepZone Logo
PrepZone

Vertical vs Horizontal Scaling

Scale up, scale out, and know when each approach stops working.

Vertical scaling (scale up)

Vertical scaling increases CPU, RAM, or disk on a single instance. It is simple: no code changes, no distributed-systems complexity.

Vertical scaling pros and cons

  • Pros — zero architecture change, no network partition issues, easy ops.
  • Cons — hard ceiling (largest instance size), single point of failure, expensive at the top end.
  • Best for — early StreamHub MVP, managed databases before sharding, stateful workloads that resist distribution.
Java
# StreamHub phase 1 — vertical growth
instance: t3.medium → t3.xlarge → r6g.2xlarge
database: db.t3.medium → db.r6g.4xlarge
# Worked until ~20K concurrent users

Horizontal scaling (scale out)

Horizontal scaling adds identical nodes behind a load balancer. Stateless app servers are the classic example.

AWS ALB → EKS traffic path

TLSCLIENT
UsersHTTPS requests
NETWORK
Route 53alias → ALB
NETWORK
AWS ALBtarget group
COMPUTE
EKS pod 1streamhub-api
COMPUTE
EKS pod 2streamhub-api
COMPUTE
EKS pod NHPA scaled
L7 ALB terminates TLS, health-checks targets, and fans out to stateless pods.

StreamHub's API tier went stateless early: session data moved to Redis, uploaded files to S3, so any app instance could handle any request.

Java
# Health check endpoint every LB uses
curl -f http://app-3.internal:8080/health
# → 200 { "status": "ok", "db": "reachable", "redis": "reachable" }

Stateless vs stateful tiers

AspectStateless (easy to scale out)Stateful (harder to scale)
ExamplesAPI servers, workersPrimary DB, WebSocket sessions
ScalingAdd nodes behind LBReplication, sharding, sticky sessions
FailureKill one node, traffic reroutesFailover, leader election
StreamHubVideo API podsPostgres primary, live chat rooms
  • Examples

    Stateless (easy to scale out)API servers, workers
    Stateful (harder to scale)Primary DB, WebSocket sessions
  • Scaling

    Stateless (easy to scale out)Add nodes behind LB
    Stateful (harder to scale)Replication, sharding, sticky sessions
  • Failure

    Stateless (easy to scale out)Kill one node, traffic reroutes
    Stateful (harder to scale)Failover, leader election
  • StreamHub

    Stateless (easy to scale out)Video API pods
    Stateful (harder to scale)Postgres primary, live chat rooms

Push state out of app servers into Redis, the database, or object storage. StreamHub stores upload progress in Redis so workers and APIs share status without in-memory coupling.

Database scaling path

Databases resist horizontal scaling more than app tiers. The usual progression:

Database scaling ladder

  1. Vertical scale — bigger instance, tuned queries, indexes.
  2. Read replicas — offload read-heavy workloads to followers.
  3. Caching — Redis for hot keys (covered in later modules).
  4. Sharding — partition data across multiple primaries.
  5. Specialised stores — search index, time-series DB for analytics.
Java
-- StreamHub: route analytics queries to replica
-- Application config (conceptual)
SET transaction_read_only = on;
-- connection string points to replica.endpoint
SELECT date, view_count FROM daily_stats WHERE creator_id = $1;

Auto-scaling

Cloud auto-scaling groups adjust instance count from metrics:

Java
# Example HPA-style rule for StreamHub API
scale_target: deployment/streamhub-api
metrics:
  - type: cpu_utilization
    target: 70%
  - type: request_rate
    target: 8000 per instance
min_replicas: 3
max_replicas: 50
cooldown_seconds: 120

Scale on sustained load, not spikes — flapping (rapid scale up/down) wastes money and causes cold starts.

When vertical still wins

ComponentWhy vertical first
Postgres primary (pre-shard)ACID transactions across rows need one leader
Redis single-shard hot cacheSimpler than cluster until memory exceeds one node
Early MVPTeam size 2 — ops cost of K8s cluster not justified
GPU inference nodeModel weights do not split trivially
  • Postgres primary (pre-shard)

    Why vertical firstACID transactions across rows need one leader
  • Redis single-shard hot cache

    Why vertical firstSimpler than cluster until memory exceeds one node
  • Early MVP

    Why vertical firstTeam size 2 — ops cost of K8s cluster not justified
  • GPU inference node

    Why vertical firstModel weights do not split trivially

Cost and operational trade-offs

Horizontal scaling adds: load balancer fees, more deploy targets, log aggregation, and consistency challenges. Vertical scaling adds: bigger bills per box and downtime during resize.

StreamHub's inflection point came at roughly $8K/month on the largest single RDS instance — sharding planning started before the bill doubled again.

Quick recall

Everything you need if you only revisit this box.

  • Scale up = bigger machine; scale out = more machines — most APIs scale out.
  • Make app servers stateless; externalise sessions, files, and job state.
  • Databases scale vertically first, then replicas, cache, then shards.
  • Auto-scale on sustained metrics with cooldowns to avoid flapping.
  • Vertical scaling has a ceiling and single-point-of-failure risk.
  • StreamHub moved API tier horizontal early; database followed the replica ladder.

Test yourself

Answer these before moving on — recall is what makes it stick.