PrepZone Logo
PrepZone

Vertical vs Horizontal Scaling

Scale up until you cannot, then scale out — stateless tiers first, stateful tiers with a plan.

Why this matters

  • Every scaling discussion starts with up vs out.
  • Stateful tiers need partitioning before horizontal scale works.
  • ShardPay maxed vertical scale at 32 cores before sharding.
  • Choosing the wrong scaling direction wastes months and money.
Vertical
Bigger machineMore CPU / RAM
Single point of failure
Horizontal
More instancesLoad balanced
Fault tolerantN+1 capacity
Scale up hits hardware limits; scale out needs stateless or partitioned state.

Scale up then scale out

Vertical scaling (scale up) means bigger hardware: more CPU, RAM, and faster disks on one machine. Horizontal scaling (scale out) means more machines behind a load balancer. Vertical is simpler — no distributed state, no partition logic, no consensus overhead. Horizontal is more resilient — lose one node and the rest absorb traffic.

ShardPay's journey followed the standard path: vertical first until hardware limits, then horizontal with partitioning. The API tier scaled vertically from 8 to 32 cores, then horizontally to 40 pods. The ledger hit disk I/O limits at 32 cores and needed hash partitioning across six shards before horizontal scale delivered linear throughput gains.

Scaling dimensions

  • Vertical scaling — bigger machine, simpler ops, hard ceiling at the largest available instance. ShardPay's monolith ran on a 32-vCPU / 128GB instance before I/O saturation forced a different approach.
  • Horizontal scaling — more nodes, requires load balancing, shared-nothing design, and often data partitioning. ShardPay's API tier scales horizontally today with zero code changes because it is stateless.
  • N+1 capacity — provision one extra node's worth of headroom so losing a single instance does not drop below required throughput. ShardPay runs 12 API pods with capacity for 11 — one can fail during deploy without SLO impact.
  • Elastic scaling — horizontal scale tied to autoscaling signals (CPU, queue depth, request rate). ShardPay's API autoscales on p95 latency; ledger shards scale manually because rebalancing data is not instantaneous.

When vertical scaling wins

For stateless or lightly stateful workloads, vertical scaling is often the right first move. ShardPay's API tier went from 8 to 32 cores and handled 4× traffic with one config change and zero architectural rework. No service discovery, no consistent hashing, no split-brain risk. The operational cost is a bigger AWS bill, not a bigger engineering team.

Vertical scaling also avoids the coordination overhead of distribution. A single Postgres instance on a large machine can outperform three smaller instances with synchronous replication — until the single machine fails or hits disk limits.

Example: API tier vertical path

ShardPay's API handled 6,000 req/sec on 8 cores with p99 at 90ms. Traffic grew to 15,000 req/sec. Vertical scale to 16 cores: p99 stayed at 95ms, no architecture change. At 25,000 req/sec on 32 cores, p99 hit 140ms — diminishing returns. Horizontal scale to 20 pods at 16 cores each: p99 returned to 85ms with N+1 redundancy. The vertical phase bought two years of growth without microservices complexity.

Java
// ShardPay shard router — horizontal scaling requires partition logic
public class ShardRouter {
    private final List<LedgerShard> shards;
    private final ConsistentHashRing ring;

    public LedgerShard route(AccountId accountId) {
        int shardIndex = ring.getShard(accountId.hashCode());
        return shards.get(shardIndex);
    }

    // Vertical scaling: this class doesn't exist — one LedgerService handles all accounts
    // Horizontal scaling: every transfer must route to the correct shard
}

When horizontal scaling is required

Stateful tiers hit vertical ceilings that money cannot fix. ShardPay's ledger saturated disk IOPS at 32 cores — more CPU did not help because the bottleneck was fsync, not compute. Horizontal sharding split accounts across six nodes, each with its own disk, multiplying I/O capacity linearly.

Horizontal scaling also provides fault isolation. A bug in one shard affects only accounts on that shard, not the entire ledger. ShardPay's blast radius per shard is ~16% of accounts — a trade-off accepted for resilience and throughput.

The scaling decision matrix

ShardPay uses a simple matrix: if the tier is stateless and under vertical limits, scale up. If stateless and past vertical limits, scale out. If stateful, partition first, then scale out shards. If partitioning is not yet justified, scale up the stateful tier and accept the single-point-of-failure risk temporarily.

ShardPay scaling history

  • API gateway — vertical to 32 cores, then horizontal to 40 pods. Stateless; scales elastically on latency signals.
  • Ledger — vertical to 32 cores (I/O wall), then horizontal sharding to 6 shards. Stateful; each shard is a Raft group with 3 replicas.
  • Fraud scoring — horizontal from day one on GPU instances. Stateless inference; models loaded from object storage.
  • Notification worker — horizontal with queue-based consumption. Stateless; Kafka consumer group handles partition assignment.

Quick recall

Everything you need if you only revisit this box.

  1. Vertical is simpler until hardware limits; horizontal needs stateless or partitioned design.
  2. ShardPay scaled API vertically first, then horizontally; ledger needed sharding.
  3. N+1 headroom means losing one node does not breach SLO.

Test yourself

Answer these before moving on — recall is what makes it stick.