PrepZone Logo
PrepZone

Performance vs Scalability

Why a fast system on one machine is not the same as a system that survives ten million users.

Performance: optimising the single path

Performance improvements make each request cheaper: better algorithms, fewer round trips, connection pooling, and hardware upgrades on one machine.

Performance levers

  • Algorithmic efficiency — O(n log n) sort instead of O(n²).
  • Database indexes — avoid full table scans on hot queries.
  • Connection pooling — reuse TCP connections to the database.
  • Payload size — compress JSON, paginate large lists.
  • Hardware — faster CPU, NVMe SSD, more RAM on one host.
Java
-- Performance fix: index the hot lookup
CREATE INDEX idx_videos_creator_id ON videos (creator_id, created_at DESC);

EXPLAIN ANALYZE
SELECT * FROM videos WHERE creator_id = 'u_42' ORDER BY created_at DESC LIMIT 20;
-- Index Scan instead of Seq Scan

On StreamHub at launch, a single app server and Postgres instance served video metadata in under 50 ms p99. Performance tuning alone was enough — there was no need to scale out yet.

Scalability: growing with demand

Scalability means adding capacity (usually horizontally) so throughput increases roughly linearly with resources. The goal is handling 100× traffic without 100× latency.

Stage 1EC2 + Postgres on one box
Stage 2RDS + separate EC2 app
Stage 3ALB + ASG + ElastiCache Redis
Stage 4CloudFront + S3 media origin
Stage 5MSK Kafka + EKS worker fleet
Stage 6Sharded RDS + multi-region EKS
Each stage adds one AWS building block when the previous bottleneck appears.

When StreamHub hit 50K concurrent viewers during a live event, one server's CPU pegged at 100%. Performance tweaks (gzip, query caching) bought a few percent — horizontal scaling bought 10× headroom.

The relationship between the two

AspectPerformanceScalability
QuestionHow fast is one request?How does speed change with load?
Typical fixIndex, cache, faster diskMore servers, sharding, queue
LimitHardware ceiling on one machineCoordination overhead, consistency
Measurep99 latency at fixed loadThroughput vs latency curve
  • Question

    PerformanceHow fast is one request?
    ScalabilityHow does speed change with load?
  • Typical fix

    PerformanceIndex, cache, faster disk
    ScalabilityMore servers, sharding, queue
  • Limit

    PerformanceHardware ceiling on one machine
    ScalabilityCoordination overhead, consistency
  • Measure

    Performancep99 latency at fixed load
    ScalabilityThroughput vs latency curve

Optimise performance first on a small system; invest in scalability when load outgrows one box.

Latency vs throughput

  • Latency — time for one operation (ms per request).
  • Throughput — operations per second (QPS).

They often trade off. Batching 100 writes into one transaction improves throughput but increases per-request latency. StreamHub's upload API batches thumbnail generation asynchronously so the user's "upload complete" response stays under 200 ms while processing continues in the background.

Java
# StreamHub upload path — latency vs throughput split
sync_path:
  - validate metadata
  - store raw file reference
  - return 202 Accepted          # low latency for user
async_path:
  - transcode video variants     # high throughput via worker pool
  - generate thumbnails
  - update search index

When performance tuning stops working

SignalLikely bottleneckScalability response
CPU at 100% on app tierCompute-bound handlersHorizontal app scaling + LB
DB connections exhaustedConnection pool saturationRead replicas, PgBouncer
Disk I/O wait highStorage throughputSSD tier, read replicas, cache
Single thread maxedVertical limit on one corePartition work across nodes
  • CPU at 100% on app tier

    Likely bottleneckCompute-bound handlers
    Scalability responseHorizontal app scaling + LB
  • DB connections exhausted

    Likely bottleneckConnection pool saturation
    Scalability responseRead replicas, PgBouncer
  • Disk I/O wait high

    Likely bottleneckStorage throughput
    Scalability responseSSD tier, read replicas, cache
  • Single thread maxed

    Likely bottleneckVertical limit on one core
    Scalability responsePartition work across nodes

Measuring both dimensions

Java
# Quick load test sketch (not a full benchmark)
wrk -t4 -c100 -d30s https://api.streamhub.example/v1/feed

# Watch: requests/sec (throughput) AND latency distribution (p50, p99)

A system is scalable if doubling servers roughly doubles sustainable QPS at the same p99. If p99 doubles when QPS doubles, you have a scalability problem, not just a performance tuning gap.

StreamHub early lesson

At 500 DAU, StreamHub ran on a $40/month VPS. At 500K DAU, the same code on bigger hardware still failed — not because the code was slow, but because one Postgres primary could not serve 80K read QPS. The fix was read replicas and Redis, not a faster CPU.

Quick recall

Everything you need if you only revisit this box.

  • Performance = speed of one request; scalability = behaviour as load grows.
  • Tune indexes, pooling, and payloads first on a small system.
  • Horizontal scaling, caching, and async processing address scalability limits.
  • Latency and throughput often trade off — separate sync user paths from batch work.
  • Profile to find the real bottleneck before adding infrastructure.
  • StreamHub outgrew vertical scaling when read QPS exceeded one database's capacity.

Test yourself

Answer these before moving on — recall is what makes it stick.