Performance: optimising the single path
Performance improvements make each request cheaper: better algorithms, fewer round trips, connection pooling, and hardware upgrades on one machine.
Performance levers
- Algorithmic efficiency — O(n log n) sort instead of O(n²).
- Database indexes — avoid full table scans on hot queries.
- Connection pooling — reuse TCP connections to the database.
- Payload size — compress JSON, paginate large lists.
- Hardware — faster CPU, NVMe SSD, more RAM on one host.
-- Performance fix: index the hot lookup
CREATE INDEX idx_videos_creator_id ON videos (creator_id, created_at DESC);
EXPLAIN ANALYZE
SELECT * FROM videos WHERE creator_id = 'u_42' ORDER BY created_at DESC LIMIT 20;
-- Index Scan instead of Seq Scan
On StreamHub at launch, a single app server and Postgres instance served video metadata in under 50 ms p99. Performance tuning alone was enough — there was no need to scale out yet.
Scalability: growing with demand
Scalability means adding capacity (usually horizontally) so throughput increases roughly linearly with resources. The goal is handling 100× traffic without 100× latency.
When StreamHub hit 50K concurrent viewers during a live event, one server's CPU pegged at 100%. Performance tweaks (gzip, query caching) bought a few percent — horizontal scaling bought 10× headroom.
The relationship between the two
| Aspect | Performance | Scalability |
|---|---|---|
| Question | How fast is one request? | How does speed change with load? |
| Typical fix | Index, cache, faster disk | More servers, sharding, queue |
| Limit | Hardware ceiling on one machine | Coordination overhead, consistency |
| Measure | p99 latency at fixed load | Throughput vs latency curve |
Question
PerformanceHow fast is one request?ScalabilityHow does speed change with load?Typical fix
PerformanceIndex, cache, faster diskScalabilityMore servers, sharding, queueLimit
PerformanceHardware ceiling on one machineScalabilityCoordination overhead, consistencyMeasure
Performancep99 latency at fixed loadScalabilityThroughput vs latency curve
Optimise performance first on a small system; invest in scalability when load outgrows one box.
Latency vs throughput
- Latency — time for one operation (ms per request).
- Throughput — operations per second (QPS).
They often trade off. Batching 100 writes into one transaction improves throughput but increases per-request latency. StreamHub's upload API batches thumbnail generation asynchronously so the user's "upload complete" response stays under 200 ms while processing continues in the background.
# StreamHub upload path — latency vs throughput split
sync_path:
- validate metadata
- store raw file reference
- return 202 Accepted # low latency for user
async_path:
- transcode video variants # high throughput via worker pool
- generate thumbnails
- update search index
When performance tuning stops working
| Signal | Likely bottleneck | Scalability response |
|---|---|---|
| CPU at 100% on app tier | Compute-bound handlers | Horizontal app scaling + LB |
| DB connections exhausted | Connection pool saturation | Read replicas, PgBouncer |
| Disk I/O wait high | Storage throughput | SSD tier, read replicas, cache |
| Single thread maxed | Vertical limit on one core | Partition work across nodes |
CPU at 100% on app tier
Likely bottleneckCompute-bound handlersScalability responseHorizontal app scaling + LBDB connections exhausted
Likely bottleneckConnection pool saturationScalability responseRead replicas, PgBouncerDisk I/O wait high
Likely bottleneckStorage throughputScalability responseSSD tier, read replicas, cacheSingle thread maxed
Likely bottleneckVertical limit on one coreScalability responsePartition work across nodes
Measuring both dimensions
# Quick load test sketch (not a full benchmark)
wrk -t4 -c100 -d30s https://api.streamhub.example/v1/feed
# Watch: requests/sec (throughput) AND latency distribution (p50, p99)
A system is scalable if doubling servers roughly doubles sustainable QPS at the same p99. If p99 doubles when QPS doubles, you have a scalability problem, not just a performance tuning gap.
StreamHub early lesson
At 500 DAU, StreamHub ran on a $40/month VPS. At 500K DAU, the same code on bigger hardware still failed — not because the code was slow, but because one Postgres primary could not serve 80K read QPS. The fix was read replicas and Redis, not a faster CPU.
Quick recall
Everything you need if you only revisit this box.
- Performance = speed of one request; scalability = behaviour as load grows.
- Tune indexes, pooling, and payloads first on a small system.
- Horizontal scaling, caching, and async processing address scalability limits.
- Latency and throughput often trade off — separate sync user paths from batch work.
- Profile to find the real bottleneck before adding infrastructure.
- StreamHub outgrew vertical scaling when read QPS exceeded one database's capacity.
Test yourself
Answer these before moving on — recall is what makes it stick.