PrepZone Logo
PrepZone

Latency, Throughput, and Little's Law

Why optimizing average latency hides tail problems — and how L = λW governs ShardPay queue depth.

Why this matters

  • Average latency hides tail problems that lose customers.
  • Throughput without latency budget creates hidden queues.
  • Interviewers ask you to find the bottleneck in a pipeline.
  • Little's Law lets you predict latency from queue depth without profiling every thread.
Gateway20ms
Auth15ms
Ledger80ms
Fraud40ms
Response200ms total
Each hop consumes part of the SLA. Sum of hop timeouts must stay under the total budget.

Latency vs throughput

Latency is the time for one operation to complete — the customer's wait from tap "Send" to confirmation. Throughput is operations completed per unit time — transfers per second across the fleet. They are related but not interchangeable: you can raise throughput by batching 100 transfers into one disk write, but each individual transfer waits longer in the batch queue.

ShardPay caps batch size at 50 transfers so p99 stays under 200ms. Without the cap, throughput graphs looked healthy while customers abandoned slow confirmations. Optimizing one metric in isolation often harms the other.

Key metrics

  • Latency — end-to-end time for a single operation, usually tracked at p50, p99, and p999. ShardPay's transfer SLO is p99 < 200ms; the mean can be 45ms while p99 is 180ms — the tail is what users remember.
  • Throughput — completed operations per second across the system. ShardPay's ledger sustains 8,000 transfers/sec at peak; the API tier sustains 25,000 req/sec because reads outnumber writes 5:1.
  • Little's Law — L = λ × W: average items in the system equals arrival rate times average wait time. If ShardPay's transfer queue holds 400 items (L) at 2,000 arrivals/sec (λ), average wait is 200ms (W = L/λ).
  • Tail latency — high percentiles (p99, p999) matter more than mean at scale. One slow shard in a fan-out query sets the tail for the entire request; ShardPay monitors per-shard p99 separately.

Little's Law in practice

Little's Law is deceptively simple but powerful. You do not need to know queue discipline or service time distribution — if you measure any two of L, λ, and W, you get the third. During a Black Friday spike, ShardPay's transfer queue depth jumped from 50 to 800. Arrival rate was 3,500/sec. Little's Law predicted average wait: 800 / 3500 ≈ 229ms — matching the observed p50 latency spike without any code profiling.

The law also explains why adding capacity does not always help. If the bottleneck is a single leader shard, adding API pods increases λ to the shard without reducing W. The queue grows on the shard, not the API tier.

Walkthrough: finding the hidden queue

ShardPay's dashboard showed API CPU at 40% and ledger CPU at 55% — both looked healthy. But transfer p99 hit 450ms. Little's Law on the ledger write queue: L = 900, λ = 2,000/sec → W = 450ms. The bottleneck was disk fsync on the leader, not CPU. Adding API instances made it worse by increasing λ. The fix was faster disks and write batching, not more pods.

Java
// ShardPay metrics — expose queue depth for Little's Law monitoring
@Component
public class TransferQueueMetrics {
    private final AtomicInteger queueDepth = new AtomicInteger();

    public void enqueue(Transfer transfer) {
        queueDepth.incrementAndGet();
        ledgerPool.submit(() -> {
            process(transfer);
            queueDepth.decrementAndGet();
        });
    }

    @Scheduled(fixedRate = 10_000)
    public void reportQueueStats() {
        int depth = queueDepth.get();
        double arrivalRate = rateCounter.transfersPerSecond();  // λ
        if (arrivalRate > 0) {
            double predictedWaitMs = (depth / arrivalRate) * 1000;  // W = L/λ
            metrics.gauge("transfer.queue.predicted_wait_ms", predictedWaitMs);
            metrics.gauge("transfer.queue.depth", depth);
        }
    }
}

Tail latency and the fan-out problem

When a request fans out to N shards or services, the overall latency is dominated by the slowest responder, not the average. With 10 parallel calls each at p99 = 50ms, the probability that at least one exceeds 50ms approaches certainty as N grows. ShardPay limits fan-out to three shards per transfer and uses deadline propagation so a slow shard does not consume the entire budget.

Coordinated omission is another tail trap: if the client times out at 200ms, it stops recording latency for requests that would have taken 500ms. Your p99 looks artificially low. ShardPay's load generator records "timeout" as 200ms+ in histograms to avoid lying to itself.

Optimizing the bottleneck

Improving non-bottleneck components wastes effort. ShardPay's transfer path is: gateway (15ms) → auth (20ms) → ledger (80ms) → fraud (40ms) → response. Optimizing auth from 20ms to 10ms saves 10ms; optimizing ledger from 80ms to 50ms saves 30ms. The ledger is the bottleneck — every optimization dollar goes there first.

Bottleneck signals

  • Rising queue depth with flat CPU — the bottleneck is I/O, lock contention, or downstream dependency, not compute. ShardPay saw this when WAL fsync saturated disk IOPS before CPU.
  • Flat throughput as load increases — the system hit a ceiling. Adding load only increases W (latency), not λ (throughput). ShardPay's ledger plateaus at ~8k writes/sec per shard regardless of client count.
  • Tail grows faster than mean — contention or GC pauses affect p999 disproportionately. ShardPay's p999/p50 ratio above 10× triggers investigation even if the mean looks fine.

Quick recall

Everything you need if you only revisit this box.

  1. Little's Law: L = λW — queue depth predicts latency from arrival rate.
  2. Optimize the bottleneck hop, not the average.
  3. Tail latency (p99/p999) drives SLOs; fan-out amplifies tails.

Test yourself

Answer these before moving on — recall is what makes it stick.