PrepZone Logo
PrepZone

Capacity Planning and Headroom

Never run production at 100% — spikes, failover capacity, and ShardPay's Black Friday playbook.

Why this matters

  • Capacity planning prevents heroic firefighting during sales events.
  • Headroom covers failover when one AZ goes dark.
  • FinOps and reliability intersect at utilization targets.
  • Autoscaling is not instant — headroom bridges the gap between signal and new capacity.
Utilization over time
Steady state~60% CPU / connections
Peak traffic~85% — within headroom
Failover burstExtra load on survivors
Run below 70% steady-state so spikes and failovers have room without paging on-call.

Plan for spikes and failover

Capacity planning answers: how much traffic can we handle, and what happens when we exceed it? ShardPay models three scenarios: steady state (60% utilization), peak event (Black Friday at 2.5× steady), and failure (one AZ down, losing 33% of capacity). The system must survive peak-on-failure without breaching SLO.

Black Friday doubled ShardPay traffic to 16,000 transfers/sec. With 40% headroom and pre-warmed cells, p99 stayed at 185ms. The previous year without headroom, queue depth exploded per Little's Law and p99 hit 2 seconds — customers abandoned transfers mid-flow.

Capacity concepts

  • Headroom — unused capacity buffer above steady-state load, typically 30–50%. ShardPay targets 40% headroom on the API tier and 30% on ledger shards so a traffic spike or slow deploy does not immediately breach SLO.
  • N+1 redundancy — capacity to absorb the loss of one node (or one AZ) without dropping below required throughput. ShardPay runs 12 API pods with capacity for 11; losing one during a deploy leaves 92% capacity.
  • Load testing — prove capacity before the event, not during. ShardPay runs monthly load tests at 3× steady state in a staging cell, measuring p99 latency and error rate at each multiplier.
  • Autoscaling lag — new instances take 2–5 minutes to provision, pull images, pass health checks, and warm JVM caches. ShardPay pre-warms 50% extra capacity one hour before known events; autoscaling handles the rest.

Steady-state utilization targets

Running at 100% utilization means zero margin for error. A single slow query, a GC pause, or a deploy rolling one pod out of service pushes the system into queueing and latency explosion. ShardPay's target is 60% steady-state CPU on API pods and 70% on ledger shards (which are I/O-bound and tolerate higher CPU).

FinOps teams want high utilization to reduce waste. Reliability teams want low utilization for headroom. The compromise is tiered targets: stateless tiers at 60% (cheap to scale out), stateful tiers at 70% (expensive to rebalance), and burst capacity via pre-warming for known events.

Walkthrough: Black Friday capacity plan

ShardPay's steady state is 6,000 transfers/sec. Black Friday forecast: 15,000 transfers/sec (2.5×). One AZ failure during the event: effective capacity drops to 67% of fleet.

Planning math: need capacity for 15,000/sec at p99 < 200ms. Load test showed each ledger shard handles 2,800 writes/sec at p99 < 150ms. Need 6 shards × 2,800 = 16,800 capacity. With one AZ down (lose 2 shards): 4 × 2,800 = 11,200 — insufficient. Solution: pre-warm 8 shards (22,400 capacity), lose 2 → 16,800, still above 15,000 with 12% headroom.

Java
// ShardPay capacity controller — pre-warm before known events
@Component
public class CapacityPreWarmer {
    private final KubernetesClient k8s;
    private final LoadTestClient loadTest;

    @Scheduled(cron = "0 0 4 * * NOV")  // 4 AM on Black Friday
    public void preWarmBlackFriday() {
        int targetPods = capacityPlan.requiredPods(PEAK_MULTIPLIER);
        int currentPods = k8s.countPods("api-tier");

        if (currentPods < targetPods) {
            k8s.scale("api-tier", targetPods);
            loadTest.warmUp(Duration.ofMinutes(10));  // JIT, connection pools, caches
            log.info("Pre-warmed API from {} to {} pods", currentPods, targetPods);
        }
    }
}

Measuring and forecasting capacity

Capacity is not a one-time calculation. ShardPay tracks weekly growth rate, models seasonal curves, and revises forecasts monthly. Key metrics: requests per second at p99 SLO, queue depth trends (Little's Law early warning), and per-shard utilization on the ledger.

Load tests must be realistic: same payload sizes, same fan-out patterns, same idempotency key distribution. ShardPay's first load test used uniform account IDs and missed hot-shard effects — 80% of test traffic hit one shard because test accounts were sequential, not hashed.

Capacity signals

  • Queue depth trending up with flat CPU — approaching I/O or lock bottleneck, not compute limit. ShardPay adds disk IOPS before adding CPU when this pattern appears.
  • p99 latency rising before error rate — the canary signal. Latency degrades before failures start; alert on p99, not just 5xx rate.
  • Autoscaling events per hour — frequent scale-up/down indicates insufficient headroom or too-aggressive thresholds. ShardPay aims for fewer than two scaling events per day in steady state.
  • Error budget burn rate — if capacity is tight, even small spikes burn SLO budget. ShardPay correlates budget burn with utilization to justify headroom spending to FinOps.

Headroom costs money — spend it wisely

Headroom is insurance. Like all insurance, it has a premium. ShardPay spends ~$40k/month on idle capacity across API and ledger tiers. The ROI: zero SLO breaches during Black Friday and sub-10-minute recovery when an AZ fails. The alternative — running at 95% and firefighting during peaks — cost $200k in lost transfers and merchant credits one year.

The art is right-sizing per tier: cheap stateless pods get more headroom (scale in minutes); expensive stateful shards get less (rebalance takes hours). Known events get pre-warming; unknown spikes rely on autoscaling with acknowledged lag.

Quick recall

Everything you need if you only revisit this box.

  1. Target 60% steady-state utilization; 100% means any spike becomes an outage.
  2. Headroom covers spikes, failover, and autoscaling lag — it costs money but prevents SLO breaches.
  3. Load test at 3× steady state before peak events; pre-warm for known spikes.

Test yourself

Answer these before moving on — recall is what makes it stick.