PrepZone Logo
PrepZone

Golden Signals and Tail Latency

Latency, traffic, errors, saturation — why p99 matters more than average for ShardPay.

Why this matters

  • Average latency hides the 1% of transfers merchants remember — p99 and p999 matter more than mean.
  • Saturation predicts incidents before error rates spike: a connection pool at 95% fills before timeouts appear in SLIs.
  • Dashboards that fit one screen keep on-call effective; ShardPay's transfer war room has exactly four panels.
  • Interviewers ask you to pick metrics for a new service — golden signals are the structured starting point.
Latencyp50 / p99
TrafficRequests/sec
Errors5xx rate
SaturationCPU, pool, queue
Latency, traffic, errors, and saturation — the four signals Google SRE recommends for every service.

The four golden signals

Each signal answers one question about system health. Together they cover user experience and capacity headroom.

Latency — how long requests take. ShardPay tracks p50, p95, p99, and p999 for transfers at the gateway. A rising p50 suggests systemic slowdown; a rising p999 with stable p50 suggests tail outliers (GC pauses, hot shards).

Traffic — demand on the system. ShardPay measures transfers/sec, API requests/sec, and Kafka consumer lag. Traffic spikes during Black Friday are expected; unexpected 3x traffic may indicate a retry storm or attack.

Errors — rate of failed requests. ShardPay splits HTTP 5xx (server fault), 4xx business rejections (insufficient funds), and timeout errors (client or server). Alerting on 5xx rate, not total error rate — business rejections are healthy behavior.

Saturation — how full resources are. CPU, memory, disk I/O, connection pools, thread pools, Kafka consumer group lag. ShardPay pages when any shard's JDBC pool exceeds 85% for 5 minutes — before timeouts affect the SLI.

Key points

  • Latency distribution — never alert on mean alone. ShardPay's p99 transfer latency SLO is 200ms while p50 is 45ms; mean of 60ms would hide p99 violations.
  • Traffic as context — errors during low traffic may be benign (deploy glitch); errors during peak traffic are merchant-visible. ShardPay normalizes error rate against traffic volume in dashboards.
  • Error taxonomy — separate infra errors from business errors. A 402 Insufficient Funds is not an SLO miss; a 503 Shard Unavailable is.
  • Saturation as leading indicator — pool exhaustion, disk 90%, CPU throttling predict errors 5–15 minutes ahead. ShardPay's most prevented incidents came from saturation alerts, not error alerts.

ShardPay transfer dashboard

On-call's primary dashboard has four panels — one per golden signal — scoped to the transfer path (gateway → orchestrator → shards).

Latency panel

Heatmap of p50/p99/p999 over 24 hours, broken down by shard. ShardPay overlays deploy markers to correlate latency shifts with releases. When p999 spikes on shard 7 only, the problem is shard-local (GC, hot account), not systemic.

Traffic panel

Transfers/sec globally and per cell. ShardPay runs 4 cells; traffic panel shows per-cell share to detect routing imbalance. A cell handling 40% of traffic when designed for 25% triggers rebalance investigation.

Errors panel

Stacked area chart: 5xx by service, timeout count, saga compensation rate. ShardPay alerts when 5xx rate exceeds 0.01% for 5 minutes OR when saga compensation rate doubles baseline — compensation is a leading indicator of partial failures.

Saturation panel

JDBC pool utilization per shard, orchestrator thread pool, Kafka consumer lag for settlement topic. ShardPay's Black Friday 2025 incident was caught here: shard 12 pool at 92% thirty minutes before first timeout.

Java
// ShardPay Micrometer metrics — golden signals instrumentation
@Component
public class TransferMetrics {
    private final Timer transferLatency;
    private final Counter transferErrors;
    private final Counter transferTotal;
    private final Gauge poolUtilization;

    public TransferMetrics(MeterRegistry registry, DataSource ds) {
        this.transferLatency = Timer.builder("shardpay.transfer.latency")
            .publishPercentiles(0.5, 0.95, 0.99, 0.999)
            .register(registry);
        this.transferErrors = Counter.builder("shardpay.transfer.errors")
            .tag("type", "5xx").register(registry);
        this.transferTotal = Counter.builder("shardpay.transfer.total")
            .register(registry);
        this.poolUtilization = Gauge.builder("shardpay.pool.utilization",
            ds, this::measurePoolUtilization).register(registry);
    }

    public void recordTransfer(long latencyMs, boolean success) {
        transferLatency.record(latencyMs, TimeUnit.MILLISECONDS);
        transferTotal.increment();
        if (!success) transferErrors.increment();
    }
}

Alerting philosophy

ShardPay alerts on symptoms merchants feel (SLO burn, p99 latency) backed by golden signals for diagnosis — not on every metric threshold.

Walkthrough: saturation → error cascade

  1. 14:00 — saturation alert: shard 4 JDBC pool at 88%.
  2. 14:08 — latency alert: p99 transfer latency 280ms (SLO is 200ms).
  3. 14:12 — error alert: 5xx rate 0.03% on orchestrator timeouts waiting for shard 4.
  4. On-call scales shard 4 connection pool and adds read replica for balance queries.
  5. 14:25 — saturation drops to 60%, p99 returns to 120ms, errors clear.

Without the saturation alert at 14:00, on-call would have started at 14:12 with less time to prevent SLI burn.

Key points

  • USE method (supplement) — Utilization, Saturation, Errors for resources. ShardPay applies USE to databases (utilization = active connections / max) alongside golden signals at the service level.
  • RED method (supplement) — Rate, Errors, Duration for request-driven services. Maps directly to traffic, errors, latency — ShardPay's gateway dashboards follow RED.
  • Dashboard hierarchy — L1: golden signals (on-call). L2: per-service breakdown (investigation). L3: JVM/GC/OS (root cause). ShardPay avoids jumping to L3 before checking L1.
  • Cardinality control — high-cardinality labels (merchantId) on metrics explode storage. ShardPay labels by shard, cell, and merchant_tier — not individual merchant.

Quick recall

Everything you need if you only revisit this box.

  1. Golden signals: latency, traffic, errors, saturation — one panel each.
  2. Alert on tail latency (p99/p999), not averages.
  3. Saturation is a leading indicator — it predicts errors before SLI burn.

Test yourself

Answer these before moving on — recall is what makes it stick.