Why this matters
- Average latency hides the 1% of transfers merchants remember — p99 and p999 matter more than mean.
- Saturation predicts incidents before error rates spike: a connection pool at 95% fills before timeouts appear in SLIs.
- Dashboards that fit one screen keep on-call effective; ShardPay's transfer war room has exactly four panels.
- Interviewers ask you to pick metrics for a new service — golden signals are the structured starting point.
The four golden signals
Each signal answers one question about system health. Together they cover user experience and capacity headroom.
Latency — how long requests take. ShardPay tracks p50, p95, p99, and p999 for transfers at the gateway. A rising p50 suggests systemic slowdown; a rising p999 with stable p50 suggests tail outliers (GC pauses, hot shards).
Traffic — demand on the system. ShardPay measures transfers/sec, API requests/sec, and Kafka consumer lag. Traffic spikes during Black Friday are expected; unexpected 3x traffic may indicate a retry storm or attack.
Errors — rate of failed requests. ShardPay splits HTTP 5xx (server fault), 4xx business rejections (insufficient funds), and timeout errors (client or server). Alerting on 5xx rate, not total error rate — business rejections are healthy behavior.
Saturation — how full resources are. CPU, memory, disk I/O, connection pools, thread pools, Kafka consumer group lag. ShardPay pages when any shard's JDBC pool exceeds 85% for 5 minutes — before timeouts affect the SLI.
Key points
- Latency distribution — never alert on mean alone. ShardPay's p99 transfer latency SLO is 200ms while p50 is 45ms; mean of 60ms would hide p99 violations.
- Traffic as context — errors during low traffic may be benign (deploy glitch); errors during peak traffic are merchant-visible. ShardPay normalizes error rate against traffic volume in dashboards.
- Error taxonomy — separate infra errors from business errors. A 402 Insufficient Funds is not an SLO miss; a 503 Shard Unavailable is.
- Saturation as leading indicator — pool exhaustion, disk 90%, CPU throttling predict errors 5–15 minutes ahead. ShardPay's most prevented incidents came from saturation alerts, not error alerts.
ShardPay transfer dashboard
On-call's primary dashboard has four panels — one per golden signal — scoped to the transfer path (gateway → orchestrator → shards).
Latency panel
Heatmap of p50/p99/p999 over 24 hours, broken down by shard. ShardPay overlays deploy markers to correlate latency shifts with releases. When p999 spikes on shard 7 only, the problem is shard-local (GC, hot account), not systemic.
Traffic panel
Transfers/sec globally and per cell. ShardPay runs 4 cells; traffic panel shows per-cell share to detect routing imbalance. A cell handling 40% of traffic when designed for 25% triggers rebalance investigation.
Errors panel
Stacked area chart: 5xx by service, timeout count, saga compensation rate. ShardPay alerts when 5xx rate exceeds 0.01% for 5 minutes OR when saga compensation rate doubles baseline — compensation is a leading indicator of partial failures.
Saturation panel
JDBC pool utilization per shard, orchestrator thread pool, Kafka consumer lag for settlement topic. ShardPay's Black Friday 2025 incident was caught here: shard 12 pool at 92% thirty minutes before first timeout.
// ShardPay Micrometer metrics — golden signals instrumentation
@Component
public class TransferMetrics {
private final Timer transferLatency;
private final Counter transferErrors;
private final Counter transferTotal;
private final Gauge poolUtilization;
public TransferMetrics(MeterRegistry registry, DataSource ds) {
this.transferLatency = Timer.builder("shardpay.transfer.latency")
.publishPercentiles(0.5, 0.95, 0.99, 0.999)
.register(registry);
this.transferErrors = Counter.builder("shardpay.transfer.errors")
.tag("type", "5xx").register(registry);
this.transferTotal = Counter.builder("shardpay.transfer.total")
.register(registry);
this.poolUtilization = Gauge.builder("shardpay.pool.utilization",
ds, this::measurePoolUtilization).register(registry);
}
public void recordTransfer(long latencyMs, boolean success) {
transferLatency.record(latencyMs, TimeUnit.MILLISECONDS);
transferTotal.increment();
if (!success) transferErrors.increment();
}
}
Alerting philosophy
ShardPay alerts on symptoms merchants feel (SLO burn, p99 latency) backed by golden signals for diagnosis — not on every metric threshold.
Walkthrough: saturation → error cascade
- 14:00 — saturation alert: shard 4 JDBC pool at 88%.
- 14:08 — latency alert: p99 transfer latency 280ms (SLO is 200ms).
- 14:12 — error alert: 5xx rate 0.03% on orchestrator timeouts waiting for shard 4.
- On-call scales shard 4 connection pool and adds read replica for balance queries.
- 14:25 — saturation drops to 60%, p99 returns to 120ms, errors clear.
Without the saturation alert at 14:00, on-call would have started at 14:12 with less time to prevent SLI burn.
Key points
- USE method (supplement) — Utilization, Saturation, Errors for resources. ShardPay applies USE to databases (utilization = active connections / max) alongside golden signals at the service level.
- RED method (supplement) — Rate, Errors, Duration for request-driven services. Maps directly to traffic, errors, latency — ShardPay's gateway dashboards follow RED.
- Dashboard hierarchy — L1: golden signals (on-call). L2: per-service breakdown (investigation). L3: JVM/GC/OS (root cause). ShardPay avoids jumping to L3 before checking L1.
- Cardinality control — high-cardinality labels (merchantId) on metrics explode storage. ShardPay labels by shard, cell, and merchant_tier — not individual merchant.
Quick recall
Everything you need if you only revisit this box.
- Golden signals: latency, traffic, errors, saturation — one panel each.
- Alert on tail latency (p99/p999), not averages.
- Saturation is a leading indicator — it predicts errors before SLI burn.
Test yourself
Answer these before moving on — recall is what makes it stick.