PrepZone Logo
PrepZone

SLI, SLO, and Error Budgets

Measure what users feel, set targets, and spend error budget on risky launches.

Why this matters

  • SRE interviews center on SLI/SLO design — not "we aim for 99.9%" without defining what is measured.
  • Error budgets gate risky releases: when ShardPay burns budget on transfer failures, feature freezes until reliability recovers.
  • SLAs with financial remedies (merchant credits) must be backed by SLOs you actually monitor — not aspirational dashboards.
  • Picking SLIs users feel separates senior answers from uptime-of-CPU metrics nobody notices.
SLI99.2% transfers succeed
SLOTarget 99.95%
SLAContract + credits
SLI is what you measure; SLO is your internal target; SLA is the customer contract with consequences.

From metric to contract

The hierarchy flows downward: define what users experience (SLI), set an internal target (SLO), then optionally contract with remedies (SLA). ShardPay's transfer path uses this chain consistently.

SLI example: (successful transfers within 200ms) / (total transfer attempts) measured over a rolling 30-day window at the API gateway — the user-facing edge, not internal shard latency.

SLO example: 99.95% of transfer attempts succeed within 200ms, measured monthly. This is ShardPay's internal target; on-call pages when burn rate threatens the monthly budget.

SLA example: Merchant enterprise contract guarantees 99.9% monthly availability with service credits of 10% MRR per 0.1% miss. The SLA is lower than the SLO — ShardPay aims for 99.95% internally to absorb variance before SLA breach.

Key points

  • SLI (Service Level Indicator) — a quantitative measure of some aspect of service behavior. ShardPay's primary transfer SLI is success rate within latency threshold, not CPU utilization or pod restart count.
  • SLO (Service Level Objective) — an internal target value for an SLI over a time window. ShardPay's transfer SLO is 99.95% monthly; fraud scoring has a separate SLO of 99.9% under 100ms.
  • SLA (Service Level Agreement) — a business contract with remedies when SLO-like thresholds are missed. ShardPay's enterprise SLA triggers credits at 99.9%; consumer merchants have best-effort terms only.
  • Error budget — the allowed bad events before SLO breach: 1 - SLO. At 99.95% SLO, ShardPay allows 0.05% failed-or-slow transfers per month — roughly 1,500 failures at current volume.

Designing ShardPay SLIs

Good SLIs are measurable, user-centric, and resistant to gaming. Bad SLIs measure internal health that users never perceive.

Transfer SLI definition

Java
SLI = count(transfers where status=SUCCESS AND latency_ms <= 200)
      / count(transfer attempts including client timeouts)

ShardPay excludes transfers where the client disconnected before response — those are counted as attempts (the server may have succeeded). This prevents the SLI from looking healthy when merchants timeout and retry.

Supporting SLIs

SLIMeasurement pointSLO
Transfer success + latencyAPI gateway99.95% / 200ms
Fraud scoring latencyFraud service ingress99.9% / 100ms
Saga convergence timeOrchestrator99.99% / 60s
Reconciliation driftNightly job0 undetected drift
  • Transfer success + latency

    Measurement pointAPI gateway
    SLO99.95% / 200ms
  • Fraud scoring latency

    Measurement pointFraud service ingress
    SLO99.9% / 100ms
  • Saga convergence time

    Measurement pointOrchestrator
    SLO99.99% / 60s
  • Reconciliation drift

    Measurement pointNightly job
    SLO0 undetected drift

ShardPay supporting SLIs

Java
// ShardPay SLI recording at gateway — simplified
public void recordTransferOutcome(TransferRequest req, TransferResponse resp,
                                 long latencyMs, Instant start) {
    boolean success = resp.status() == TransferStatus.SUCCESS;
    boolean withinSlo = latencyMs <= 200;

    sliRecorder.record("transfer", Map.of(
        "success", success,
        "within_latency_slo", success && withinSlo,
        "latency_ms", latencyMs,
        "merchant_tier", req.merchantTier()
    ));
    // Prometheus: shardpay_sli_transfer_good_total / shardpay_sli_transfer_total
}

Error budget policy

When budget remains, ShardPay ships features aggressively — including risky migrations. When budget burns, reliability work takes priority.

Walkthrough: error budget burn during shard migration

  1. ShardPay migrates merchants 50,000–60,000 to a new cell over one week.
  2. Day 3: misconfigured routing causes 0.08% transfer failures for 2 hours — burns 40% of monthly error budget in one incident.
  3. Error budget policy triggers: freeze non-critical deploys, daily reliability standup, migration rollback if burn continues.
  4. Team fixes routing, verifies SLI recovery, resumes migration at reduced batch size.
  5. Postmortem action: canary migration with SLI gate — halt if 5-minute SLI drops below 99.99%.

Key points

  • Burn rate — how fast error budget is consumed. ShardPay alerts on multi-window burn rates: 1-hour burn at 10x normal pages immediately; 6-hour burn at 5x pages during business hours.
  • Multi-window alerting — Google SRE book pattern. A brief blip burns little budget; sustained degradation burns fast. ShardPay uses 1h, 6h, and 3d windows for transfer SLO.
  • SLO vs SLA gap — internal SLO should be stricter than external SLA to provide buffer. ShardPay's 99.95% SLO vs 99.9% SLA gives ~0.05% headroom before merchant credits.
  • User-centric selection — availability SLI alone misses slow responses. ShardPay combines success AND latency into one SLI because a "successful" 5-second transfer is a failed experience for merchants.

Quick recall

Everything you need if you only revisit this box.

  1. SLI measures user-visible behavior; SLO sets the internal target; SLA is the contract.
  2. Error budget = 1 - SLO; burn rate alerts catch incidents before SLA breach.
  3. ShardPay's transfer SLI combines success and 200ms latency at the gateway.

Test yourself

Answer these before moving on — recall is what makes it stick.