Why this matters
- SRE interviews center on SLI/SLO design — not "we aim for 99.9%" without defining what is measured.
- Error budgets gate risky releases: when ShardPay burns budget on transfer failures, feature freezes until reliability recovers.
- SLAs with financial remedies (merchant credits) must be backed by SLOs you actually monitor — not aspirational dashboards.
- Picking SLIs users feel separates senior answers from uptime-of-CPU metrics nobody notices.
From metric to contract
The hierarchy flows downward: define what users experience (SLI), set an internal target (SLO), then optionally contract with remedies (SLA). ShardPay's transfer path uses this chain consistently.
SLI example: (successful transfers within 200ms) / (total transfer attempts) measured over a rolling 30-day window at the API gateway — the user-facing edge, not internal shard latency.
SLO example: 99.95% of transfer attempts succeed within 200ms, measured monthly. This is ShardPay's internal target; on-call pages when burn rate threatens the monthly budget.
SLA example: Merchant enterprise contract guarantees 99.9% monthly availability with service credits of 10% MRR per 0.1% miss. The SLA is lower than the SLO — ShardPay aims for 99.95% internally to absorb variance before SLA breach.
Key points
- SLI (Service Level Indicator) — a quantitative measure of some aspect of service behavior. ShardPay's primary transfer SLI is success rate within latency threshold, not CPU utilization or pod restart count.
- SLO (Service Level Objective) — an internal target value for an SLI over a time window. ShardPay's transfer SLO is 99.95% monthly; fraud scoring has a separate SLO of 99.9% under 100ms.
- SLA (Service Level Agreement) — a business contract with remedies when SLO-like thresholds are missed. ShardPay's enterprise SLA triggers credits at 99.9%; consumer merchants have best-effort terms only.
- Error budget — the allowed bad events before SLO breach:
1 - SLO. At 99.95% SLO, ShardPay allows 0.05% failed-or-slow transfers per month — roughly 1,500 failures at current volume.
Designing ShardPay SLIs
Good SLIs are measurable, user-centric, and resistant to gaming. Bad SLIs measure internal health that users never perceive.
Transfer SLI definition
SLI = count(transfers where status=SUCCESS AND latency_ms <= 200)
/ count(transfer attempts including client timeouts)
ShardPay excludes transfers where the client disconnected before response — those are counted as attempts (the server may have succeeded). This prevents the SLI from looking healthy when merchants timeout and retry.
Supporting SLIs
| SLI | Measurement point | SLO |
|---|---|---|
| Transfer success + latency | API gateway | 99.95% / 200ms |
| Fraud scoring latency | Fraud service ingress | 99.9% / 100ms |
| Saga convergence time | Orchestrator | 99.99% / 60s |
| Reconciliation drift | Nightly job | 0 undetected drift |
Transfer success + latency
Measurement pointAPI gatewaySLO99.95% / 200msFraud scoring latency
Measurement pointFraud service ingressSLO99.9% / 100msSaga convergence time
Measurement pointOrchestratorSLO99.99% / 60sReconciliation drift
Measurement pointNightly jobSLO0 undetected drift
ShardPay supporting SLIs
// ShardPay SLI recording at gateway — simplified
public void recordTransferOutcome(TransferRequest req, TransferResponse resp,
long latencyMs, Instant start) {
boolean success = resp.status() == TransferStatus.SUCCESS;
boolean withinSlo = latencyMs <= 200;
sliRecorder.record("transfer", Map.of(
"success", success,
"within_latency_slo", success && withinSlo,
"latency_ms", latencyMs,
"merchant_tier", req.merchantTier()
));
// Prometheus: shardpay_sli_transfer_good_total / shardpay_sli_transfer_total
}
Error budget policy
When budget remains, ShardPay ships features aggressively — including risky migrations. When budget burns, reliability work takes priority.
Walkthrough: error budget burn during shard migration
- ShardPay migrates merchants 50,000–60,000 to a new cell over one week.
- Day 3: misconfigured routing causes 0.08% transfer failures for 2 hours — burns 40% of monthly error budget in one incident.
- Error budget policy triggers: freeze non-critical deploys, daily reliability standup, migration rollback if burn continues.
- Team fixes routing, verifies SLI recovery, resumes migration at reduced batch size.
- Postmortem action: canary migration with SLI gate — halt if 5-minute SLI drops below 99.99%.
Key points
- Burn rate — how fast error budget is consumed. ShardPay alerts on multi-window burn rates: 1-hour burn at 10x normal pages immediately; 6-hour burn at 5x pages during business hours.
- Multi-window alerting — Google SRE book pattern. A brief blip burns little budget; sustained degradation burns fast. ShardPay uses 1h, 6h, and 3d windows for transfer SLO.
- SLO vs SLA gap — internal SLO should be stricter than external SLA to provide buffer. ShardPay's 99.95% SLO vs 99.9% SLA gives ~0.05% headroom before merchant credits.
- User-centric selection — availability SLI alone misses slow responses. ShardPay combines success AND latency into one SLI because a "successful" 5-second transfer is a failed experience for merchants.
Quick recall
Everything you need if you only revisit this box.
- SLI measures user-visible behavior; SLO sets the internal target; SLA is the contract.
- Error budget = 1 - SLO; burn rate alerts catch incidents before SLA breach.
- ShardPay's transfer SLI combines success and 200ms latency at the gateway.
Test yourself
Answer these before moving on — recall is what makes it stick.