PrepZone Logo
PrepZone

SLOs, SLIs, and Error Budgets

Define reliability targets, measure SLIs, and balance feature velocity with stability.

Why this matters

  • This topic directly affects how reliably BookStore reaches production — slos and error budgets balance bookstore feature velocity against reliability targets..
  • Interviewers connect hands-on commands and manifests to real delivery stories, not buzzwords.
  • Later modules assume you can explain both the why and the concrete file or command involved.
  • Platform maturity shows up when teams automate this instead of relying on tribal knowledge.
LogsStructured JSON, Loki/ELK
MetricsCounters, gauges, Prometheus
TracesSpans, OpenTelemetry

All three correlate via trace ID and service labels

Logs answer what happened. Metrics answer how much. Traces answer where time was spent.

Define SLI first

What users feel: BookStore engineers treat this as part of the standard path from laptop to slos slis error budgets readiness. Document decisions in the team runbook so on-call knows which knobs exist.

Key ideas

  • Define SLI first — primary idea for slos-slis-error-budgets
  • BookStore context — catalog API, checkout, and inventory services share the same pattern
  • Automation — prefer pipeline jobs over manual SSH steps
  • Verification — staging must prove the change before prod traffic

Error budget policy

Freeze features when burned: When staging matches production architecture, BookStore catches misconfigurations early. Pair this section's practice with observability dashboards to confirm behavior under load.

Key ideas

  • Error budget policy — operational detail
  • Rollback — know how to revert without rebuilding artefacts
  • Security — least privilege for deploy roles
  • Documentation — link runbooks from the service README

Production checklist

Before promoting BookStore changes tied to this topic, run automated tests, inspect artefact immutability (image digest or JAR checksum), execute a staging smoke test on /actuator/health, and watch error-rate dashboards for thirty minutes after prod rollout.

Key ideas

  • Staging soak — validate under synthetic load
  • Change ticket — attach pipeline URL and artefact digest
  • On-call — page owner stays on dashboards during rollout
  • Post-deploy — record metrics baseline for comparison
Java
SLI: successful checkout API responses / total requests
SLO: 99.9% over 30 days
Error budget: 0.1% ≈ 43 minutes downtime/month

Quick recall

Everything you need if you only revisit this box.

  • BookStore uses slos slis error budgets as a standard delivery practice.
  • Prefer automation and versioned config over manual server changes.
  • Staging proves changes before customer-facing promotion.
  • Observability confirms success — do not rely on silence alone.
  • Rollback plans must be tested, not invented during an outage.
  • Security and least privilege apply to every pipeline and cluster role.

Test yourself

Answer these before moving on — recall is what makes it stick.