PrepZone Logo
PrepZone

On-Call, Runbooks, and Postmortems

Incident lifecycle, blameless postmortems, runbook templates, and on-call best practices.

Why this matters

  • This topic directly affects how reliably BookStore reaches production — runbooks and blameless postmortems guide bookstore on-call from alert to stable service..
  • Interviewers connect hands-on commands and manifests to real delivery stories, not buzzwords.
  • Later modules assume you can explain both the why and the concrete file or command involved.
  • Platform maturity shows up when teams automate this instead of relying on tribal knowledge.
LogsStructured JSON, Loki/ELK
MetricsCounters, gauges, Prometheus
TracesSpans, OpenTelemetry

All three correlate via trace ID and service labels

Logs answer what happened. Metrics answer how much. Traces answer where time was spent.

Incident lifecycle

Detect, mitigate, learn: BookStore engineers treat this as part of the standard path from laptop to incident response runbooks readiness. Document decisions in the team runbook so on-call knows which knobs exist.

Key ideas

  • Incident lifecycle — primary idea for incident-response-runbooks
  • BookStore context — catalog API, checkout, and inventory services share the same pattern
  • Automation — prefer pipeline jobs over manual SSH steps
  • Verification — staging must prove the change before prod traffic

Blameless postmortem

Action items with owners: When staging matches production architecture, BookStore catches misconfigurations early. Pair this section's practice with observability dashboards to confirm behavior under load.

Key ideas

  • Blameless postmortem — operational detail
  • Rollback — know how to revert without rebuilding artefacts
  • Security — least privilege for deploy roles
  • Documentation — link runbooks from the service README

Production checklist

Before promoting BookStore changes tied to this topic, run automated tests, inspect artefact immutability (image digest or JAR checksum), execute a staging smoke test on /actuator/health, and watch error-rate dashboards for thirty minutes after prod rollout.

Key ideas

  • Staging soak — validate under synthetic load
  • Change ticket — attach pipeline URL and artefact digest
  • On-call — page owner stays on dashboards during rollout
  • Post-deploy — record metrics baseline for comparison
Java
## BookStore checkout latency
1. Ack page in 5m
2. Check Grafana: p95 latency panel
3. kubectl logs deploy/bookstore-api --since=10m
4. Roll back if error budget burn > 2x

Quick recall

Everything you need if you only revisit this box.

  • BookStore uses incident response runbooks as a standard delivery practice.
  • Prefer automation and versioned config over manual server changes.
  • Staging proves changes before customer-facing promotion.
  • Observability confirms success — do not rely on silence alone.
  • Rollback plans must be tested, not invented during an outage.
  • Security and least privilege apply to every pipeline and cluster role.

Test yourself

Answer these before moving on — recall is what makes it stick.