PrepZone Logo
PrepZone

Rolling Updates, Rollbacks, and Probes

maxSurge, maxUnavailable, readiness/liveness probes, and instant rollback.

Why this matters

  • This topic directly affects how reliably BookStore reaches production — rolling updates and rollbacks change bookstore versions with bounded blast radius..
  • Interviewers connect hands-on commands and manifests to real delivery stories, not buzzwords.
  • Later modules assume you can explain both the why and the concrete file or command involved.
  • Platform maturity shows up when teams automate this instead of relying on tribal knowledge.
Rolling
v1.0 pods3 running
v1.1 pod addedOne at a time
v1.0 pods removedZero downtime
Blue/Green
Blue (v1.0)Live traffic
Green (v1.1)Validated, idle
Switch trafficInstant cutover
Rolling replaces pods gradually. Blue/green switches traffic between two full environments.

maxSurge and maxUnavailable

Surge without downtime: BookStore engineers treat this as part of the standard path from laptop to rolling updates rollbacks readiness. Document decisions in the team runbook so on-call knows which knobs exist.

Key ideas

  • maxSurge and maxUnavailable — primary idea for rolling-updates-rollbacks
  • BookStore context — catalog API, checkout, and inventory services share the same pattern
  • Automation — prefer pipeline jobs over manual SSH steps
  • Verification — staging must prove the change before prod traffic

Instant rollback

kubectl rollout undo: When staging matches production architecture, BookStore catches misconfigurations early. Pair this section's practice with observability dashboards to confirm behavior under load.

Key ideas

  • Instant rollback — operational detail
  • Rollback — know how to revert without rebuilding artefacts
  • Security — least privilege for deploy roles
  • Documentation — link runbooks from the service README

Production checklist

Before promoting BookStore changes tied to this topic, run automated tests, inspect artefact immutability (image digest or JAR checksum), execute a staging smoke test on /actuator/health, and watch error-rate dashboards for thirty minutes after prod rollout.

Key ideas

  • Staging soak — validate under synthetic load
  • Change ticket — attach pipeline URL and artefact digest
  • On-call — page owner stays on dashboards during rollout
  • Post-deploy — record metrics baseline for comparison
Java
kubectl rollout status deployment/bookstore-api
kubectl rollout undo deployment/bookstore-api

Quick recall

Everything you need if you only revisit this box.

  • BookStore uses rolling updates rollbacks as a standard delivery practice.
  • Prefer automation and versioned config over manual server changes.
  • Staging proves changes before customer-facing promotion.
  • Observability confirms success — do not rely on silence alone.
  • Rollback plans must be tested, not invented during an outage.
  • Security and least privilege apply to every pipeline and cluster role.

Test yourself

Answer these before moving on — recall is what makes it stick.