PrepZone Logo
PrepZone

HPA, Resource Requests/Limits, and QoS

Horizontal Pod Autoscaler, CPU/memory limits, and Quality of Service classes.

Why this matters

  • This topic directly affects how reliably BookStore reaches production — hpa and resource limits keep bookstore pods scaled and schedulable under load..
  • Interviewers connect hands-on commands and manifests to real delivery stories, not buzzwords.
  • Later modules assume you can explain both the why and the concrete file or command involved.
  • Platform maturity shows up when teams automate this instead of relying on tribal knowledge.
Control plane
API serverkubectl / controllers
etcdCluster state store
SchedulerPod → node placement
Controller managerReconcile loops
Worker node
kubeletPod lifecycle
kube-proxyService networking
Container runtimecontainerd
BookStore podYour workload
The control plane manages desired state. Worker nodes run your containerised workloads.

Requests and limits

Schedulable and capped: BookStore engineers treat this as part of the standard path from laptop to hpa resource limits readiness. Document decisions in the team runbook so on-call knows which knobs exist.

Key ideas

  • Requests and limits — primary idea for hpa-resource-limits
  • BookStore context — catalog API, checkout, and inventory services share the same pattern
  • Automation — prefer pipeline jobs over manual SSH steps
  • Verification — staging must prove the change before prod traffic

HPA on CPU

Scale BookStore API: When staging matches production architecture, BookStore catches misconfigurations early. Pair this section's practice with observability dashboards to confirm behavior under load.

Key ideas

  • HPA on CPU — operational detail
  • Rollback — know how to revert without rebuilding artefacts
  • Security — least privilege for deploy roles
  • Documentation — link runbooks from the service README

Production checklist

Before promoting BookStore changes tied to this topic, run automated tests, inspect artefact immutability (image digest or JAR checksum), execute a staging smoke test on /actuator/health, and watch error-rate dashboards for thirty minutes after prod rollout.

Key ideas

  • Staging soak — validate under synthetic load
  • Change ticket — attach pipeline URL and artefact digest
  • On-call — page owner stays on dashboards during rollout
  • Post-deploy — record metrics baseline for comparison
Java
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: bookstore-api
spec:
  minReplicas: 2
  maxReplicas: 10
  metrics:
    - type: Resource
      resource:
        name: cpu
        target:
          type: Utilization
          averageUtilization: 70

Quick recall

Everything you need if you only revisit this box.

  • BookStore uses hpa resource limits as a standard delivery practice.
  • Prefer automation and versioned config over manual server changes.
  • Staging proves changes before customer-facing promotion.
  • Observability confirms success — do not rely on silence alone.
  • Rollback plans must be tested, not invented during an outage.
  • Security and least privilege apply to every pipeline and cluster role.

Test yourself

Answer these before moving on — recall is what makes it stick.