PrepZone Logo
PrepZone

Logs, Metrics, and Distributed Tracing

The three pillars, Prometheus/Grafana stack, structured logging, and OpenTelemetry.

Why this matters

  • This topic directly affects how reliably BookStore reaches production — logs, metrics, and traces explain bookstore behavior under load and during incidents..
  • Interviewers connect hands-on commands and manifests to real delivery stories, not buzzwords.
  • Later modules assume you can explain both the why and the concrete file or command involved.
  • Platform maturity shows up when teams automate this instead of relying on tribal knowledge.
LogsStructured JSON, Loki/ELK
MetricsCounters, gauges, Prometheus
TracesSpans, OpenTelemetry

All three correlate via trace ID and service labels

Logs answer what happened. Metrics answer how much. Traces answer where time was spent.

Metrics and logs

Prometheus and structured JSON: BookStore engineers treat this as part of the standard path from laptop to observability three pillars readiness. Document decisions in the team runbook so on-call knows which knobs exist.

Key ideas

  • Metrics and logs — primary idea for observability-three-pillars
  • BookStore context — catalog API, checkout, and inventory services share the same pattern
  • Automation — prefer pipeline jobs over manual SSH steps
  • Verification — staging must prove the change before prod traffic

Traces

Follow checkout request path: When staging matches production architecture, BookStore catches misconfigurations early. Pair this section's practice with observability dashboards to confirm behavior under load.

Key ideas

  • Traces — operational detail
  • Rollback — know how to revert without rebuilding artefacts
  • Security — least privilege for deploy roles
  • Documentation — link runbooks from the service README

Production checklist

Before promoting BookStore changes tied to this topic, run automated tests, inspect artefact immutability (image digest or JAR checksum), execute a staging smoke test on /actuator/health, and watch error-rate dashboards for thirty minutes after prod rollout.

Key ideas

  • Staging soak — validate under synthetic load
  • Change ticket — attach pipeline URL and artefact digest
  • On-call — page owner stays on dashboards during rollout
  • Post-deploy — record metrics baseline for comparison
Java
# Prometheus scrape annotation on BookStore pods
metadata:
  annotations:
    prometheus.io/scrape: "true"
    prometheus.io/port: "8080"
    prometheus.io/path: "/actuator/prometheus"

Quick recall

Everything you need if you only revisit this box.

  • BookStore uses observability three pillars as a standard delivery practice.
  • Prefer automation and versioned config over manual server changes.
  • Staging proves changes before customer-facing promotion.
  • Observability confirms success — do not rely on silence alone.
  • Rollback plans must be tested, not invented during an outage.
  • Security and least privilege apply to every pipeline and cluster role.

Test yourself

Answer these before moving on — recall is what makes it stick.