PrepZone Logo
PrepZone

On-Call Runbook and Command Cheat Sheet

The commands and checks every VaultCommerce on-call engineer runs before escalating a Kafka incident.

Why this matters

  • VaultCommerce L1 on-call runs this checklist in under 10 minutes.
  • Prevents premature offset resets and unnecessary broker restarts.
  • Documents escalation path to platform team with captured evidence.
  • Living runbook updated after every postmortem.
Kafka brokersJMX
Prometheus
Grafana + PagerDuty
JMX metrics exported to Prometheus. Lag alerts route to on-call. Logs correlate via trace IDs.

First five minutes

Acknowledge page. Open Grafana Kafka dashboard. Note alert name and time. Check #deploys channel for correlated releases. Run cluster-id and quorum describe. If ActiveControllerCount != 1, escalate immediately to platform.

Key points

  • Triage order — cluster → brokers → consumers → clients
  • Evidence bundle — lag describe, log snippet, deploy correlation
  • Escalation threshold — URP > 0 for 10m, controller loss, disk > 90%
  • Safe actions — scale consumers (≤ partitions), throttle non-critical producers
  • Unsafe without approval — offset reset, broker restart, retention change

Before escalating or acting

Capture: consumer group lag output, URP count, broker disk %, last 50 broker log lines, affected service deploy version. Do NOT reset offsets without platform approval. Do NOT restart multiple brokers. VaultCommerce ticket template requires all captures attached.

VaultCommerce rollout checklist

Before promoting changes that touch the VaultCommerce order and payment event backbone, run the staging KRaft cluster (Kafka 3.7+, Schema Registry 7.x) through a 10k events/min soak test. Compare producer request latency p99 and consumer lag per group against the pre-deploy baseline. Cluster quorum and controller before consumers. Document the change in the internal topic registry, attach Grafana screenshots to the change ticket, and keep an engineer on lag dashboards for 30 minutes after production rollout — roll back the service release before altering broker-level settings if lag or under-replicated partitions spike.

Java
# VaultCommerce on-call cheat sheet — copy/paste block
export BS=kafka-1.vaultcommerce.internal:9092

kafka-cluster.sh cluster-id --bootstrap-server $BS
kafka-metadata-quorum.sh --bootstrap-server $BS describe --status
kafka-consumer-groups.sh --bootstrap-server $BS --list
kafka-consumer-groups.sh --bootstrap-server $BS --describe --group inventory-service
kafka-log-dirs.sh --bootstrap-server $BS --describe --human-readable | tail -20

# JMX offline replicas (via metrics or):
kafka-topics.sh --bootstrap-server $BS --describe | grep -i "Leader: none"

# Escalate to #kafka-platform with output + deploy timeline

Quick recall

Everything you need if you only revisit this box.

  1. Cluster quorum and controller before consumers.
  2. Capture evidence before changes.
  3. No offset reset or mass restart without approval.

Test yourself

Answer these before moving on — recall is what makes it stick.