Why this matters
- VaultCommerce L1 on-call runs this checklist in under 10 minutes.
- Prevents premature offset resets and unnecessary broker restarts.
- Documents escalation path to platform team with captured evidence.
- Living runbook updated after every postmortem.
First five minutes
Acknowledge page. Open Grafana Kafka dashboard. Note alert name and time. Check #deploys channel for correlated releases. Run cluster-id and quorum describe. If ActiveControllerCount != 1, escalate immediately to platform.
Key points
- Triage order — cluster → brokers → consumers → clients
- Evidence bundle — lag describe, log snippet, deploy correlation
- Escalation threshold — URP > 0 for 10m, controller loss, disk > 90%
- Safe actions — scale consumers (≤ partitions), throttle non-critical producers
- Unsafe without approval — offset reset, broker restart, retention change
Before escalating or acting
Capture: consumer group lag output, URP count, broker disk %, last 50 broker log lines, affected service deploy version. Do NOT reset offsets without platform approval. Do NOT restart multiple brokers. VaultCommerce ticket template requires all captures attached.
VaultCommerce rollout checklist
Before promoting changes that touch the VaultCommerce order and payment event backbone, run the staging KRaft cluster (Kafka 3.7+, Schema Registry 7.x) through a 10k events/min soak test. Compare producer request latency p99 and consumer lag per group against the pre-deploy baseline. Cluster quorum and controller before consumers. Document the change in the internal topic registry, attach Grafana screenshots to the change ticket, and keep an engineer on lag dashboards for 30 minutes after production rollout — roll back the service release before altering broker-level settings if lag or under-replicated partitions spike.
# VaultCommerce on-call cheat sheet — copy/paste block
export BS=kafka-1.vaultcommerce.internal:9092
kafka-cluster.sh cluster-id --bootstrap-server $BS
kafka-metadata-quorum.sh --bootstrap-server $BS describe --status
kafka-consumer-groups.sh --bootstrap-server $BS --list
kafka-consumer-groups.sh --bootstrap-server $BS --describe --group inventory-service
kafka-log-dirs.sh --bootstrap-server $BS --describe --human-readable | tail -20
# JMX offline replicas (via metrics or):
kafka-topics.sh --bootstrap-server $BS --describe | grep -i "Leader: none"
# Escalate to #kafka-platform with output + deploy timeline
Quick recall
Everything you need if you only revisit this box.
- Cluster quorum and controller before consumers.
- Capture evidence before changes.
- No offset reset or mass restart without approval.
Test yourself
Answer these before moving on — recall is what makes it stick.