PrepZone Logo
PrepZone

Troubleshooting: Lag, Rebalance, Disk

Step-by-step diagnosis for the five Kafka incidents VaultCommerce sees most often in production.

Why this matters

  • VaultCommerce on-call resolves 80% of Kafka pages with five playbooks.
  • Start with consumer lag vs broker health — different fixes.
  • Recent deploy correlation cuts MTTR in half.
  • This article is the capstone for production module.
Incident reported
High lag?Scale consumers
Broker down?Check ISR
Disk full?Retention / capacity
Start with consumer lag, then check broker health, then partition leadership and disk.

Consumer lag playbook

  1. Is lag all partitions or one? (skew → hot key). 2) Consumer count vs partitions? 3) Handler slow? Check DB traces. 4) Producer surge? Check ingress rate. VaultCommerce scales consumers first, optimizes SQL second, adds partitions last.

Key points

  • Symptom clustering — lag + rebalance often = slow poll
  • Partition skew — single partition lagging implicates key distribution
  • URP — under-replicated partitions need broker recovery
  • Disk full — segment delete cannot run if disk 100%
  • Deploy correlation — check release timeline first

Broker and disk playbook

URP > 0: identify offline broker, check disk and logs, restart if stuck. Disk full: expand volume or drop retention on non-critical topic. Controller failover: verify ActiveControllerCount=1. VaultCommerce runbook links Grafana panel per step.

VaultCommerce rollout checklist

Before promoting changes that touch the VaultCommerce order and payment event backbone, run the staging KRaft cluster (Kafka 3.7+, Schema Registry 7.x) through a 10k events/min soak test. Compare producer request latency p99 and consumer lag per group against the pre-deploy baseline. Lag: skew vs uniform → different root causes. Document the change in the internal topic registry, attach Grafana screenshots to the change ticket, and keep an engineer on lag dashboards for 30 minutes after production rollout — roll back the service release before altering broker-level settings if lag or under-replicated partitions spike.

Java
# VaultCommerce incident triage — run in order
# 1. Cluster health
kafka-cluster.sh cluster-id --bootstrap-server kafka-1:9092
kafka-metadata-quorum.sh --bootstrap-server kafka-1:9092 describe --status

# 2. Consumer lag
kafka-consumer-groups.sh --bootstrap-server kafka-1:9092 --describe --group inventory-service

# 3. Disk / URP
kafka-log-dirs.sh --bootstrap-server kafka-1:9092 --describe | grep -E "size|offline"

# 4. Recent errors
kubectl logs -l app=kafka-broker --tail=200 | grep -iE "error|exception|fatal"

Quick recall

Everything you need if you only revisit this box.

  1. Lag: skew vs uniform → different root causes.
  2. URP and disk before restarting random brokers.
  3. Never reset offsets without understanding why.

Test yourself

Answer these before moving on — recall is what makes it stick.