Why this matters
- VaultCommerce pages on-call when inventory lag exceeds 10k for 5 minutes during sales.
- Scaling past partition count does not reduce lag — add partitions (carefully) or optimize handlers.
- Lag growth rate matters more than absolute value for capacity planning.
- Backpressure upstream prevents unbounded lag from drowning downstream Postgres.
Measuring lag
kafka-consumer-groups.sh --describe shows CURRENT-OFFSET vs LOG-END-OFFSET per partition. Burrow or Kafka lag exporter pushes to Prometheus. VaultCommerce dashboards aggregate sum(lag) per service with partition breakdown for skew detection.
Key points
- Consumer lag — LOG-END-OFFSET minus CURRENT-OFFSET per partition
- Lag spike — deploy bug, slow DB, or traffic surge
- Partition skew — one partition lagging while others are fine (hot key)
- HPA ceiling — min(replicas, partition count) for Kafka consumers
- Backpressure — slow producers or shed load when consumers saturated
Scaling and backpressure
Scale consumers to partition count. If still lagging, optimize handler (batch DB writes, cache lookups). If producer outpaces forever, add partitions (requires key planning) or throttle producers. VaultCommerce enables rate limiting on analytics publish during inventory catch-up. During Black Friday 2025, temporarily pausing non-critical publishers dropped inventory lag from 85k to 12k in 18 minutes while handlers caught up.
VaultCommerce rollout checklist
Before promoting changes that touch inventory-service and payment-service consumer groups, run the staging KRaft cluster (Kafka 3.7+, Schema Registry 7.x) through a 10k events/min soak test. Compare producer request latency p99 and consumer lag per group against the pre-deploy baseline. Lag = how far behind the log consumers are. Document the change in the internal topic registry, attach Grafana screenshots to the change ticket, and keep an engineer on lag dashboards for 30 minutes after production rollout — roll back the service release before altering broker-level settings if lag or under-replicated partitions spike.
# VaultCommerce on-call — lag check per service
kafka-consumer-groups.sh --bootstrap-server kafka-1:9092 \
--describe --group inventory-service | awk 'NR>1 {lag=$6-$5; if(lag>1000) print $0, "LAG="lag}'
# Prometheus alert: kafka_consumergroup_lag_sum{group="inventory-service"} > 10000
Quick recall
Everything you need if you only revisit this box.
- Lag = how far behind the log consumers are.
- Scale consumers to partition count, then optimize handlers.
- Alert on lag growth rate, not just absolute value.
Test yourself
Answer these before moving on — recall is what makes it stick.