Why this matters
- VaultCommerce production: 6 brokers, 2TB NVMe each, 12 partitions per hot topic.
- Kafka is I/O bound — heap 6GB, G1GC; rely on OS cache for page reads.
- Network: 10 Gbps NIC; replication traffic doubles write volume with RF=3.
- Under-sized clusters show up as disk full and fetch latency long before CPU maxes.
Disk and retention math
Daily ingress × retention days × RF = raw disk. VaultCommerce: 500GB/day × 7 × 3 = 10.5TB cluster raw; spread across 6 brokers ≈ 1.75TB/broker + 30% headroom. Monitor kafka.log:type=Log,name=Size.
Key points
- Retention sizing — ingress rate × days × replication factor
- Page cache — Kafka relies on OS cache, not JVM heap for hot data
- Network bandwidth — producer ingress + replication + consumer fetch
- Partition count — target < 4000 partitions per broker for metadata health
- num.io.threads / num.network.threads — broker request parallelism
JVM and OS tuning
KAFKA_HEAP_OPTS=-Xmx6g -Xms6g. Avoid huge heaps — long GC pauses cause ISR shrink. log.dirs on dedicated NVMe. vm.swappiness=1. VaultCommerce disables transparent huge pages per Confluent guidance.
VaultCommerce rollout checklist
Before promoting changes that touch the VaultCommerce order and payment event backbone, run the staging KRaft cluster (Kafka 3.7+, Schema Registry 7.x) through a 10k events/min soak test. Compare producer request latency p99 and consumer lag per group against the pre-deploy baseline. Disk = daily volume × retention × RF + headroom. Document the change in the internal topic registry, attach Grafana screenshots to the change ticket, and keep an engineer on lag dashboards for 30 minutes after production rollout — roll back the service release before altering broker-level settings if lag or under-replicated partitions spike.
# VaultCommerce broker JVM and disk check
echo $KAFKA_HEAP_OPTS # -Xmx6g -Xms6g
df -h /var/kafka/data
kafka-broker-api-versions.sh --bootstrap-server localhost:9092 2>/dev/null | head -1
# Rough partition load per broker
kafka-metadata.sh --snapshot /var/kafka/data/__cluster_metadata-0/*.checkpoint \
--command "partition-count" 2>/dev/null || \
kafka-log-dirs.sh --bootstrap-server localhost:9092 --describe | grep -c partition
Quick recall
Everything you need if you only revisit this box.
- Disk = daily volume × retention × RF + headroom.
- Small JVM heap; big OS page cache.
- Partitions for parallelism, not unlimited.
Test yourself
Answer these before moving on — recall is what makes it stick.