PrepZone Logo
PrepZone

Broker Sizing, Partitions, and JVM Tuning

Disk, network, heap, and partition count guidelines for a VaultCommerce production cluster.

Why this matters

  • VaultCommerce production: 6 brokers, 2TB NVMe each, 12 partitions per hot topic.
  • Kafka is I/O bound — heap 6GB, G1GC; rely on OS cache for page reads.
  • Network: 10 Gbps NIC; replication traffic doubles write volume with RF=3.
  • Under-sized clusters show up as disk full and fetch latency long before CPU maxes.
ProducersOrder, payment services
Broker cluster3+ brokers, KRaft
Consumer groupsInventory, analytics
Topic: vaultcommerce.orders.placed.v1
Partition 0Leader on broker-1
Partition 1Leader on broker-2
Partition 2Leader on broker-3
Producers write to topic partitions on brokers. Consumers read via consumer groups.

Disk and retention math

Daily ingress × retention days × RF = raw disk. VaultCommerce: 500GB/day × 7 × 3 = 10.5TB cluster raw; spread across 6 brokers ≈ 1.75TB/broker + 30% headroom. Monitor kafka.log:type=Log,name=Size.

Key points

  • Retention sizing — ingress rate × days × replication factor
  • Page cache — Kafka relies on OS cache, not JVM heap for hot data
  • Network bandwidth — producer ingress + replication + consumer fetch
  • Partition count — target < 4000 partitions per broker for metadata health
  • num.io.threads / num.network.threads — broker request parallelism

JVM and OS tuning

KAFKA_HEAP_OPTS=-Xmx6g -Xms6g. Avoid huge heaps — long GC pauses cause ISR shrink. log.dirs on dedicated NVMe. vm.swappiness=1. VaultCommerce disables transparent huge pages per Confluent guidance.

VaultCommerce rollout checklist

Before promoting changes that touch the VaultCommerce order and payment event backbone, run the staging KRaft cluster (Kafka 3.7+, Schema Registry 7.x) through a 10k events/min soak test. Compare producer request latency p99 and consumer lag per group against the pre-deploy baseline. Disk = daily volume × retention × RF + headroom. Document the change in the internal topic registry, attach Grafana screenshots to the change ticket, and keep an engineer on lag dashboards for 30 minutes after production rollout — roll back the service release before altering broker-level settings if lag or under-replicated partitions spike.

Java
# VaultCommerce broker JVM and disk check
echo $KAFKA_HEAP_OPTS   # -Xmx6g -Xms6g
df -h /var/kafka/data
kafka-broker-api-versions.sh --bootstrap-server localhost:9092 2>/dev/null | head -1

# Rough partition load per broker
kafka-metadata.sh --snapshot /var/kafka/data/__cluster_metadata-0/*.checkpoint \
  --command "partition-count" 2>/dev/null || \
  kafka-log-dirs.sh --bootstrap-server localhost:9092 --describe | grep -c partition

Quick recall

Everything you need if you only revisit this box.

  1. Disk = daily volume × retention × RF + headroom.
  2. Small JVM heap; big OS page cache.
  3. Partitions for parallelism, not unlimited.

Test yourself

Answer these before moving on — recall is what makes it stick.