PrepZone Logo
PrepZone

State Stores, RocksDB, and Changelog Topics

Fault-tolerant local state in Kafka Streams — RocksDB on disk backed by compacted changelog topics.

Why this matters

  • VaultCommerce fraud aggregates store session counts in vaultcommerce-fraud-detector-order-count-store.
  • Standby replicas (num.standby.replicas=1) reduce recovery time on failover.
  • State store disk must be sized — RocksDB grows with distinct keys.
  • Interactive queries expose state via REST for debugging (use carefully in prod).
Stream task
RocksDBLocal state
Compacted topic
Local RocksDB state is rebuilt from a compacted changelog topic on restart or rebalance.

RocksDB and changelog

Every state store maps to internal changelog topic applicationId-storeName-changelog. On restore, consumer replays changelog from beginning (or checkpoint). VaultCommerce monitors changelog lag during redeploys — large state = long restore.

Key points

  • State store — local key-value materialization of aggregations/joins
  • Changelog topic — compacted backup of state store updates
  • RocksDB — embedded LSM storage under Kafka Streams
  • Standby replica — warm copy for faster task migration
  • Restore consumer — replays changelog on task assignment

Standby and caching

num.standby.replicas=1 keeps hot copy on another instance. statestore.cache.max.bytes bounds in-memory cache before flush to RocksDB. VaultCommerce 8GB ephemeral disk per pod for state-heavy topologies.

VaultCommerce rollout checklist

Before promoting changes that touch the VaultCommerce order and payment event backbone, run the staging KRaft cluster (Kafka 3.7+, Schema Registry 7.x) through a 10k events/min soak test. Compare producer request latency p99 and consumer lag per group against the pre-deploy baseline. RocksDB local; changelog topic for fault tolerance. Document the change in the internal topic registry, attach Grafana screenshots to the change ticket, and keep an engineer on lag dashboards for 30 minutes after production rollout — roll back the service release before altering broker-level settings if lag or under-replicated partitions spike.

Java
// VaultCommerce streams config — state store tuning
props.put(StreamsConfig.APPLICATION_ID_CONFIG, "vaultcommerce-fraud-detector");
props.put(StreamsConfig.STATE_DIR_CONFIG, "/var/kafka-streams/state");
props.put(StreamsConfig.NUM_STANDBY_REPLICAS_CONFIG, 1);
props.put(StreamsConfig.CACHE_MAX_BYTES_BUFFERING_CONFIG, 10 * 1024 * 1024L);
props.put(StreamsConfig.PROCESSING_GUARANTEE_CONFIG, StreamsConfig.EXACTLY_ONCE_V2);

Quick recall

Everything you need if you only revisit this box.

  1. RocksDB local; changelog topic for fault tolerance.
  2. Size disk for state store growth.
  3. Standby replicas speed up failover.

Test yourself

Answer these before moving on — recall is what makes it stick.