Why this matters
- VaultCommerce fraud aggregates store session counts in
vaultcommerce-fraud-detector-order-count-store. - Standby replicas (
num.standby.replicas=1) reduce recovery time on failover. - State store disk must be sized — RocksDB grows with distinct keys.
- Interactive queries expose state via REST for debugging (use carefully in prod).
RocksDB and changelog
Every state store maps to internal changelog topic applicationId-storeName-changelog. On restore, consumer replays changelog from beginning (or checkpoint). VaultCommerce monitors changelog lag during redeploys — large state = long restore.
Key points
- State store — local key-value materialization of aggregations/joins
- Changelog topic — compacted backup of state store updates
- RocksDB — embedded LSM storage under Kafka Streams
- Standby replica — warm copy for faster task migration
- Restore consumer — replays changelog on task assignment
Standby and caching
num.standby.replicas=1 keeps hot copy on another instance. statestore.cache.max.bytes bounds in-memory cache before flush to RocksDB. VaultCommerce 8GB ephemeral disk per pod for state-heavy topologies.
VaultCommerce rollout checklist
Before promoting changes that touch the VaultCommerce order and payment event backbone, run the staging KRaft cluster (Kafka 3.7+, Schema Registry 7.x) through a 10k events/min soak test. Compare producer request latency p99 and consumer lag per group against the pre-deploy baseline. RocksDB local; changelog topic for fault tolerance. Document the change in the internal topic registry, attach Grafana screenshots to the change ticket, and keep an engineer on lag dashboards for 30 minutes after production rollout — roll back the service release before altering broker-level settings if lag or under-replicated partitions spike.
// VaultCommerce streams config — state store tuning
props.put(StreamsConfig.APPLICATION_ID_CONFIG, "vaultcommerce-fraud-detector");
props.put(StreamsConfig.STATE_DIR_CONFIG, "/var/kafka-streams/state");
props.put(StreamsConfig.NUM_STANDBY_REPLICAS_CONFIG, 1);
props.put(StreamsConfig.CACHE_MAX_BYTES_BUFFERING_CONFIG, 10 * 1024 * 1024L);
props.put(StreamsConfig.PROCESSING_GUARANTEE_CONFIG, StreamsConfig.EXACTLY_ONCE_V2);
Quick recall
Everything you need if you only revisit this box.
- RocksDB local; changelog topic for fault tolerance.
- Size disk for state store growth.
- Standby replicas speed up failover.
Test yourself
Answer these before moving on — recall is what makes it stick.