Why this matters
- VaultCommerce lost a broker during a zone outage — ISR-aware producers continued without data loss.
- Under-replicated partitions are the leading indicator of imminent unavailability.
- Leaders redistribute automatically via KRaft-controlled elections after broker failure.
- Understanding ISR explains why
acks=allalone is insufficient withoutmin.insync.replicas.
Leader, followers, and fetch
Producers and consumers talk only to the leader for a partition. Followers replicate by fetching from the leader. When a follower catches up within replica.lag.time.max.ms, it joins the ISR. VaultCommerce alerts when ISR shrinks below 2 for any production topic.
Key points
- ISR (in-sync replicas) — followers within lag threshold; only ISR members count for
acks=all - Leader election — KRaft controller promotes a new leader from ISR on failure
- Under-replicated partition — leader has fewer ISR replicas than replication factor
- min.insync.replicas — minimum ISR size required to accept writes
- Preferred leader —
kafka-leader-electionrebalances leaders across brokers for even load
Unclean leader election
unclean.leader.election.enable=false (default) means if all ISR replicas die, the partition goes offline rather than electing an out-of-sync follower and losing data. VaultCommerce accepts brief unavailability over silent truncation.
VaultCommerce rollout checklist
Before promoting changes that touch the VaultCommerce order and payment event backbone, run the staging KRaft cluster (Kafka 3.7+, Schema Registry 7.x) through a 10k events/min soak test. Compare producer request latency p99 and consumer lag per group against the pre-deploy baseline. ISR defines which replicas count for durable acks. Document the change in the internal topic registry, attach Grafana screenshots to the change ticket, and keep an engineer on lag dashboards for 30 minutes after production rollout — roll back the service release before altering broker-level settings if lag or under-replicated partitions spike.
# VaultCommerce broker server.properties excerpt — replication safety
default.replication.factor: 3
min.insync.replicas: 2
unclean.leader.election.enable: false
replica.lag.time.max.ms: 30000
auto.leader.rebalance.enable: true
# Topic override for critical payment events
# min.insync.replicas=2 on vaultcommerce.payments.captured.v1
Quick recall
Everything you need if you only revisit this box.
- ISR defines which replicas count for durable acks.
- Never enable unclean leader election for financial topics.
- Alert on under-replicated partitions before leaders fail.
Test yourself
Answer these before moving on — recall is what makes it stick.