PrepZone Logo
PrepZone

Replication, ISR, Leaders, and Followers

How Kafka survives broker loss without losing committed messages — leaders, followers, and the in-sync replica set.

Why this matters

  • VaultCommerce lost a broker during a zone outage — ISR-aware producers continued without data loss.
  • Under-replicated partitions are the leading indicator of imminent unavailability.
  • Leaders redistribute automatically via KRaft-controlled elections after broker failure.
  • Understanding ISR explains why acks=all alone is insufficient without min.insync.replicas.
Producer
Leader P0Broker 1
Follower P0Broker 2
Follower P0Broker 3
Each partition has one leader and N-1 followers. Only the ISR set can become leader on failover.

Leader, followers, and fetch

Producers and consumers talk only to the leader for a partition. Followers replicate by fetching from the leader. When a follower catches up within replica.lag.time.max.ms, it joins the ISR. VaultCommerce alerts when ISR shrinks below 2 for any production topic.

Key points

  • ISR (in-sync replicas) — followers within lag threshold; only ISR members count for acks=all
  • Leader election — KRaft controller promotes a new leader from ISR on failure
  • Under-replicated partition — leader has fewer ISR replicas than replication factor
  • min.insync.replicas — minimum ISR size required to accept writes
  • Preferred leader — kafka-leader-election rebalances leaders across brokers for even load

Unclean leader election

unclean.leader.election.enable=false (default) means if all ISR replicas die, the partition goes offline rather than electing an out-of-sync follower and losing data. VaultCommerce accepts brief unavailability over silent truncation.

VaultCommerce rollout checklist

Before promoting changes that touch the VaultCommerce order and payment event backbone, run the staging KRaft cluster (Kafka 3.7+, Schema Registry 7.x) through a 10k events/min soak test. Compare producer request latency p99 and consumer lag per group against the pre-deploy baseline. ISR defines which replicas count for durable acks. Document the change in the internal topic registry, attach Grafana screenshots to the change ticket, and keep an engineer on lag dashboards for 30 minutes after production rollout — roll back the service release before altering broker-level settings if lag or under-replicated partitions spike.

Java
# VaultCommerce broker server.properties excerpt — replication safety
default.replication.factor: 3
min.insync.replicas: 2
unclean.leader.election.enable: false
replica.lag.time.max.ms: 30000
auto.leader.rebalance.enable: true

# Topic override for critical payment events
# min.insync.replicas=2 on vaultcommerce.payments.captured.v1

Quick recall

Everything you need if you only revisit this box.

  1. ISR defines which replicas count for durable acks.
  2. Never enable unclean leader election for financial topics.
  3. Alert on under-replicated partitions before leaders fail.

Test yourself

Answer these before moving on — recall is what makes it stick.