PrepZone Logo
PrepZone

KRaft: Kafka Without ZooKeeper

Modern Kafka runs metadata in KRaft quorum — faster failover, simpler ops, and the only path forward on Kafka 4.x.

Why this matters

  • VaultCommerce migrated from ZooKeeper to KRaft in 2025 — failover dropped from ~15s to under 5s.
  • Kafka 4.x removes ZooKeeper entirely; KRaft is the only supported metadata mode going forward.
  • Simpler ops: one fewer system to patch, monitor, and secure.
  • Interviewers now expect KRaft architecture, not ZK-based explanations.
KRaft controllers
Controller 1
Controller 2
Controller 3
Data brokers
Broker 1Partitions
Broker 2Partitions
Broker 3Partitions
Controller quorum stores cluster metadata. Brokers handle data partitions — no ZooKeeper required.

Controllers and metadata log

A subset of brokers (or dedicated nodes) run as KRaft controllers. They replicate metadata records — topic configs, partition assignments, broker registrations — via Raft. On broker startup, it reads the metadata log to know which partitions to host. VaultCommerce runs 3 controllers (odd quorum) across AZs.

Key points

  • KRaft quorum — odd number of controller nodes (typically 3 or 5) for Raft consensus
  • Metadata log — internal topic storing cluster topology changes
  • Controller leader — single active controller handling metadata mutations
  • Combined mode — broker and controller on same node (dev/small prod)
  • Dedicated controllers — separate JVMs for large clusters (VaultCommerce at 6+ brokers)

Migration and compatibility

Kafka 3.7 supports KRaft-only new clusters and ZK migration tooling. VaultCommerce used kafka-metadata-quorum describe to verify quorum health. Client code is unchanged — only broker and admin tooling differ.

VaultCommerce rollout checklist

Before promoting changes that touch the VaultCommerce order and payment event backbone, run the staging KRaft cluster (Kafka 3.7+, Schema Registry 7.x) through a 10k events/min soak test. Compare producer request latency p99 and consumer lag per group against the pre-deploy baseline. KRaft replaces ZooKeeper with internal Raft metadata. Document the change in the internal topic registry, attach Grafana screenshots to the change ticket, and keep an engineer on lag dashboards for 30 minutes after production rollout — roll back the service release before altering broker-level settings if lag or under-replicated partitions spike.

Java
# VaultCommerce — verify KRaft quorum health after deploy
kafka-metadata-quorum.sh --bootstrap-server kafka-1:9092 describe --status

# QuorumId: abc123...
# LeaderId: 1  LeaderEpoch: 42
# Voters: [1, 2, 3]  Observers: []

kafka-broker-api-versions.sh --bootstrap-server kafka-1:9092 | head -5

Quick recall

Everything you need if you only revisit this box.

  1. KRaft replaces ZooKeeper with internal Raft metadata.
  2. Use odd-numbered controller quorums across AZs.
  3. Clients are unaware — migration is broker-side.

Test yourself

Answer these before moving on — recall is what makes it stick.