PrepZone Logo
PrepZone

Paxos and ZAB — Conceptual Overview

The classics behind modern consensus — enough depth to read ZooKeeper and Kafka KRaft docs.

Why this matters

  • Legacy systems and vendor docs still reference Paxos and ZAB — you will operate them even if you implement Raft elsewhere.
  • Architect interviews may compare consensus algorithms and ask when ZooKeeper vs etcd fits.
  • You do not implement Paxos from scratch — you understand guarantees and failure modes of systems that use it.
  • ShardPay uses ZooKeeper for legacy fraud-rule coordination while migrating critical paths to etcd/Raft.
Phase 1: PrepareProposer asks acceptors
Phase 2: AcceptQuorum accepts value
LearnValue chosen
Proposers compete; acceptors vote; learners apply chosen value once a quorum agrees.

Conceptual comparison

Paxos proves that distributed processes can agree on a single value despite failures — via propose, prepare, and accept phases. Multi-Paxos elects a stable leader to run continuous Paxos instances as a log. ZAB (ZooKeeper Atomic Broadcast) orders state updates for ZooKeeper's replicated tree — optimized for primary-backup with crash recovery. Raft repackages similar safety ideas with explicit leader election and log matching rules designed for implementability.

ShardPay engineers reading ZK docs encounter zxid (ZAB transaction id) and ephemeral sequential nodes — this article maps those terms to concepts you already know from Raft.

Algorithm families

  • Paxos (single decree) — Agree on one value in one round. Rarely used alone in production — building blocks for Multi-Paxos. ShardPay's early prototype used a Paxos library for lock service before migrating to etcd.
  • Multi-Paxos — Sequence of slots, each decided by Paxos; stable leader optimizes by skipping prepare phase. Conceptually similar to Raft log replication — leader proposes, majority accepts, learners apply.
  • ZAB — ZooKeeper's protocol: primary order broadcasts, followers ack, quorum sync before commit. ShardPay's fraud-rule engine watches /rules/merchant-99/limit znodes — updates propagate in ZAB order to all ZK followers.
  • Raft — Explicit terms, vote grants, log consistency check on election. ShardPay chose Raft (via etcd) for new services because on-call engineers can trace leader elections in structured logs.
  • Safety vs liveness — All guarantee safety (no two decisions conflict) given majority survival; liveness requires eventual leader election. ShardPay monitors ZK ruok and etcd health — loss of quorum blocks writes (CP) on both.

ZooKeeper and ZAB in ShardPay

ShardPay's legacy fraud coordination stores dynamic velocity limits in ZooKeeper. Rule updates are sequential; watchers fire in order. During ZK leader failover, writes pause briefly (< 2s) while ZAB elects new primary — fraud checks fail open to "allow with audit flag" per runbook.

Walkthrough: Updating a fraud limit via ZooKeeper

  1. Risk engineer sets /shardpay/fraud/limits/merchant-42 to dailyMax=100000.
  2. ZK primary (leader) assigns zxid=0x10000000a3 and broadcasts to followers.
  3. Majority of followers persist and ack; change is committed in ZAB order.
  4. All ShardPay fraud-evaluator instances watching the znode receive watch event in order.
  5. Evaluators reload limit atomically — no evaluator sees 50000 after 100000 (monotonic reads on ZAB stream).
Java
// ShardPay fraud service — ZooKeeper watch (legacy path)
zkClient.getData("/shardpay/fraud/limits/" + merchantId, (event) -> {
    if (event.getType() == Watcher.Event.EventType.NodeDataChanged) {
        limitsCache.reload(merchantId);
    }
}, null);

When to use which system

Operational mapping

  • ZooKeeper + ZAB — Mature coordination: locks, leader election, config trees, watches. ShardPay keeps ZK for fraud limits until etcd migration completes; ZK excels at ephemeral nodes and ordered watches.
  • etcd + Raft — Kubernetes and cloud-native stacks; gRPC API; transactional compare-and-swap. ShardPay's shard directory and feature flags live in etcd — same Raft mental model as this module's Raft article.
  • Kafka KRaft — Raft-based metadata quorum replacing ZooKeeper for Kafka. ShardPay's event bus uses KRaft mode — partition leadership is separate from KRaft controller but metadata commits are Raft-ordered.
  • Custom Paxos — Avoid unless you are a database vendor. ShardPay rule: buy consensus, do not build — operate etcd/ZK/KRaft and test with Jepsen.
  • Migration path — Dual-write fraud limits to ZK and etcd during cutover; compare watch ordering; retire ZK watchers after 30-day parity check.

Example: Misreading Paxos "phase 1"

An engineer assumed ZooKeeper reads are always linearizable from any follower. ZK follower reads without sync() may be stale. ShardPay's fraud path uses zk.sync(path) before reading critical limits after deploy — ensures local ZK session has caught up to leader's committed zxid.

Quick recall

Everything you need if you only revisit this box.

  1. Paxos is the classic foundation; Multi-Paxos and Raft are continuous-log variants.
  2. ZAB powers ZooKeeper ordering — understand zxid and primary failover.
  3. Operate consensus systems (ZK, etcd, KRaft); rarely build Paxos from scratch.

Test yourself

Answer these before moving on — recall is what makes it stick.