Why this matters
- Legacy systems and vendor docs still reference Paxos and ZAB — you will operate them even if you implement Raft elsewhere.
- Architect interviews may compare consensus algorithms and ask when ZooKeeper vs etcd fits.
- You do not implement Paxos from scratch — you understand guarantees and failure modes of systems that use it.
- ShardPay uses ZooKeeper for legacy fraud-rule coordination while migrating critical paths to etcd/Raft.
Conceptual comparison
Paxos proves that distributed processes can agree on a single value despite failures — via propose, prepare, and accept phases. Multi-Paxos elects a stable leader to run continuous Paxos instances as a log. ZAB (ZooKeeper Atomic Broadcast) orders state updates for ZooKeeper's replicated tree — optimized for primary-backup with crash recovery. Raft repackages similar safety ideas with explicit leader election and log matching rules designed for implementability.
ShardPay engineers reading ZK docs encounter zxid (ZAB transaction id) and ephemeral sequential nodes — this article maps those terms to concepts you already know from Raft.
Algorithm families
- Paxos (single decree) — Agree on one value in one round. Rarely used alone in production — building blocks for Multi-Paxos. ShardPay's early prototype used a Paxos library for lock service before migrating to etcd.
- Multi-Paxos — Sequence of slots, each decided by Paxos; stable leader optimizes by skipping prepare phase. Conceptually similar to Raft log replication — leader proposes, majority accepts, learners apply.
- ZAB — ZooKeeper's protocol: primary order broadcasts, followers ack, quorum sync before commit. ShardPay's fraud-rule engine watches
/rules/merchant-99/limitznodes — updates propagate in ZAB order to all ZK followers. - Raft — Explicit terms, vote grants, log consistency check on election. ShardPay chose Raft (via etcd) for new services because on-call engineers can trace leader elections in structured logs.
- Safety vs liveness — All guarantee safety (no two decisions conflict) given majority survival; liveness requires eventual leader election. ShardPay monitors ZK
ruokand etcdhealth— loss of quorum blocks writes (CP) on both.
ZooKeeper and ZAB in ShardPay
ShardPay's legacy fraud coordination stores dynamic velocity limits in ZooKeeper. Rule updates are sequential; watchers fire in order. During ZK leader failover, writes pause briefly (< 2s) while ZAB elects new primary — fraud checks fail open to "allow with audit flag" per runbook.
Walkthrough: Updating a fraud limit via ZooKeeper
- Risk engineer sets
/shardpay/fraud/limits/merchant-42todailyMax=100000. - ZK primary (leader) assigns
zxid=0x10000000a3and broadcasts to followers. - Majority of followers persist and ack; change is committed in ZAB order.
- All ShardPay fraud-evaluator instances watching the znode receive watch event in order.
- Evaluators reload limit atomically — no evaluator sees
50000after100000(monotonic reads on ZAB stream).
// ShardPay fraud service — ZooKeeper watch (legacy path)
zkClient.getData("/shardpay/fraud/limits/" + merchantId, (event) -> {
if (event.getType() == Watcher.Event.EventType.NodeDataChanged) {
limitsCache.reload(merchantId);
}
}, null);
When to use which system
Operational mapping
- ZooKeeper + ZAB — Mature coordination: locks, leader election, config trees, watches. ShardPay keeps ZK for fraud limits until etcd migration completes; ZK excels at ephemeral nodes and ordered watches.
- etcd + Raft — Kubernetes and cloud-native stacks; gRPC API; transactional compare-and-swap. ShardPay's shard directory and feature flags live in etcd — same Raft mental model as this module's Raft article.
- Kafka KRaft — Raft-based metadata quorum replacing ZooKeeper for Kafka. ShardPay's event bus uses KRaft mode — partition leadership is separate from KRaft controller but metadata commits are Raft-ordered.
- Custom Paxos — Avoid unless you are a database vendor. ShardPay rule: buy consensus, do not build — operate etcd/ZK/KRaft and test with Jepsen.
- Migration path — Dual-write fraud limits to ZK and etcd during cutover; compare watch ordering; retire ZK watchers after 30-day parity check.
Example: Misreading Paxos "phase 1"
An engineer assumed ZooKeeper reads are always linearizable from any follower. ZK follower reads without sync() may be stale. ShardPay's fraud path uses zk.sync(path) before reading critical limits after deploy — ensures local ZK session has caught up to leader's committed zxid.
Quick recall
Everything you need if you only revisit this box.
- Paxos is the classic foundation; Multi-Paxos and Raft are continuous-log variants.
- ZAB powers ZooKeeper ordering — understand zxid and primary failover.
- Operate consensus systems (ZK, etcd, KRaft); rarely build Paxos from scratch.
Test yourself
Answer these before moving on — recall is what makes it stick.