Why this matters
- VaultCommerce evaluated RabbitMQ, SQS, Redis Streams, and Kafka before standardizing on Kafka for the order event backbone.
- Picking the wrong tool means paying operational tax — SQS cannot replay last Tuesday's orders for a new analytics pipeline.
- Interviewers test whether you know why Kafka fits high-throughput logs vs when RabbitMQ's routing is simpler.
- Each technology optimizes for different delivery, ordering, and retention guarantees.
Task queues vs distributed logs
RabbitMQ and SQS excel at work distribution: one message, one worker, message deleted after ack. Kafka retains messages in an append-only log — multiple consumer groups read independently and new groups can rewind to offset 0. VaultCommerce's fraud team joined months later by consuming the full order history from the beginning.
Key points
- Task queue — message removed after single consumer acknowledges
- Pub/sub log — messages retained; subscribers track their own offset
- RabbitMQ — flexible routing (direct, topic, fanout); moderate throughput
- SQS — managed, at-least-once, no ordering unless FIFO (300 TPS/partition limit)
- Kafka — partition-scaled throughput, replay, log compaction for changelog topics
Where each tool fits VaultCommerce
RabbitMQ: low-volume admin notifications with complex routing keys. SQS: serverless Lambda hooks with minimal ops. Redis Streams: sub-millisecond inventory lock coordination inside a region. Kafka: order lifecycle, CDC from Postgres, and stream processing at 50k events/sec peak.
VaultCommerce rollout checklist
Before promoting changes that touch the VaultCommerce order and payment event backbone, run the staging KRaft cluster (Kafka 3.7+, Schema Registry 7.x) through a 10k events/min soak test. Compare producer request latency p99 and consumer lag per group against the pre-deploy baseline. Queues delete; logs retain and replay. Document the change in the internal topic registry, attach Grafana screenshots to the change ticket, and keep an engineer on lag dashboards for 30 minutes after production rollout — roll back the service release before altering broker-level settings if lag or under-replicated partitions spike.
# VaultCommerce — compare peak throughput during load test (Black Friday sim)
# RabbitMQ: ~12k msg/s with 4-node cluster, single queue bottleneck
# SQS: throttled at account limits around 3k/s without batching
# Kafka: 48k msg/s on vaultcommerce.orders.placed.v1 (12 partitions, RF=3)
kafka-producer-perf-test.sh --topic vaultcommerce.orders.placed.v1 \
--num-records 1000000 --record-size 512 --throughput -1 \
--producer-props bootstrap.servers=kafka-1:9092 acks=all
Quick recall
Everything you need if you only revisit this box.
- Queues delete; logs retain and replay.
- Kafka scales via partitions; classic queues scale via competing consumers.
- Match the tool to retention, replay, and throughput — not hype.
Test yourself
Answer these before moving on — recall is what makes it stick.