Why this matters
- VaultCommerce's Redis footprint grew past 32 GB of hot keys — Cluster spreads slots across nodes; Sentinel alone cannot shard.
- Hot keys (a viral SKU counter) saturate one slot even in Cluster — application-level sharding or local caching still matters.
- Production hardening spans TLS, ACLs, memory limits, and slow-log monitoring — not just topology choice.
Primary-replica with Sentinel
One primary accepts writes; replicas serve reads and stand by for promotion. Three or more Sentinel processes quorum-vote on failures.
# sentinel.conf
sentinel monitor vaultcommerce-redis 10.0.1.10 6379 2
sentinel down-after-milliseconds vaultcommerce-redis 5000
sentinel failover-timeout vaultcommerce-redis 60000
SENTINEL masters
SENTINEL replicas vaultcommerce-redis
SENTINEL get-master-addr-by-name vaultcommerce-redis
Clients must resolve the current primary after failover — use a Sentinel-aware driver or managed endpoint.
Sentinel vs Cluster
- Sentinel — HA for a single dataset; one primary at a time; simpler ops.
- Cluster — 16,384 hash slots sharded across primaries; multi-master writes.
- Both — use replicas for read scaling; Cluster adds slot migration for growth.
Redis Cluster essentials
Keys map to slots via CRC16(key) mod 16384. Multi-key commands require {hash-tag} in the same slot:
CLUSTER INFO
CLUSTER SLOTS
SET product:{8842}:views 100
SET product:{8842}:meta "{\"name\":\"Trail Pack\"}"
MGET product:{8842}:views product:{8842}:meta
VaultCommerce tags flash-sale keys with {sale:FLASH10} so related counters land on one slot for atomic Lua scripts.
Hot-key mitigation
A viral SKU can overload one slot/node despite Cluster:
# Spread read counters across N shard keys, aggregate on read
import random
SHARDS = 8
def incr_views(sku: str) -> None:
shard = random.randint(0, SHARDS - 1)
r.incr(f"views:{sku}:{shard}")
def total_views(sku: str) -> int:
keys = [f"views:{sku}:{i}" for i in range(SHARDS)]
return sum(int(v or 0) for v in r.mget(keys))
Local in-process caching for the hottest keys adds another layer when Redis itself becomes the bottleneck.
Production checklist
redis-cli ACL LIST
redis-cli CONFIG GET requirepass
redis-cli SLOWLOG GET 10
redis-cli INFO memory | grep used_memory_human
Quick recall
Everything you need if you only revisit this box.
- Sentinel: automatic failover for primary-replica — one writable primary.
- Cluster: 16,384 slots across nodes — scale out when RAM or QPS exceeds one host.
- Hash tags
{tag}co-locate related keys for multi-key operations. - Hot keys need app-level sharding or local cache even with Cluster.
- Enable ACLs, TLS, slow-log alerts, and memory limits in production.
- VaultCommerce uses Cluster for large catalog cache; Sentinel suffices for sessions-only tiers.
Test yourself
Answer these before moving on — recall is what makes it stick.