PrepZone Logo
PrepZone

Distributed Cache Deep Dive

Redis clusters, eviction policies, hot keys and cache stampede prevention at scale.

Read these first

StreamHub's Redis cluster holds stream metadata, session tokens, and rate-limit counters for every API pod. When one shard goes hot during a viral stream, the difference between a well-tuned cluster and a naive single-node Redis is the difference between 5 ms and timeout errors.

Redis cluster architecture

Hash ring0 … 2^32-1
Node Akeys 0–25%
Node Bkeys 25–50%
Node Ckeys 50–75%
Node Dkeys 75–100%

New node E joins → only ~1/N keys move, not all keys.

Adding a node only remaps keys adjacent to it on the ring.

Redis Cluster essentials

  • 16384 hash slots partitioned across primary nodes; each key maps to one slot.
  • Primary + replica pairs per shard — replica promotes on primary failure.
  • Client-side routing: Redis clients follow MOVED/ASK redirects to the correct shard.
  • No cross-slot transactions unless keys share the same hash tag {user}:profile and {user}:settings.
Java
# ElastiCache Redis cluster (pseudo-config)
engine: redis-7
node_type: cache.r7g.xlarge
num_shards: 12
replicas_per_shard: 1
maxmemory_policy: allkeys-lru
cluster_mode: enabled

Eviction policies in practice

PolicyBehaviourWhen to use
allkeys-lruEvict any key by LRU when memory fullGeneral cache — StreamHub metadata
volatile-lruEvict only keys with TTL setMixed cache + session store
allkeys-lfuEvict least frequently used keysStable hot sets (leaderboards)
noevictionReturn errors when memory fullRate limit counters — never silently drop
  • allkeys-lru

    BehaviourEvict any key by LRU when memory full
    When to useGeneral cache — StreamHub metadata
  • volatile-lru

    BehaviourEvict only keys with TTL set
    When to useMixed cache + session store
  • allkeys-lfu

    BehaviourEvict least frequently used keys
    When to useStable hot sets (leaderboards)
  • noeviction

    BehaviourReturn errors when memory full
    When to useRate limit counters — never silently drop

Hot key problem

A single key receiving millions of reads per second overwhelms one shard regardless of consistent hashing.

Hot key mitigations

  • Local cache in app: Brief in-process cache (1–5 s) in front of Redis for celebrity stream IDs.
  • Key replication: Store the same value under stream:123:{0..9} and pick a random suffix on read.
  • Read replicas: Route hot reads to replica nodes; primary handles writes only.
  • Pre-warming: Background job refreshes hot keys before TTL expiry.
Java
# Hot key: random suffix sharding on read
def get_hot_stream(stream_id: str) -> StreamMeta:
    suffix = random.randint(0, 9)
    key = f"stream:{stream_id}:{suffix}"
    data = redis.get(key)
    if not data:
        data = load_and_replicate(stream_id, replicas=10)
    return StreamMeta.parse(data)

Cache stampede prevention

Java
import threading

_locks: dict[str, threading.Lock] = {}

def get_with_single_flight(key: str) -> str:
    if val := redis.get(key):
        return val
    lock = _locks.setdefault(key, threading.Lock())
    with lock:
        if val := redis.get(key):          # double-check after acquiring lock
            return val
        val = db.load(key)
        redis.setex(key, 300, val)
        return val

Probabilistic early expiration refreshes keys in the background before they expire — only one pod "wins" the refresh lottery per window.

Monitoring and operations

AspectHealthy cluster signalProblem signal
Hit ratio> 90% for metadata cache< 70% — TTL too short or wrong keys cached
Latency p99< 2 ms per GET> 10 ms — hot shard or network issue
MemorySteady with LRU churnFlat at max — eviction not keeping up
ConnectionsStable per podSpike — connection leak or missing pooling
  • Hit ratio

    Healthy cluster signal> 90% for metadata cache
    Problem signal< 70% — TTL too short or wrong keys cached
  • Latency p99

    Healthy cluster signal< 2 ms per GET
    Problem signal> 10 ms — hot shard or network issue
  • Memory

    Healthy cluster signalSteady with LRU churn
    Problem signalFlat at max — eviction not keeping up
  • Connections

    Healthy cluster signalStable per pod
    Problem signalSpike — connection leak or missing pooling

Quick recall

Everything you need if you only revisit this box.

  • Redis Cluster uses 16384 hash slots with consistent hashing across primaries.
  • Hot keys break sharding — replicate locally or shard the key with random suffixes.
  • Cache stampede: mitigate with single-flight locks, jittered TTL, or background refresh.
  • allkeys-lru for general cache; noeviction only when losing keys is unacceptable.
  • Monitor hit ratio, p99 latency, and memory — not just uptime.
  • Hash tags {user} keep related keys on the same shard for multi-key operations.

Test yourself

Answer these before moving on — recall is what makes it stick.