When StreamHub's catalog API crossed 10K reads per second, Postgres became the bottleneck long before CPU did. A Redis layer in front of stream metadata cut p99 latency from 180 ms to under 5 ms — but only after picking the right cache pattern and invalidation strategy.
Why caching matters at scale
What a cache buys you
- Latency: Memory reads are microseconds; disk-backed DB reads are milliseconds.
- Throughput: Offloading reads frees database connections for writes.
- Cost: Fewer DB replicas needed when 80–90% of reads hit cache.
- Resilience: A warm cache absorbs traffic spikes during viral moments.
Cache-aside with ElastiCache
Core cache patterns
| Pattern | Write path | Best for |
|---|---|---|
| Cache-aside (lazy load) | App reads cache; on miss, loads DB and populates | General read-heavy workloads |
| Read-through | Cache library fetches from DB on miss transparently | When app code should stay cache-agnostic |
| Write-through | Every write updates cache and DB synchronously | Strong consistency on reads after writes |
| Write-back | Write to cache; async flush to DB later | Very high write throughput; brief loss risk |
| Write-around | Write directly to DB, skip cache | Write-heavy, rarely read data (audit logs) |
Cache-aside (lazy load)
Write pathApp reads cache; on miss, loads DB and populatesBest forGeneral read-heavy workloadsRead-through
Write pathCache library fetches from DB on miss transparentlyBest forWhen app code should stay cache-agnosticWrite-through
Write pathEvery write updates cache and DB synchronouslyBest forStrong consistency on reads after writesWrite-back
Write pathWrite to cache; async flush to DB laterBest forVery high write throughput; brief loss riskWrite-around
Write pathWrite directly to DB, skip cacheBest forWrite-heavy, rarely read data (audit logs)
Pick the pattern that matches your read/write ratio and consistency needs.
StreamHub uses cache-aside for stream metadata: the API checks Redis first, falls back to Postgres on miss, and sets a TTL of 300 seconds.
def get_stream(stream_id: str) -> StreamMeta:
cached = redis.get(f"stream:{stream_id}")
if cached:
return StreamMeta.parse(cached)
row = db.query("SELECT * FROM streams WHERE id = %s", stream_id)
redis.setex(f"stream:{stream_id}", 300, row.to_json())
return row
Cache layer topology
Layers from browser to database
- Browser cache: HTTP
Cache-ControlandETagheaders — free performance for static assets. - CDN: Edge caches for thumbnails, JS bundles, and API GET responses at PoPs worldwide.
- API gateway cache: Idempotent GET caching at the gateway for public catalog endpoints.
- App-level cache: In-process Caffeine/Guava — nanosecond access, per-pod only.
- Distributed cache: Redis or Memcached shared across all StreamHub API pods.
- Database buffer pool: InnoDB buffer pool or Postgres
shared_buffers— automatic last line.
Eviction and TTL policies
| Aspect | LRU (Least Recently Used) | LFU (Least Frequently Used) |
|---|---|---|
| Eviction rule | Drop the key accessed longest ago | Drop the key accessed least often |
| Best for | General workloads with shifting hot sets | Stable hot keys (trending streams) |
| Weakness | One-time scans can evict real hot keys | New keys start cold and may be evicted early |
| StreamHub use | Default for user session data | Trending stream leaderboard cache |
Eviction rule
LRU (Least Recently Used)Drop the key accessed longest agoLFU (Least Frequently Used)Drop the key accessed least oftenBest for
LRU (Least Recently Used)General workloads with shifting hot setsLFU (Least Frequently Used)Stable hot keys (trending streams)Weakness
LRU (Least Recently Used)One-time scans can evict real hot keysLFU (Least Frequently Used)New keys start cold and may be evicted earlyStreamHub use
LRU (Least Recently Used)Default for user session dataLFU (Least Frequently Used)Trending stream leaderboard cache
Always set TTL with jitter — if every key expires at exactly 300 s, a thundering herd hits the database simultaneously.
# Redis TTL with jitter (pseudo-config)
default_ttl_seconds: 300
jitter_range_seconds: 30 # actual TTL = 270–330 s
Cache invalidation strategies
Keeping cache fresh
- Delete on write: Simplest — on DB update, delete the cache key; next read repopulates.
- TTL-only: Accept staleness up to N seconds; good for low-risk catalog data.
- Event-driven: CDC from Postgres → Kafka → cache invalidator service (Debezium pattern).
- Version tags: Store a version integer; bump on write; readers detect stale entries.
StreamHub cache-aside in production
StreamHub production architecture (AWS)
For StreamHub's live viewer count, we combine cache-aside with a 10-second TTL — stale by a few viewers is acceptable, but a DB query per page load is not.
Quick recall
Everything you need if you only revisit this box.
- Cache-aside is the default: app manages reads, writes, and population on miss.
- Write-through guarantees fresh reads after writes at the cost of slower writes.
- Layer caches from browser → CDN → gateway → Redis → DB buffer pool.
- TTL with jitter prevents thundering herd on simultaneous expiry.
- Invalidate on write for strong consistency; TTL-only for eventually fresh data.
- Hot keys need replication or background refresh — eviction policy alone won't save you.
- Always state what happens when the cache is down: fail open to DB or circuit-break.
Test yourself
Answer these before moving on — recall is what makes it stick.
- Explain caching strategies — cache-aside, read/write-through, write-behind, TTL, invalidation.
- Given an access pattern, walk through choosing between SQL / KV / Document / Wide-column / Search / Graph / Vector.
- Walk through the scalability journey — vertical, horizontal, sharding, replication, async, CDN.