StreamHub's public API allows 100 requests per minute on the free tier and 1,000 on Pro. Without a distributed limiter, one abusive scraper could starve legitimate viewers — and without atomic counters, two gateway pods could each allow 100, letting 200 through.
Requirements
Functional requirements
- Enforce quotas keyed by user ID, API key, IP, or tenant.
- Support multiple tiers (Free 100/min, Pro 1K/min, Enterprise custom).
- Return HTTP 429 with
Retry-Afterand rate-limit headers. - Allow hot-reload of rules without gateway restart.
Non-functional requirements
- Latency: add less than 10 ms per request.
- Accuracy: ~99% in distributed mode; slight over-count is acceptable.
- Scale: 15K+ peak QPS across many gateway pods globally.
- Failure mode: fail-open for public reads; fail-closed for payment endpoints.
Back-of-the-envelope estimation
- 10M DAU × 50 requests/day = 500M requests/day → ~5.8K avg QPS, ~15K peak.
- Counter state:
(user_id, minute_bucket)→ integer. ~10M users × 60 buckets ≈ 600M keys/hour. - Redis
INCRis O(1); one shard handles 100K+ ops/sec; cluster scales linearly.
No tokens → 429 Too Many Requests.
Algorithm comparison
| Algorithm | Burst handling | Memory | Accuracy |
|---|---|---|---|
| Token bucket | Allows bursts up to bucket size | O(1) per key | Exact |
| Leaky bucket | Smooth fixed output rate | O(1) per key | Exact |
| Fixed window | Spike at window boundary (2×) | O(1) per key | Approximate at edges |
| Sliding window log | No burst beyond limit | O(requests) per key | Exact |
| Sliding window counter | Weighted current + previous window | O(1) per key | Near-exact |
Token bucket
Burst handlingAllows bursts up to bucket sizeMemoryO(1) per keyAccuracyExactLeaky bucket
Burst handlingSmooth fixed output rateMemoryO(1) per keyAccuracyExactFixed window
Burst handlingSpike at window boundary (2×)MemoryO(1) per keyAccuracyApproximate at edgesSliding window log
Burst handlingNo burst beyond limitMemoryO(requests) per keyAccuracyExactSliding window counter
Burst handlingWeighted current + previous windowMemoryO(1) per keyAccuracyNear-exact
Token bucket for burst tolerance; sliding window counter for accuracy without memory blow-up.
StreamHub uses token bucket for API tiers (creators need burst capacity during go-live) and sliding window counter for auth endpoints where boundary spikes matter.
Distributed implementation with Redis
-- Token bucket via Lua (atomic on one key)
local key = KEYS[1]
local rate = tonumber(ARGV[1]) -- tokens per second
local capacity = tonumber(ARGV[2])
local now = tonumber(ARGV[3])
local data = redis.call('HMGET', key, 'tokens', 'last_refill')
local tokens = tonumber(data[1]) or capacity
local last = tonumber(data[2]) or now
local elapsed = math.max(0, now - last)
tokens = math.min(capacity, tokens + elapsed * rate)
if tokens >= 1 then
tokens = tokens - 1
redis.call('HMSET', key, 'tokens', tokens, 'last_refill', now)
redis.call('EXPIRE', key, 3600)
return 1 -- allowed
end
return 0 -- denied
HTTP/1.1 429 Too Many Requests
X-RateLimit-Limit: 100
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 1727314860
Retry-After: 12
High-level architecture
AWS ALB → EKS traffic path
Component roles
- API Gateway: runs limiter middleware before routing to backend services.
- Redis Cluster: holds atomic counters; Lua scripts prevent race conditions across pods.
- Rules service: pushes tier configs via Kafka/Pub/Sub for hot-reload without restart.
- Metrics: emit allowed/denied counts per tier for alerting on abuse patterns.
Quick recall
Everything you need if you only revisit this box.
- Rate limiters protect APIs from abuse and enforce tiered quotas per identity.
- Token bucket allows bursts; sliding window counter avoids boundary spikes.
- Redis + Lua gives atomic increments across distributed gateway pods.
- Return 429 with
Retry-After,X-RateLimit-Limit, andX-RateLimit-Remaining. - Fail-open for public APIs when Redis is down; fail-closed for payments.
- Target under 10 ms overhead; counters are O(1) per request.
Test yourself
Answer these before moving on — recall is what makes it stick.