PrepZone Logo
PrepZone

Rate Limiter Design

Token bucket, sliding window and distributed counters that protect APIs from abuse.

StreamHub's public API allows 100 requests per minute on the free tier and 1,000 on Pro. Without a distributed limiter, one abusive scraper could starve legitimate viewers — and without atomic counters, two gateway pods could each allow 100, letting 200 through.

Requirements

Functional requirements

  • Enforce quotas keyed by user ID, API key, IP, or tenant.
  • Support multiple tiers (Free 100/min, Pro 1K/min, Enterprise custom).
  • Return HTTP 429 with Retry-After and rate-limit headers.
  • Allow hot-reload of rules without gateway restart.

Non-functional requirements

  • Latency: add less than 10 ms per request.
  • Accuracy: ~99% in distributed mode; slight over-count is acceptable.
  • Scale: 15K+ peak QPS across many gateway pods globally.
  • Failure mode: fail-open for public reads; fail-closed for payment endpoints.

Back-of-the-envelope estimation

  • 10M DAU × 50 requests/day = 500M requests/day → ~5.8K avg QPS, ~15K peak.
  • Counter state: (user_id, minute_bucket) → integer. ~10M users × 60 buckets ≈ 600M keys/hour.
  • Redis INCR is O(1); one shard handles 100K+ ops/sec; cluster scales linearly.
Request
Token bucket100 req/min refill
API

No tokens → 429 Too Many Requests.

Tokens refill at a fixed rate; each request consumes one token.

Algorithm comparison

AlgorithmBurst handlingMemoryAccuracy
Token bucketAllows bursts up to bucket sizeO(1) per keyExact
Leaky bucketSmooth fixed output rateO(1) per keyExact
Fixed windowSpike at window boundary (2×)O(1) per keyApproximate at edges
Sliding window logNo burst beyond limitO(requests) per keyExact
Sliding window counterWeighted current + previous windowO(1) per keyNear-exact
  • Token bucket

    Burst handlingAllows bursts up to bucket size
    MemoryO(1) per key
    AccuracyExact
  • Leaky bucket

    Burst handlingSmooth fixed output rate
    MemoryO(1) per key
    AccuracyExact
  • Fixed window

    Burst handlingSpike at window boundary (2×)
    MemoryO(1) per key
    AccuracyApproximate at edges
  • Sliding window log

    Burst handlingNo burst beyond limit
    MemoryO(requests) per key
    AccuracyExact
  • Sliding window counter

    Burst handlingWeighted current + previous window
    MemoryO(1) per key
    AccuracyNear-exact

Token bucket for burst tolerance; sliding window counter for accuracy without memory blow-up.

StreamHub uses token bucket for API tiers (creators need burst capacity during go-live) and sliding window counter for auth endpoints where boundary spikes matter.

Distributed implementation with Redis

Java
-- Token bucket via Lua (atomic on one key)
local key = KEYS[1]
local rate = tonumber(ARGV[1])      -- tokens per second
local capacity = tonumber(ARGV[2])
local now = tonumber(ARGV[3])

local data = redis.call('HMGET', key, 'tokens', 'last_refill')
local tokens = tonumber(data[1]) or capacity
local last = tonumber(data[2]) or now

local elapsed = math.max(0, now - last)
tokens = math.min(capacity, tokens + elapsed * rate)

if tokens >= 1 then
  tokens = tokens - 1
  redis.call('HMSET', key, 'tokens', tokens, 'last_refill', now)
  redis.call('EXPIRE', key, 3600)
  return 1   -- allowed
end
return 0       -- denied
Java
HTTP/1.1 429 Too Many Requests
X-RateLimit-Limit: 100
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 1727314860
Retry-After: 12

High-level architecture

AWS ALB → EKS traffic path

TLSCLIENT
UsersHTTPS requests
NETWORK
Route 53alias → ALB
NETWORK
AWS ALBtarget group
COMPUTE
EKS pod 1streamhub-api
COMPUTE
EKS pod 2streamhub-api
COMPUTE
EKS pod NHPA scaled
L7 ALB terminates TLS, health-checks targets, and fans out to stateless pods.

Component roles

  • API Gateway: runs limiter middleware before routing to backend services.
  • Redis Cluster: holds atomic counters; Lua scripts prevent race conditions across pods.
  • Rules service: pushes tier configs via Kafka/Pub/Sub for hot-reload without restart.
  • Metrics: emit allowed/denied counts per tier for alerting on abuse patterns.

Quick recall

Everything you need if you only revisit this box.

  • Rate limiters protect APIs from abuse and enforce tiered quotas per identity.
  • Token bucket allows bursts; sliding window counter avoids boundary spikes.
  • Redis + Lua gives atomic increments across distributed gateway pods.
  • Return 429 with Retry-After, X-RateLimit-Limit, and X-RateLimit-Remaining.
  • Fail-open for public APIs when Redis is down; fail-closed for payments.
  • Target under 10 ms overhead; counters are O(1) per request.

Test yourself

Answer these before moving on — recall is what makes it stick.