Design a service like sho.rt/abc12 → https://streamhub.com/stream/live_9912. Reads dominate writes 100:1 — cache every redirect, generate short codes without collisions, and handle 100M+ new URLs per month.
Requirements
Functional requirements
- Shorten a long URL; return a unique short code (6–8 base62 characters).
- Redirect short URL to original with HTTP 301 (permanent) or 302 (analytics tracking).
- Optional custom aliases for premium users.
- Optional expiration and click analytics.
Non-functional requirements
- Latency: redirect p99 under 50 ms globally.
- Availability: 99.99% — redirects are the hot path.
- Scale: 100:1 read/write ratio; 100M new URLs/month; 10B redirects/day.
- Uniqueness: No two long URLs share a code (or define last-write-wins policy).
Estimation
- 100M URLs/month × 500 bytes ≈ 50 GB/month storage — trivial for modern DBs.
- 10B redirects/day ÷ 86400 ≈ 115K redirect QPS average, ~350K peak.
- Cache hit ratio target: 90%+ → ~35K DB reads/sec at peak — still needs cache + sharding.
URL shortener (AWS)
API design
POST /v1/urls:
body: { "long_url": "https://streamhub.com/stream/live_9912", "custom_code": null }
response: { "short_code": "abc12", "short_url": "https://sho.rt/abc12" }
GET /abc12:
response: 301 Location: https://streamhub.com/stream/live_9912
CREATE TABLE url_mappings (
id BIGINT PRIMARY KEY,
short_code VARCHAR(8) UNIQUE NOT NULL,
long_url TEXT NOT NULL,
user_id BIGINT,
created_at TIMESTAMPTZ DEFAULT now(),
expires_at TIMESTAMPTZ
);
CREATE INDEX idx_short_code ON url_mappings (short_code);
Short code generation
| Aspect | Hash-based (MD5/SHA + truncate) | Counter + Base62 encode |
|---|---|---|
| Mechanism | Hash long URL, take first 7 chars | Snowflake counter → base62 string |
| Collision | Possible — must check DB and rehash | Impossible if counter is unique |
| Custom URLs | Hard — hash doesn't support aliases | Easy — reserve custom codes separately |
| StreamHub pick | Not used | Counter + base62 via Snowflake ID |
Mechanism
Hash-based (MD5/SHA + truncate)Hash long URL, take first 7 charsCounter + Base62 encodeSnowflake counter → base62 stringCollision
Hash-based (MD5/SHA + truncate)Possible — must check DB and rehashCounter + Base62 encodeImpossible if counter is uniqueCustom URLs
Hash-based (MD5/SHA + truncate)Hard — hash doesn't support aliasesCounter + Base62 encodeEasy — reserve custom codes separatelyStreamHub pick
Hash-based (MD5/SHA + truncate)Not usedCounter + Base62 encodeCounter + base62 via Snowflake ID
BASE62 = "0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz"
def to_base62(num: int) -> str:
if num == 0:
return BASE62[0]
result = []
while num:
result.append(BASE62[num % 62])
num //= 62
return "".join(reversed(result))
Architecture
StreamHub production architecture (AWS)
Component roles
- Write API: Accepts long URL, generates code via Snowflake, inserts to DB, warms cache.
- Redirect service: Separate lightweight service — cache-first, DB fallback, 301 response.
- Redis cache:
short_code → long_urlwith no TTL for active links (invalidate on expiry). - Analytics: Async — log click events to Kafka; don't block redirect path.
def redirect(short_code: str) -> str:
if long_url := redis.get(f"url:{short_code}"):
kafka.publish("click.event", {"code": short_code}) # async, non-blocking
return long_url
row = db.get_by_code(short_code)
if not row:
raise NotFound()
redis.set(f"url:{short_code}", row.long_url)
return row.long_url
Quick recall
Everything you need if you only revisit this box.
- 100:1 read/write ratio — cache every redirect; separate read-optimised service.
- Generate codes via Snowflake counter + base62, not hash truncation (collision risk).
- Redis cache: short_code → long_url; async Kafka for click analytics.
- 10B redirects/day ≈ 115K QPS — cache must handle 90%+ to keep DB manageable.
- 301 for permanent (browser-cached); 302 if you need server-side click counting.
- DB sharding by short_code prefix when mappings exceed single-node capacity.
Test yourself
Answer these before moving on — recall is what makes it stick.