PrepZone Logo
PrepZone

URL Shortener

Hash long URLs, serve redirects at scale and handle billions of clicks with cache and sharding.

Read these first

Design a service like sho.rt/abc12 → https://streamhub.com/stream/live_9912. Reads dominate writes 100:1 — cache every redirect, generate short codes without collisions, and handle 100M+ new URLs per month.

Requirements

Functional requirements

  • Shorten a long URL; return a unique short code (6–8 base62 characters).
  • Redirect short URL to original with HTTP 301 (permanent) or 302 (analytics tracking).
  • Optional custom aliases for premium users.
  • Optional expiration and click analytics.

Non-functional requirements

  • Latency: redirect p99 under 50 ms globally.
  • Availability: 99.99% — redirects are the hot path.
  • Scale: 100:1 read/write ratio; 100M new URLs/month; 10B redirects/day.
  • Uniqueness: No two long URLs share a code (or define last-write-wins policy).

Estimation

  • 100M URLs/month × 500 bytes ≈ 50 GB/month storage — trivial for modern DBs.
  • 10B redirects/day ÷ 86400 ≈ 115K redirect QPS average, ~350K peak.
  • Cache hit ratio target: 90%+ → ~35K DB reads/sec at peak — still needs cache + sharding.

URL shortener (AWS)

miss302CLIENT
Click sho.rt/…
NETWORK
CloudFront + …
DATABASE
ElastiCachehot codes
DATABASE
DynamoDBcode → URL
EXTERNAL
Destination302 redirect
CloudFront → ElastiCache → DynamoDB on redirect; API writes on create.

API design

Java
POST /v1/urls:
  body: { "long_url": "https://streamhub.com/stream/live_9912", "custom_code": null }
  response: { "short_code": "abc12", "short_url": "https://sho.rt/abc12" }

GET /abc12:
  response: 301 Location: https://streamhub.com/stream/live_9912
Java
CREATE TABLE url_mappings (
    id          BIGINT PRIMARY KEY,
    short_code  VARCHAR(8) UNIQUE NOT NULL,
    long_url    TEXT NOT NULL,
    user_id     BIGINT,
    created_at  TIMESTAMPTZ DEFAULT now(),
    expires_at  TIMESTAMPTZ
);
CREATE INDEX idx_short_code ON url_mappings (short_code);

Short code generation

AspectHash-based (MD5/SHA + truncate)Counter + Base62 encode
MechanismHash long URL, take first 7 charsSnowflake counter → base62 string
CollisionPossible — must check DB and rehashImpossible if counter is unique
Custom URLsHard — hash doesn't support aliasesEasy — reserve custom codes separately
StreamHub pickNot usedCounter + base62 via Snowflake ID
  • Mechanism

    Hash-based (MD5/SHA + truncate)Hash long URL, take first 7 chars
    Counter + Base62 encodeSnowflake counter → base62 string
  • Collision

    Hash-based (MD5/SHA + truncate)Possible — must check DB and rehash
    Counter + Base62 encodeImpossible if counter is unique
  • Custom URLs

    Hash-based (MD5/SHA + truncate)Hard — hash doesn't support aliases
    Counter + Base62 encodeEasy — reserve custom codes separately
  • StreamHub pick

    Hash-based (MD5/SHA + truncate)Not used
    Counter + Base62 encodeCounter + base62 via Snowflake ID
Java
BASE62 = "0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz"

def to_base62(num: int) -> str:
    if num == 0:
        return BASE62[0]
    result = []
    while num:
        result.append(BASE62[num % 62])
        num //= 62
    return "".join(reversed(result))

Architecture

StreamHub production architecture (AWS)

HTTPSstaticmissAPICLIENT
Mobile / WebStreamHub cli…
NETWORK
Route 53GeoDNS routing
NETWORK
CloudFrontCDN + WAF edge
NETWORK
AWS ALBTLS terminati…
NETWORK
API GatewayJWT · rate li…
STORAGE
Amazon S3media origin
COMPUTE
Amazon EKSAPI · auth · …
DATABASE
ElastiCachesessions · ho…
DATABASE
RDS Postgresprimary + rep…
INTEGRATION
Amazon MSKdomain events
ANALYTICS
OpenSearchstream discov…
OPS
CloudWatchmetrics · X-R…
End-to-end path from user to data — reference this when placing any new service.

Component roles

  • Write API: Accepts long URL, generates code via Snowflake, inserts to DB, warms cache.
  • Redirect service: Separate lightweight service — cache-first, DB fallback, 301 response.
  • Redis cache: short_code → long_url with no TTL for active links (invalidate on expiry).
  • Analytics: Async — log click events to Kafka; don't block redirect path.
Java
def redirect(short_code: str) -> str:
    if long_url := redis.get(f"url:{short_code}"):
        kafka.publish("click.event", {"code": short_code})  # async, non-blocking
        return long_url
    row = db.get_by_code(short_code)
    if not row:
        raise NotFound()
    redis.set(f"url:{short_code}", row.long_url)
    return row.long_url

Quick recall

Everything you need if you only revisit this box.

  • 100:1 read/write ratio — cache every redirect; separate read-optimised service.
  • Generate codes via Snowflake counter + base62, not hash truncation (collision risk).
  • Redis cache: short_code → long_url; async Kafka for click analytics.
  • 10B redirects/day ≈ 115K QPS — cache must handle 90%+ to keep DB manageable.
  • 301 for permanent (browser-cached); 302 if you need server-side click counting.
  • DB sharding by short_code prefix when mappings exceed single-node capacity.

Test yourself

Answer these before moving on — recall is what makes it stick.