PrepZone Logo
PrepZone

Web Crawler Design

Discover, fetch and index billions of pages with politeness, deduplication and distributed workers.

Read these first

Design StreamIndex — StreamHub's web crawler that discovers streamer profile pages across the internet, extracts metadata, and feeds the search index. Billions of pages, distributed workers, and strict politeness constraints.

Requirements

Functional requirements

  • Given a seed URL, discover and fetch linked pages up to N hops deep.
  • Extract structured data: title, description, Open Graph tags, links.
  • Store raw HTML and parsed metadata for the search indexer.
  • Respect robots.txt and crawl-delay directives.
  • Deduplicate URLs — never fetch the same canonical URL twice.

Non-functional requirements

  • Scale: Crawl 1B pages/month; 10B total indexed pages.
  • Politeness: Max 1 request/sec per domain (configurable).
  • Freshness: Re-crawl popular pages every 24 h; long-tail every 30 days.
  • Availability: 99% — crawler downtime delays index freshness, not user-facing latency.
  • Storage: Raw HTML averages 50 KB/page → 10B × 50 KB = 500 TB compressed.

Out of scope: JavaScript rendering (headless browser farm is v2), login-gated pages, real-time crawl.

Estimation

  • 1B pages/month ÷ 30 days ÷ 86400 ≈ 400 pages/sec sustained crawl rate.
  • At 1 req/sec/domain limit: need 400+ domains crawled in parallel simultaneously.
  • URL frontier: 10B URLs × 100 bytes ≈ 1 TB for URL store (Bloom filter + disk).
  • Bandwidth: 400 pages/sec × 50 KB ≈ 20 MB/sec ≈ 1.7 TB/day ingress.

Web crawler architecture

COMPUTE
EventBridgescheduler
INTEGRATION
SQS frontierURL queue
COMPUTE
Crawler workersLambda / EKS
STORAGE
Amazon S3raw pages
ANALYTICS
OpenSearchindex
EventBridge scheduler → SQS → workers → S3 raw HTML → OpenSearch index.

API design

Java
# Internal crawl management
POST /internal/v1/crawl/seeds:
  body: { "urls": ["https://example.com/streamer/jane"], "priority": "high" }

GET  /internal/v1/crawl/status/{job_id}
POST /internal/v1/crawl/pause/{domain}     # emergency politeness stop

# Output to search indexer
Kafka topic: crawl.pages.parsed
  payload: { "url": "...", "title": "...", "links": [...], "fetched_at": "..." }
Java
CREATE TABLE url_frontier (
    url_hash     VARCHAR(64) PRIMARY KEY,
    url          TEXT NOT NULL,
    domain       VARCHAR(256) NOT NULL,
    priority     INT DEFAULT 0,
    next_fetch   TIMESTAMPTZ NOT NULL,
    fetch_count  INT DEFAULT 0,
    status       VARCHAR(16) DEFAULT 'pending'
);
CREATE INDEX idx_frontier_domain ON url_frontier (domain, next_fetch);

CREATE TABLE crawl_history (
    url_hash     VARCHAR(64) PRIMARY KEY,
    http_status  INT,
    content_hash VARCHAR(64),     -- detect content changes
    fetched_at   TIMESTAMPTZ NOT NULL,
    html_s3_key  VARCHAR(256)
);

High-level architecture

Crawler components

  • URL frontier: Priority queue of URLs to fetch — partitioned by domain.
  • Fetcher workers: Pull URLs from frontier; HTTP GET with timeout; store raw HTML in S3.
  • Robots.txt cache: Pre-fetch and cache robots.txt per domain (TTL 24 h).
  • Politeness scheduler: Token bucket per domain — max 1 req/sec default.
  • Link extractor: Parse HTML; extract <a href> links; normalise and deduplicate URLs.
  • Duplication detector: Bloom filter for fast "seen?" check; disk-backed URL hash store for certainty.
  • Content processor: Parse metadata; publish to Kafka for search indexer consumption.

Deep dive: URL frontier and politeness

AspectBFS frontier (FIFO)Priority frontier
OrderingFirst discovered, first fetchedHigh-priority domains/URLs first
Use caseBroad discoveryFreshness for important streamer sites
RiskLow-priority pages delay high-priorityStarvation without aging boost
StreamHubDefault for general crawlPartner domains and trending streamers
  • Ordering

    BFS frontier (FIFO)First discovered, first fetched
    Priority frontierHigh-priority domains/URLs first
  • Use case

    BFS frontier (FIFO)Broad discovery
    Priority frontierFreshness for important streamer sites
  • Risk

    BFS frontier (FIFO)Low-priority pages delay high-priority
    Priority frontierStarvation without aging boost
  • StreamHub

    BFS frontier (FIFO)Default for general crawl
    Priority frontierPartner domains and trending streamers
Java
def schedule_fetch(url: str, domain: str):
    if bloom_filter.might_contain(url):
        if db.url_already_fetched(url):
            return  # definite duplicate
    if not robots_cache.is_allowed(domain, url):
        return
    if rate_limiter.can_request(domain):
        frontier.enqueue(url, priority=compute_priority(url))
        rate_limiter.consume(domain)
    else:
        frontier.enqueue_delayed(url, delay=rate_limiter.retry_after(domain))

Deep dive: deduplication and content change detection

Avoiding wasted work

  • URL normalisation: Lowercase, remove fragments, resolve relative paths, canonicalise www..
  • Bloom filter: 10B URLs with 1% false positive rate ≈ 12 GB bitmap — fast in-memory check.
  • Content hash: SHA-256 of page body; skip re-indexing if hash unchanged since last crawl.
  • Near-duplicate detection: Simhash for pages with minor changes (ads, timestamps).

Quick recall

Everything you need if you only revisit this box.

  • URL frontier as priority queue; partition by domain for politeness scheduling.
  • Token bucket per domain: default 1 req/sec; cache robots.txt with 24 h TTL.
  • Bloom filter for fast dedup; disk-backed hash store for certainty.
  • Store raw HTML in S3; publish parsed metadata to Kafka for search indexer.
  • 1B pages/month ≈ 400 pages/sec — distribute fetchers with domain-aware partitioning.
  • Content hash skips re-indexing when page body unchanged since last crawl.

Test yourself

Answer these before moving on — recall is what makes it stick.