Design StreamIndex — StreamHub's web crawler that discovers streamer profile pages across the internet, extracts metadata, and feeds the search index. Billions of pages, distributed workers, and strict politeness constraints.
Requirements
Functional requirements
- Given a seed URL, discover and fetch linked pages up to N hops deep.
- Extract structured data: title, description, Open Graph tags, links.
- Store raw HTML and parsed metadata for the search indexer.
- Respect robots.txt and crawl-delay directives.
- Deduplicate URLs — never fetch the same canonical URL twice.
Non-functional requirements
- Scale: Crawl 1B pages/month; 10B total indexed pages.
- Politeness: Max 1 request/sec per domain (configurable).
- Freshness: Re-crawl popular pages every 24 h; long-tail every 30 days.
- Availability: 99% — crawler downtime delays index freshness, not user-facing latency.
- Storage: Raw HTML averages 50 KB/page → 10B × 50 KB = 500 TB compressed.
Out of scope: JavaScript rendering (headless browser farm is v2), login-gated pages, real-time crawl.
Estimation
- 1B pages/month ÷ 30 days ÷ 86400 ≈ 400 pages/sec sustained crawl rate.
- At 1 req/sec/domain limit: need 400+ domains crawled in parallel simultaneously.
- URL frontier: 10B URLs × 100 bytes ≈ 1 TB for URL store (Bloom filter + disk).
- Bandwidth: 400 pages/sec × 50 KB ≈ 20 MB/sec ≈ 1.7 TB/day ingress.
Web crawler architecture
API design
# Internal crawl management
POST /internal/v1/crawl/seeds:
body: { "urls": ["https://example.com/streamer/jane"], "priority": "high" }
GET /internal/v1/crawl/status/{job_id}
POST /internal/v1/crawl/pause/{domain} # emergency politeness stop
# Output to search indexer
Kafka topic: crawl.pages.parsed
payload: { "url": "...", "title": "...", "links": [...], "fetched_at": "..." }
CREATE TABLE url_frontier (
url_hash VARCHAR(64) PRIMARY KEY,
url TEXT NOT NULL,
domain VARCHAR(256) NOT NULL,
priority INT DEFAULT 0,
next_fetch TIMESTAMPTZ NOT NULL,
fetch_count INT DEFAULT 0,
status VARCHAR(16) DEFAULT 'pending'
);
CREATE INDEX idx_frontier_domain ON url_frontier (domain, next_fetch);
CREATE TABLE crawl_history (
url_hash VARCHAR(64) PRIMARY KEY,
http_status INT,
content_hash VARCHAR(64), -- detect content changes
fetched_at TIMESTAMPTZ NOT NULL,
html_s3_key VARCHAR(256)
);
High-level architecture
Crawler components
- URL frontier: Priority queue of URLs to fetch — partitioned by domain.
- Fetcher workers: Pull URLs from frontier; HTTP GET with timeout; store raw HTML in S3.
- Robots.txt cache: Pre-fetch and cache robots.txt per domain (TTL 24 h).
- Politeness scheduler: Token bucket per domain — max 1 req/sec default.
- Link extractor: Parse HTML; extract
<a href>links; normalise and deduplicate URLs. - Duplication detector: Bloom filter for fast "seen?" check; disk-backed URL hash store for certainty.
- Content processor: Parse metadata; publish to Kafka for search indexer consumption.
Deep dive: URL frontier and politeness
| Aspect | BFS frontier (FIFO) | Priority frontier |
|---|---|---|
| Ordering | First discovered, first fetched | High-priority domains/URLs first |
| Use case | Broad discovery | Freshness for important streamer sites |
| Risk | Low-priority pages delay high-priority | Starvation without aging boost |
| StreamHub | Default for general crawl | Partner domains and trending streamers |
Ordering
BFS frontier (FIFO)First discovered, first fetchedPriority frontierHigh-priority domains/URLs firstUse case
BFS frontier (FIFO)Broad discoveryPriority frontierFreshness for important streamer sitesRisk
BFS frontier (FIFO)Low-priority pages delay high-priorityPriority frontierStarvation without aging boostStreamHub
BFS frontier (FIFO)Default for general crawlPriority frontierPartner domains and trending streamers
def schedule_fetch(url: str, domain: str):
if bloom_filter.might_contain(url):
if db.url_already_fetched(url):
return # definite duplicate
if not robots_cache.is_allowed(domain, url):
return
if rate_limiter.can_request(domain):
frontier.enqueue(url, priority=compute_priority(url))
rate_limiter.consume(domain)
else:
frontier.enqueue_delayed(url, delay=rate_limiter.retry_after(domain))
Deep dive: deduplication and content change detection
Avoiding wasted work
- URL normalisation: Lowercase, remove fragments, resolve relative paths, canonicalise
www.. - Bloom filter: 10B URLs with 1% false positive rate ≈ 12 GB bitmap — fast in-memory check.
- Content hash: SHA-256 of page body; skip re-indexing if hash unchanged since last crawl.
- Near-duplicate detection: Simhash for pages with minor changes (ads, timestamps).
Quick recall
Everything you need if you only revisit this box.
- URL frontier as priority queue; partition by domain for politeness scheduling.
- Token bucket per domain: default 1 req/sec; cache robots.txt with 24 h TTL.
- Bloom filter for fast dedup; disk-backed hash store for certainty.
- Store raw HTML in S3; publish parsed metadata to Kafka for search indexer.
- 1B pages/month ≈ 400 pages/sec — distribute fetchers with domain-aware partitioning.
- Content hash skips re-indexing when page body unchanged since last crawl.
Test yourself
Answer these before moving on — recall is what makes it stick.