Design StreamHub's activity feed: when a followed streamer goes live or posts a clip, subscribers see it in their feed within seconds. Celebrities with 10M followers break naive push fan-out — the hybrid model solves this.
Requirements
Functional requirements
- Home feed: chronological or ranked list of posts from followed streamers.
- Publish post: streamer creates a clip highlight, schedule announcement, or go-live event.
- Follow/unfollow streamers; feed updates accordingly.
- Like and comment on feed items (basic engagement).
- Feed pagination: infinite scroll with cursor.
Non-functional requirements
- Latency: feed load under 300 ms p99.
- Scale: 100M DAU; avg 200 follows per user; 10M posts/day.
- Fan-out: Celebrity streamers with up to 10M followers.
- Freshness: New posts appear within 5 seconds for active users.
- Availability: 99.9% — feed is the home screen.
Out of scope: ML ranking model internals, ads insertion, stories/ephemeral content.
Estimation
- 100M DAU × 10 feed loads/day = 1B feed reads/day ≈ 12K QPS average, ~40K peak.
- 10M posts/day × 200 avg followers = 2B fan-out writes/day if pure push — 23K writes/sec.
- Celebrity post (10M followers): 10M writes in one fan-out — must not be synchronous.
- Feed cache: 200 post IDs × 8 bytes × 100M users = 160 GB if fully materialised — needs tiering.
News feed architecture (AWS)
API design
GET /v1/feed?cursor=xxx&limit=20
POST /v1/posts:
body: { "type": "clip", "clip_id": "clip_9912", "caption": "Insane play!" }
POST /v1/users/{id}/follow
DELETE /v1/users/{id}/follow
GET /v1/posts/{id}/comments?cursor=xxx
CREATE TABLE posts (
id BIGINT PRIMARY KEY,
author_id BIGINT NOT NULL,
type VARCHAR(16) NOT NULL,
content_ref VARCHAR(64), -- clip_id, stream_id
caption TEXT,
created_at TIMESTAMPTZ NOT NULL
);
CREATE INDEX idx_posts_author ON posts (author_id, created_at DESC);
-- Fan-out table (push model): pre-computed feed entries
CREATE TABLE feed_timeline (
user_id BIGINT NOT NULL,
post_id BIGINT NOT NULL,
author_id BIGINT NOT NULL,
created_at TIMESTAMPTZ NOT NULL,
PRIMARY KEY (user_id, created_at DESC, post_id)
);
High-level architecture
StreamHub production architecture (AWS)
| Aspect | Pull model (read-time merge) | Push model (write-time fan-out) |
|---|---|---|
| Feed load | Fetch followed authors' recent posts; merge + sort | Read pre-built timeline table — O(1) |
| Post publish | O(1) — just write the post | O(followers) — write to each follower's timeline |
| Best for | Users who follow many; celebrities (write side) | Users with few follows; regular streamers |
| Celebrity problem | Celebrity post is O(1) on write | 10M fan-out writes — must hybridise |
Feed load
Pull model (read-time merge)Fetch followed authors' recent posts; merge + sortPush model (write-time fan-out)Read pre-built timeline table — O(1)Post publish
Pull model (read-time merge)O(1) — just write the postPush model (write-time fan-out)O(followers) — write to each follower's timelineBest for
Pull model (read-time merge)Users who follow many; celebrities (write side)Push model (write-time fan-out)Users with few follows; regular streamersCelebrity problem
Pull model (read-time merge)Celebrity post is O(1) on writePush model (write-time fan-out)10M fan-out writes — must hybridise
Hybrid fan-out strategy
StreamHub hybrid approach
- Regular streamers (< 100K followers): Push fan-out on post — write to each follower's
feed_timeline. - Celebrities (> 100K followers): Skip push — store post only; followers merge celebrity posts at read time (pull).
- Active users: Pre-computed timeline cached in Redis (200 most recent post IDs).
- Inactive users: Compute feed on next login; no wasted fan-out writes.
Deep dive: feed read path
def get_feed(user_id: int, cursor: str, limit: int) -> FeedPage:
if cached := redis.get(f"feed:{user_id}"):
return paginate(cached, cursor, limit)
# Pull: fetch from pre-built timeline (push entries)
push_posts = db.query_timeline(user_id, cursor, limit)
# Pull: merge celebrity posts from followed celebrities
celebrities = get_followed_celebrities(user_id)
celeb_posts = db.query_recent_posts(celebrities, since=cursor)
merged = merge_and_sort(push_posts, celeb_posts)[:limit]
redis.setex(f"feed:{user_id}", 60, merged)
return merged
Deep dive: ranking (optional extension)
Chronological is the MVP. For ranked feeds, pre-compute scores in a batch job (Spark/Flink) or use a lightweight online ranker:
| Signal | Weight | Source |
|---|---|---|
| Recency | High | Post created_at — exponential decay |
| Engagement | Medium | Likes + comments in first hour |
| Relationship | High | Close friends / frequent watchers boosted |
| Content type | Low | Clips ranked above schedule announcements |
Recency
WeightHighSourcePost created_at — exponential decayEngagement
WeightMediumSourceLikes + comments in first hourRelationship
WeightHighSourceClose friends / frequent watchers boostedContent type
WeightLowSourceClips ranked above schedule announcements
Quick recall
Everything you need if you only revisit this box.
- Pull: merge on read (cheap writes, expensive reads) — good for celebrities.
- Push: fan-out on write (expensive writes, cheap reads) — good for regular users.
- Hybrid: push if followers < 100K; pull merge for celebrities at read time.
- Skip fan-out for inactive users — compute on next login.
- Cache top 200 post IDs in Redis per active user for sub-300 ms feed loads.
- 1B feed reads/day ≈ 12K QPS — cache and pre-computation are mandatory.
Test yourself
Answer these before moving on — recall is what makes it stick.