Design StreamHub's notification service: when a followed streamer goes live, notify subscribers via push, email, or SMS based on their preferences — 10M events/day, 100M deliveries/day, under 30 seconds latency for push.
Requirements
Functional requirements
- Send push, email, and SMS notifications triggered by platform events.
- User preference filters: opt-in/out per channel and notification type.
- Template rendering with dynamic content (streamer name, stream title).
- Delivery status tracking (sent, delivered, failed, bounced).
- Rate limiting: max N notifications per user per hour.
Non-functional requirements
- Latency: push under 30 s from event; email under 5 min; SMS under 1 min.
- Throughput: 100M deliveries/day ≈ 1.2K deliveries/sec average, ~5K peak.
- Reliability: at-least-once delivery with deduplication.
- Availability: 99.9% — degraded email acceptable; push is priority.
Estimation
- 10M events/day × avg 10 recipients = 100M deliveries/day.
- Push payload ~500 bytes; email ~5 KB with HTML template.
- Storage for delivery logs: 100M × 200 bytes ≈ 20 GB/day — partition by date, TTL 90 days.
Notification platform (AWS)
API design
# Internal event ingestion (not public)
POST /internal/v1/notifications:
body:
event_type: "stream.started"
actor_id: "sh_4420"
payload:
stream_id: "live_9912"
title: "Friday Night Ranked"
# User preference management (public)
GET /v1/users/me/notification-preferences
PUT /v1/users/me/notification-preferences:
body:
push: { stream_started: true, new_follower: false }
email: { stream_started: false, weekly_digest: true }
sms: { stream_started: false }
CREATE TABLE notification_preferences (
user_id BIGINT NOT NULL,
channel VARCHAR(16) NOT NULL, -- push, email, sms
event_type VARCHAR(64) NOT NULL,
enabled BOOLEAN DEFAULT true,
PRIMARY KEY (user_id, channel, event_type)
);
CREATE TABLE delivery_log (
id BIGINT PRIMARY KEY,
user_id BIGINT NOT NULL,
channel VARCHAR(16) NOT NULL,
event_type VARCHAR(64) NOT NULL,
status VARCHAR(16) NOT NULL, -- pending, sent, failed
provider_id VARCHAR(128),
created_at TIMESTAMPTZ DEFAULT now()
);
Architecture
StreamHub production architecture (AWS)
Pipeline stages
- Event ingestion: Kafka consumer receives
stream.startedevents. - Recipient resolver: Query follower list; filter by preferences and rate limits.
- Fan-out queue: Separate queues per channel (push, email, SMS) for independent scaling.
- Template engine: Render channel-specific content from Jinja/Handlebars templates.
- Provider adapters: FCM/APNs (push), SendGrid (email), Twilio (SMS) — swappable.
- Delivery tracker: Update status; retry failures with exponential backoff; DLQ after max attempts.
def process_event(event: StreamStarted):
followers = db.get_followers(event.streamer_id, limit=100_000)
for batch in chunks(followers, 1000):
for user in batch:
prefs = get_preferences(user.id, "stream.started")
if prefs.push:
push_queue.publish(Notification(user, event, channel="push"))
if prefs.email:
email_queue.publish(Notification(user, event, channel="email"))
Scaling fan-out for celebrities
| Aspect | Synchronous fan-out | Queue-based fan-out |
|---|---|---|
| Latency to producer | Minutes for 10M followers | Milliseconds — publish to queue and return |
| Failure handling | One provider timeout blocks all | Per-channel retry and DLQ |
| Scaling | Limited by single thread | Scale consumers per channel independently |
Latency to producer
Synchronous fan-outMinutes for 10M followersQueue-based fan-outMilliseconds — publish to queue and returnFailure handling
Synchronous fan-outOne provider timeout blocks allQueue-based fan-outPer-channel retry and DLQScaling
Synchronous fan-outLimited by single threadQueue-based fan-outScale consumers per channel independently
Quick recall
Everything you need if you only revisit this box.
- Fan out via separate queues per channel — push, email, SMS scale independently.
- Filter by user preferences before enqueueing; rate-limit per user per hour.
- Celebrity fan-out (10M followers) requires batching and queue sharding — never synchronous.
- Template engine renders channel-specific content; provider adapters are swappable.
- At-least-once delivery with dedup by (user_id, event_id, channel).
- Delivery log with TTL for debugging; DLQ for permanently failed notifications.
Test yourself
Answer these before moving on — recall is what makes it stick.