Why this matters
- AWS, Slack, and Discord use cells to contain blast radius — a bad deploy or database corruption in Cell A does not take down Cells B through Z.
- StreamHub's 2024 us-east-1 incident affected all users because they shared one monolithic database. Cells would have limited impact to ~5% of users.
- Cells trade operational complexity (managing N independent stacks) for resilience at scale.
Cell architecture components
- Cell router — maps user ID to cell assignment; sticky for lifetime of account.
- Cell — self-contained stack: LB + app tier + database + cache.
- Control plane — shared services across cells: auth, billing, global config.
- Cell provisioning — automated pipeline to spin up a new cell when capacity fills.
- Cross-cell operations — admin tools, monitoring, and deploy orchestration across all cells.
Topology
User-to-cell assignment
Assign users to cells via consistent hashing on user ID. Once assigned, the mapping is permanent.
CELL_COUNT = 20
def get_cell_id(user_id: str) -> int:
hash_val = int(hashlib.sha256(user_id.encode()).hexdigest(), 16)
return hash_val % CELL_COUNT
# usr_42 → cell 7
# usr_99 → cell 14
{
"user_id": "usr_42",
"cell_id": 7,
"assigned_at": "2025-01-15T10:00:00Z"
}
Store the mapping in a global metadata service (small, highly available, rarely changes).
Cell internals
Each cell is a complete, independent deployment.
# Cell 7 deployment (identical across all cells)
apiVersion: apps/v1
kind: Deployment
metadata:
name: streamhub-api-cell-7
labels:
cell: "7"
spec:
replicas: 10
template:
metadata:
labels:
app: streamhub-api
cell: "7"
spec:
containers:
- name: api
image: registry.streamhub.io/api:3.2.0
env:
- name: CELL_ID
value: "7"
- name: DATABASE_URL
valueFrom:
secretKeyRef:
name: cell-7-db-credentials
key: url
Routing requests to the correct cell
def route_request(request):
user_id = extract_user_id(request)
cell_id = cell_router.get_cell(user_id)
cell_endpoint = f"https://cell-{cell_id}.streamhub.internal"
return proxy_to(cell_endpoint, request)
| Aspect | Approach | Blast radius |
|---|---|---|
| Monolith | Single DB, single deploy | 100% of users affected |
| Multi-region | Regional stacks | ~33% per region (3 regions) |
| Cell-based (20 cells) | Isolated mini-stacks | ~5% per cell |
| Cell-based (100 cells) | Fine-grained isolation | ~1% per cell |
Monolith
ApproachSingle DB, single deployBlast radius100% of users affectedMulti-region
ApproachRegional stacksBlast radius~33% per region (3 regions)Cell-based (20 cells)
ApproachIsolated mini-stacksBlast radius~5% per cellCell-based (100 cells)
ApproachFine-grained isolationBlast radius~1% per cell
More cells = smaller blast radius but higher operational overhead.
Shared control plane
Some services must be global and cannot be cell-local.
Control plane services
- Authentication — OAuth2 issuer validates tokens globally; cells trust the same JWKS.
- Billing / payments — PCI-compliant payment processing is centralised.
- Global search — federated query across cells (or per-cell index with aggregator).
- Admin dashboard — cross-cell visibility for operations and support.
- Deploy orchestration — rolling deploy across all cells with health gates.
Cell provisioning and capacity
When Cell 7 reaches 80% capacity (users, storage, or QPS), provision Cell 21.
# Automated cell provisioning pipeline
terraform apply -var="cell_id=21" -var="region=us-east-1"
kubectl apply -f cells/cell-21/
# Run smoke tests
# Update cell router to include cell 21 in hash ring
# New users assigned to cell 21; existing users stay in their cell
Quick recall
Everything you need if you only revisit this box.
- Cells are isolated mini-stacks (LB + apps + DB) that contain failure blast radius.
- Users assigned to a cell via consistent hashing; assignment is permanent.
- Control plane (auth, billing, admin) is shared; data plane is per-cell.
- More cells = smaller blast radius (~5% with 20 cells) but higher ops complexity.
- Provision new cells when existing ones reach 80% capacity.
Test yourself
Answer these before moving on — recall is what makes it stick.