Why this matters
- Meta, Apple, Slack, and AWS use cell-based architectures at scale for blast radius and capacity planning.
- Cells turn "how big can one cluster get?" into "how many cells do we need?" — a linear scaling question.
- ShardPay routes merchants to home cells; a cell outage affects ~25% of merchants, not the entire fleet.
- Interviewers at staff level ask you to design cell routing, handle cell migration, and justify shared vs cell-local services.
Anatomy of a ShardPay cell
Each cell is a complete vertical slice: ingress load balancers, API gateway, transfer orchestrator, ledger shards, Kafka cluster, and observability stack. A cell serves roughly 300,000 merchants with no runtime dependency on other cells for the transfer hot path.
Cell 1 holds merchants M-000001 through M-300000. Cell 2 holds M-300001 through M-600000. Cross-cell transfers exist (merchant in cell 1 pays merchant in cell 2) but are async sagas with higher latency budget — not the common case.
Key points
- Cell — a self-contained deployment unit with full stack for a subset of tenants. ShardPay's cell 3 runs 847 pods, 12 ledger shards, and a dedicated 9-broker Kafka cluster.
- Cell router — maps tenant to home cell at the edge. ShardPay's router uses consistent hash on
merchantIdwith virtual nodes for even distribution. - N+1 cell capacity — spare cell ready to absorb traffic from a failed cell. ShardPay maintains cell 5 as warm standby at 10% synthetic traffic for continuous validation.
- Shared services — identity, billing, merchant onboarding run globally with smaller blast radius and stricter SLOs. ShardPay accepts that identity outage blocks new logins but not in-flight transfers (tokens cached at gateway).
Routing and migration
Merchants have a home cell stored in the global registry. The edge router resolves merchantId → cellId on every request. Migration between cells happens for rebalancing or incident evacuation.
Cell router flow
- Request arrives at global anycast IP.
- Edge extracts
merchantIdfrom JWT or path. - Registry lookup:
M-442100 → cell 2. - Request proxied to cell 2's regional ingress.
- All downstream calls stay within cell 2 unless cross-cell transfer saga explicitly calls cell 1.
// ShardPay cell router — merchant to cell resolution
public RouteTarget resolve(RoutingContext ctx) {
String merchantId = ctx.authenticatedMerchantId();
CellAssignment assignment = registry.getAssignment(merchantId);
if (assignment.status() == AssignmentStatus.MIGRATING) {
// Dual-read during migration window
return RouteTarget.both(assignment.sourceCell(), assignment.targetCell());
}
return RouteTarget.single(assignment.homeCell());
}
Walkthrough: migrating a merchant between cells
Merchant M-550000 moves from cell 2 to cell 3 for load rebalancing.
- Pre-migration: mark assignment
MIGRATING; replication job copies ledger data to cell 3 shards. - Cutover window: 30-second read-only on merchant account; final delta sync.
- Switch: registry updates home cell to 3; router sends new requests to cell 3.
- Verification: reconciliation compares cell 2 and cell 3 balances — must match.
- Cleanup: remove replica from cell 2 after 7-day retention for rollback safety.
Shared vs cell-local decisions
Not everything belongs in a cell. ShardPay uses a decision matrix: if the service is on the transfer hot path and scales with merchant count, it is cell-local. If it is global policy or infrequent, it is shared.
| Service | Placement | Rationale |
|---|---|---|
| Ledger shards | Cell-local | Scales with merchant data |
| Transfer orchestrator | Cell-local | Hot path latency |
| Kafka settlement | Cell-local | Blast radius isolation |
| Identity / OAuth | Shared | Single source of truth for auth |
| Merchant onboarding | Shared | Infrequent, global workflow |
| Observability backend | Shared | Aggregated view; buffered locally |
Ledger shards
PlacementCell-localRationaleScales with merchant dataTransfer orchestrator
PlacementCell-localRationaleHot path latencyKafka settlement
PlacementCell-localRationaleBlast radius isolationIdentity / OAuth
PlacementSharedRationaleSingle source of truth for authMerchant onboarding
PlacementSharedRationaleInfrequent, global workflowObservability backend
PlacementSharedRationaleAggregated view; buffered locally
ShardPay cell vs shared service placement
Key points
- Horizontal scale by cell count — add cell 6 when average cell utilization exceeds 70%. ShardPay plans capacity in cells, not individual pod counts.
- Cross-cell operations — expensive by design to discourage tight coupling. Cross-cell transfer adds ~40ms for inter-cell gRPC vs ~5ms intra-cell.
- Cell-scoped deploys — canary within one cell before fleet-wide rollout. ShardPay's deploy pipeline requires 20-minute healthy canary in cell 1 before promoting to cells 2–4.
- Data sovereignty — EU merchants pinned to EU cells for GDPR. ShardPay's cell 4 runs entirely in
eu-west-1; routing policy prevents EU data from entering US cells.
System Design trackSee cell architecture in system design
Quick recall
Everything you need if you only revisit this box.
- Cells are full-stack isolated slices — scale by adding cells, not growing one cluster.
- Cell router maps merchantId to home cell; migration uses explicit MIGRATING state.
- Shared services (identity) are intentional blast radius — harden them separately.
Test yourself
Answer these before moving on — recall is what makes it stick.