PrepZone Logo
PrepZone

Cell-Based and Regional Isolation

Shard-by-customer cohorts for independent failure and scale — ShardPay's cell architecture.

Why this matters

  • Meta, Apple, Slack, and AWS use cell-based architectures at scale for blast radius and capacity planning.
  • Cells turn "how big can one cluster get?" into "how many cells do we need?" — a linear scaling question.
  • ShardPay routes merchants to home cells; a cell outage affects ~25% of merchants, not the entire fleet.
  • Interviewers at staff level ask you to design cell routing, handle cell migration, and justify shared vs cell-local services.
Cell 1
LB
Services
Data
Cell 2
LB
Services
Data
Cell 3
LB
Services
Data
Each cell is a mini-deployment with its own data and traffic slice.

Anatomy of a ShardPay cell

Each cell is a complete vertical slice: ingress load balancers, API gateway, transfer orchestrator, ledger shards, Kafka cluster, and observability stack. A cell serves roughly 300,000 merchants with no runtime dependency on other cells for the transfer hot path.

Cell 1 holds merchants M-000001 through M-300000. Cell 2 holds M-300001 through M-600000. Cross-cell transfers exist (merchant in cell 1 pays merchant in cell 2) but are async sagas with higher latency budget — not the common case.

Key points

  • Cell — a self-contained deployment unit with full stack for a subset of tenants. ShardPay's cell 3 runs 847 pods, 12 ledger shards, and a dedicated 9-broker Kafka cluster.
  • Cell router — maps tenant to home cell at the edge. ShardPay's router uses consistent hash on merchantId with virtual nodes for even distribution.
  • N+1 cell capacity — spare cell ready to absorb traffic from a failed cell. ShardPay maintains cell 5 as warm standby at 10% synthetic traffic for continuous validation.
  • Shared services — identity, billing, merchant onboarding run globally with smaller blast radius and stricter SLOs. ShardPay accepts that identity outage blocks new logins but not in-flight transfers (tokens cached at gateway).

Routing and migration

Merchants have a home cell stored in the global registry. The edge router resolves merchantId → cellId on every request. Migration between cells happens for rebalancing or incident evacuation.

Cell router flow

  1. Request arrives at global anycast IP.
  2. Edge extracts merchantId from JWT or path.
  3. Registry lookup: M-442100 → cell 2.
  4. Request proxied to cell 2's regional ingress.
  5. All downstream calls stay within cell 2 unless cross-cell transfer saga explicitly calls cell 1.
Java
// ShardPay cell router — merchant to cell resolution
public RouteTarget resolve(RoutingContext ctx) {
    String merchantId = ctx.authenticatedMerchantId();
    CellAssignment assignment = registry.getAssignment(merchantId);
    if (assignment.status() == AssignmentStatus.MIGRATING) {
        // Dual-read during migration window
        return RouteTarget.both(assignment.sourceCell(), assignment.targetCell());
    }
    return RouteTarget.single(assignment.homeCell());
}

Walkthrough: migrating a merchant between cells

Merchant M-550000 moves from cell 2 to cell 3 for load rebalancing.

  1. Pre-migration: mark assignment MIGRATING; replication job copies ledger data to cell 3 shards.
  2. Cutover window: 30-second read-only on merchant account; final delta sync.
  3. Switch: registry updates home cell to 3; router sends new requests to cell 3.
  4. Verification: reconciliation compares cell 2 and cell 3 balances — must match.
  5. Cleanup: remove replica from cell 2 after 7-day retention for rollback safety.

Shared vs cell-local decisions

Not everything belongs in a cell. ShardPay uses a decision matrix: if the service is on the transfer hot path and scales with merchant count, it is cell-local. If it is global policy or infrequent, it is shared.

ServicePlacementRationale
Ledger shardsCell-localScales with merchant data
Transfer orchestratorCell-localHot path latency
Kafka settlementCell-localBlast radius isolation
Identity / OAuthSharedSingle source of truth for auth
Merchant onboardingSharedInfrequent, global workflow
Observability backendSharedAggregated view; buffered locally
  • Ledger shards

    PlacementCell-local
    RationaleScales with merchant data
  • Transfer orchestrator

    PlacementCell-local
    RationaleHot path latency
  • Kafka settlement

    PlacementCell-local
    RationaleBlast radius isolation
  • Identity / OAuth

    PlacementShared
    RationaleSingle source of truth for auth
  • Merchant onboarding

    PlacementShared
    RationaleInfrequent, global workflow
  • Observability backend

    PlacementShared
    RationaleAggregated view; buffered locally

ShardPay cell vs shared service placement

Key points

  • Horizontal scale by cell count — add cell 6 when average cell utilization exceeds 70%. ShardPay plans capacity in cells, not individual pod counts.
  • Cross-cell operations — expensive by design to discourage tight coupling. Cross-cell transfer adds ~40ms for inter-cell gRPC vs ~5ms intra-cell.
  • Cell-scoped deploys — canary within one cell before fleet-wide rollout. ShardPay's deploy pipeline requires 20-minute healthy canary in cell 1 before promoting to cells 2–4.
  • Data sovereignty — EU merchants pinned to EU cells for GDPR. ShardPay's cell 4 runs entirely in eu-west-1; routing policy prevents EU data from entering US cells.

Quick recall

Everything you need if you only revisit this box.

  1. Cells are full-stack isolated slices — scale by adding cells, not growing one cluster.
  2. Cell router maps merchantId to home cell; migration uses explicit MIGRATING state.
  3. Shared services (identity) are intentional blast radius — harden them separately.

Test yourself

Answer these before moving on — recall is what makes it stick.