Why this matters
- Architects are judged on blast radius containment, not just uptime percentages.
- Correlated failures — same software version, shared control plane, single power feed — defeat naive N+1 redundancy.
- ShardPay's 2024 cell-1 incident affected 12,000 merchants, not 1.2M, because cell boundaries contained the blast.
- Staff+ interviews ask you to identify hidden shared failure domains in a diagram, not just label AZs.
Mapping failure domains
A failure domain is any scope where one trigger (power loss, network cut, bad config, CVE exploit) can knock out multiple components at once. Blast radius is how much of the business stops working when that domain fails.
ShardPay maps domains hierarchically: rack → AZ → region → cell → global control plane. Each layer should contain failures from the layer below without cascading upward unchecked.
Key points
- Failure domain — the scope of a single root-cause event. ShardPay treats one AZ as a failure domain: losing
us-east-1ashould not take downus-east-1borus-east-1c. - Blast radius — the user-visible impact when a domain fails. Cell 1 outage affects merchants 1–12,000; shared identity service outage affects all 1.2M merchants — a larger blast radius by design (with extra redundancy).
- Physical isolation — AZs have independent power and networking within a region. ShardPay runs each shard's replicas across 3 AZs with quorum writes requiring 2/3 AZs.
- Correlated failure — same bug in every instance because every instance runs the same deploy. ShardPay uses canary deploys per cell so a bad binary reaches cell 1 before cell 4.
ShardPay isolation strategy
ShardPay combines AZ redundancy within cells and cell independence across the fleet. The remaining intentional shared domains are identity, DNS, and the deployment pipeline — each with extra hardening.
Per-shard AZ placement
Each ledger shard has a primary in one AZ and replicas in two others. Write quorum (2 of 3) survives single-AZ loss. ShardPay tested this with AZ failure drills quarterly — last drill recovered write path in 47 seconds.
Cell independence
Four cells each serve ~300k merchants with full stack: load balancers, API, orchestrator, shards, Kafka cluster. Cell 2's Kafka cluster is separate from cell 1's — no cross-cell broker dependency during normal operation.
Shared control plane risks
Identity (OAuth tokens), external DNS, and the CI/CD pipeline are shared across cells. These are ShardPay's largest blast radius by design. Mitigations: multi-AZ identity replicas, DNS with anycast failover, deploy gates requiring per-cell canary success.
// ShardPay cell-aware routing — blast radius containment
public Cell routeMerchant(String merchantId) {
int cellId = consistentHash(merchantId, cellRegistry.activeCells());
Cell cell = cellRegistry.get(cellId);
if (cell.health() == CellHealth.DEGRADED) {
// Fail over only within pre-defined pairs — not global drain
return cellRegistry.failoverPartner(cellId);
}
return cell;
}
Reducing blast radius over time
Blast radius reduction is iterative. ShardPay's evolution: monolith → sharded DB → cells → per-cell Kafka → canary-by-cell deploys.
Walkthrough: bad deploy contained to one cell
- New orchestrator version deploys to cell 1 canary (5% of cell 1 traffic).
- Within 10 minutes, cell 1 transfer SLI drops from 99.97% to 99.2% — burn rate alert fires.
- Automated rollback reverts cell 1 only; cells 2–4 unaffected on previous version.
- Global SLI impact: 0.003% — within error budget because blast was cell-scoped.
- Postmortem: add integration test for the regression; extend canary observation window from 10 to 20 minutes.
Key points
- N+1 at every layer — spare capacity, not just spare instances. ShardPay maintains a warm standby cell (cell 5) that can absorb one failed cell's merchants within 30 minutes of DNS reroute.
- Bulkhead by tenant tier — enterprise merchants get dedicated cell slices with smaller neighbor blast radius. ShardPay's top 50 merchants each have isolated shard pairs.
- Dependency audit — quarterly review of shared dependencies (logging SaaS, APM vendor, cloud provider API). A Datadog outage should not prevent transfers — ShardPay runs local metric buffers.
- Failure domain testing — game days that actually kill AZs and cells, not tabletop exercises. ShardPay's Q3 game day killed cell 3 entirely; failover to cell 5 completed in 22 minutes.
Quick recall
Everything you need if you only revisit this box.
- Failure domains bound incident scope; blast radius is the user-visible impact.
- Cells contain deploy and infrastructure failures; shared services remain the largest risk.
- Canary deploys per cell limit correlated software failure blast radius.
Test yourself
Answer these before moving on — recall is what makes it stick.