PrepZone Logo
PrepZone

Chaos Engineering and Jepsen Testing

Prove your consistency claims under partition — not just in design docs.

Why this matters

  • ShardPay's untested failover path for shard metadata lost routing table updates during the first real AZ failure.
  • Jepsen history analysis exposed "linearizable" claims that lost committed transfers under network partition in other systems — ShardPay runs similar tests on ledger replicas before major upgrades.
  • Game days reveal retry policies, circuit breakers, and runbooks that unit tests never exercise.
  • Interviewers ask how you validate consistency and resilience claims without waiting for production incidents.
Steady-state hypothesis
Inject faultKill AZ, latency, partition
Observe metrics
Improve resilience
Hypothesize steady state, inject failure, observe SLI drift, fix before customers notice.

Chaos experiments with hypothesis

A chaos experiment starts with a hypothesis: "If we kill one ledger pod per shard during peak load, p99 transfer latency stays under 300ms and no committed transfer is lost." ShardPay defines blast radius (staging only, one region), abort conditions (error rate > 5%, manual kill switch), and observability dashboards before injecting failure.

Tools: Litmus or Chaos Mesh on Kubernetes for pod kills; Toxiproxy for latency and partition simulation; custom scripts for clock skew injection on staging nodes.

Walkthrough: monthly game day

ShardPay's SRE team runs a four-hour game day in staging mirroring production traffic shape. At T+60m they terminate 30% of ledger pods randomly. Hypothesis holds except fraud circuit opens from retry noise — runbook updated to widen bulkhead before next drill. At T+120m they partition a minority of Raft voters from the sequencer cluster; no duplicate settlement batch IDs; one 400ms leader election blip logged.

Key points

  • Chaos experiment — structured test with hypothesis, metrics, blast radius, and abort criteria. ShardPay files a one-page experiment doc in Confluence before every drill; post-mortem compares expected vs actual behavior.
  • Game day — scheduled multi-team drill simulating realistic failures under load. ShardPay includes payments ops, support, and comms to practice customer messaging when transfers enter pending_fraud at scale.
  • Jepsen — framework that injects network partitions and clock skew while a workload generator runs; analyzes history for linearizability and other consistency violations. ShardPay adapted Jepsen-style workloads against staging ledger replicas before migrating from sync to semi-sync replication.
  • Blast radius — limit scope: one namespace, one region, max percentage of pods. Production chaos at ShardPay stays at ≤1% canary traffic with instant rollback — never "random prod kills" without guardrails.
  • Consistency checking — compare linearization graphs from concurrent read/write histories. ShardPay's checker verifies no lost committed transfer after partition heals — if Jepsen finds an anomaly, release blocks.

Jepsen-style verification for the ledger

ShardPay's ledger shard uses majority commit before acknowledging transfer. Jepsen-style tests run a nemesis that isolates minority replicas while a client performs concurrent transfers and balance reads. After heal, checker validates: every acknowledged transfer appears in all full replicas; balances match sum of entries.

This caught a bug where async replication lag allowed stale reads on minority-follower routing — fixed by routing reads only to committed-index-matched nodes.

Walkthrough: partition during commit

Nemesis partitions leader from one follower during commit of transfer T-991. Majority still acks; client receives success. Partition heals; checker confirms T-991 on all survivors. Minority follower that missed entry truncates log and resyncs — no duplicate, no lost commit. Without this test ShardPay once served stale balance from unsynced follower in a misconfigured read pool.

Java
// ShardPay chaos hook — guarded pod failure injection (staging only)
@Component
@Profile("staging-chaos")
public class LedgerChaosHook {
    private final AtomicBoolean armed = new AtomicBoolean(false);

    public void armExperiment(String hypothesis, double maxErrorRate) {
        // SRE arms via admin API after hypothesis doc approved
        armed.set(true);
    }

    @Scheduled(fixedDelay = 30_000)
    void maybeKillRandomLedgerPod() {
        if (!armed.get()) return;
        if (metrics.errorRate() > 0.05) {
            abort("error rate exceeded abort threshold");
            return;
        }
        kubernetesClient.apps().statefulSets()
            .inNamespace("shardpay-staging")
            .withName("ledger")
            .scale(replicaCount - 1, true); // documented in runbook; auto-restores
    }
}

From staging chaos to production confidence

Start in staging with abort switches. Graduate to small production experiments (fault injection on single canary pod) only after staging passes three consecutive drills. ShardPay never runs partition tests in prod on the ledger — too much regulatory risk — but does inject latency on non-critical read paths in canaries.

Document every surprise as runbook input: who pages, what customer comms, which circuit breakers should have opened but did not.

Quick recall

Everything you need if you only revisit this box.

  1. Chaos experiments need hypothesis, blast radius, and abort conditions — not random failures.
  2. Jepsen validates consistency under partition; load tests validate latency — both required.
  3. Start in staging; graduate to tiny prod canaries only with kill switches and runbooks.

Test yourself

Answer these before moving on — recall is what makes it stick.