Why this matters
- Reliability engineering vocabulary separates what broke from what users saw.
- Retries and redundancy convert faults into invisible recoveries.
- ShardPay's SLO dashboard tracks failures, while logs track faults.
- The worst incidents return HTTP 200 with wrong data — errors, not failures.
Three layers of trouble
Reliability engineering uses precise vocabulary because each layer needs a different response. A fault is a latent defect: a bug in replication code, a flaky NIC, a misconfigured timeout. A failure is when the system deviates from its spec: a 503, a timeout, an unavailable shard. An error is incorrect output despite apparently successful operation: a stale balance, a duplicate charge, a transfer credited twice.
The diagnostic chain runs fault → failure → error, but detection runs in reverse. Customers report errors. Monitoring alerts on failures. Post-mortems dig into faults.
Walkthrough: disk fault to customer error
A ledger replica develops bad sectors on its write-ahead log disk (fault). The replica stops acknowledging writes and falls out of the Raft quorum (failure — reduced write capacity). The leader fails over cleanly. A mobile client polls a follower that has not yet replicated the latest debit and displays an old balance (error — HTTP 200, wrong number). ShardPay fixed this by routing balance reads after transfer to the leader for 30 seconds, not by adding more retries.
Definitions
- Fault — an underlying defect that may or may not surface; includes software bugs, hardware degradation, and config drift. ShardPay tracks fault indicators (disk latency p99, GC pause duration) on dashboards that do not page on-call — they inform capacity planning.
- Failure — observable deviation from the service spec: elevated 5xx rate, SLO burn, or shard unavailability. ShardPay pages when transfer success rate drops below 99.95% over five minutes.
- Error — semantically wrong output: incorrect balance, lost transfer that returned success, or duplicate idempotent replay applied twice due to a bug. Errors are the hardest class because monitoring often shows green while customers lose money.
- Fault tolerance — continuing to deliver correct service despite faults, usually through redundancy, replication, and automatic failover. ShardPay's three-replica ledger tolerates one replica fault without user-visible failure.
Detecting faults before they become failures
Proactive fault detection shrinks the blast radius. ShardPay runs synthetic transfer probes every 30 seconds, monitors replica lag against a 2-second threshold, and alerts on rising disk latency before the disk fails completely. A fault caught at "replica lag 1.8s" allows a controlled failover; a fault caught at "replica not responding" becomes a customer-facing latency spike.
Health checks must distinguish alive from correct. A pod that responds to /health but serves stale data is alive but faulty. ShardPay's readiness probe includes a check against the leader's latest committed index, not just a TCP socket open.
// ShardPay ledger — readiness distinguishes alive vs correct
@Component
public class LedgerReadinessProbe implements HealthIndicator {
private final RaftNode raft;
private final ReplicationLagMonitor lag;
@Override
public Health health() {
if (!raft.isLeader() && lag.millisecondsBehindLeader() > 2_000) {
return Health.down()
.withDetail("reason", "replica lag exceeds 2s")
.withDetail("lagMs", lag.millisecondsBehindLeader())
.build(); // remove from load balancer before serving stale data
}
return Health.up().build();
}
}
From failure to recovery
When a failure is customer-visible, the goal is bounded recovery time and no data loss. ShardPay's runbooks distinguish failover (promote a replica, ~seconds) from restore (replay backups, ~hours). Failures during failover are expected; errors after failover are not.
Retries convert transient failures into successful recoveries — but only with idempotency. A client that retries a timed-out transfer must not create a second debit. ShardPay's idempotency keys live in the failure-handling layer, not as an afterthought.
Response strategies
- Redundancy — N+1 replicas so one fault does not cause failure. ShardPay runs three ledger replicas per shard; one fault leaves quorum intact.
- Graceful degradation — return partial results when a dependency fails. ShardPay returns transfer status
PENDINGwhen fraud is unavailable rather than failing the entire transfer. - Failover automation — Raft leader election promotes a follower without human intervention. ShardPay targets sub-10-second leader failover for ledger shards.
- Error budgets — allocate acceptable failure rate before halting deploys. ShardPay burns error budget on 5xx spikes; stale-read errors count against a separate data-quality SLO.
Why errors are the hardest class
Failures are loud: dashboards turn red, pagers fire, customers see error pages. Errors are quiet: the API returns 200, the transfer ID looks valid, but the balance is wrong. ShardPay's costliest incident was a rounding bug in cross-currency transfers — every HTTP metric was green for three days until reconciliation found a $12,000 discrepancy.
Defending against errors requires invariants: properties that must always hold (total debits equal total credits per shard, idempotency keys are unique, sequence numbers are monotonic). ShardPay runs hourly reconciliation jobs that compare ledger totals against an independent audit trail, catching errors that slip past request-level monitoring.
Quick recall
Everything you need if you only revisit this box.
- Fault → failure → error is the diagnostic chain; detect errors with invariants, not just status codes.
- Fault tolerance and redundancy convert faults into invisible recoveries.
- Errors with HTTP 200 are the most dangerous class in payment systems.
Test yourself
Answer these before moving on — recall is what makes it stick.