PrepZone Logo
PrepZone

Reliability, Availability, and Durability

Nines math, RPO/RTO, and how ShardPay sets SLA targets for payment ledger durability.

Why this matters

  • You cannot promise 99.99% availability without defining what "up" means.
  • Payment systems prioritize durability over raw availability.
  • Nines math appears in every architect interview.
  • Confusing availability with correctness causes architectures that are "up" but lose money.
ReliabilityMean time between failures
AvailabilityUptime / SLA %
DurabilityData survives crash

ShardPay ledger: high durability + CP availability trade-offs

Reliability is the probability of no failure over time; availability is uptime; durability is data surviving crashes.

Definitions that matter

These three words sound interchangeable but imply different trade-offs. Availability is the fraction of time the service accepts requests and returns non-error responses. Durability is whether committed data survives crashes, disk failures, and datacenter loss. Reliability is the probability the system performs correctly over a period — combining availability, durability, and correctness.

ShardPay's ledger must never lose a committed transfer (durability). The API should respond within 200ms for most requests (availability). Over a month, both should hold without reconciliation drift (reliability). You cannot maximize all three simultaneously during a partition — CAP forces a choice.

Core metrics

  • Availability — uptime fraction, usually expressed in nines (99.9%, 99.99%). ShardPay's transfer API targets 99.99% availability (~52 minutes downtime/year), measured as successful responses divided by total requests excluding client errors.
  • Durability — committed data survives failures. ShardPay replicates every ledger write to three replicas on independent disks before acknowledging the client; durability is 11 nines against single-disk loss.
  • RPO (Recovery Point Objective) — maximum acceptable data loss window in a disaster. ShardPay's RPO is zero for committed transfers: no acknowledged transfer may be lost, even if an AZ burns down.
  • RTO (Recovery Time Objective) — maximum time to restore service after disaster. ShardPay targets 15-minute RTO for ledger failover to a secondary region, accepting read-only mode for the first five minutes.
  • 99.99% ("four nines") — ~52 minutes downtime per year, ~4.3 minutes per month. ShardPay allocates this budget across deploys, failovers, and dependency outages using error budgets.

Durability before availability for payments

ShardPay will reject a write rather than acknowledge it before replication completes. During a network partition, the ledger chooses CP (consistency + partition tolerance): some clients see 503 Service Unavailable instead of a false success. A payment system that is "available" but loses transfers is worse than one that is briefly unavailable but correct.

This is the opposite of a social media feed, where showing a stale like count for 30 seconds is acceptable. ShardPay documents durability-first vs availability-first per component: ledger is CP, notification feed is AP, receipt PDF CDN is best-effort.

Example: AZ failure during peak load

US-East-1 loses power. ShardPay's ledger quorum loses one of three replicas but maintains majority in US-East-2. Writes continue with slightly higher latency. Availability dips to 99.97% for 20 minutes (RTO window). Durability holds: every transfer acknowledged before the failure is present on surviving replicas. No RPO breach. Customers see slower transfers, not lost money.

Java
// ShardPay ledger write path — durability before client ack
public TransferResult commitTransfer(Transfer transfer) {
    // 1. Append to local WAL (durable on disk)
    long index = wal.append(transfer);

    // 2. Replicate to quorum before responding
    ReplicationResult result = raft.replicate(index, transfer);
    if (!result.quorumAcked()) {
        throw new UnavailableException("cannot achieve write quorum");  // CP: reject, don't lie
    }

    // 3. Only now tell the client success
    return TransferResult.committed(transfer.id(), index);
}

Measuring and budgeting nines

Nines are multiplicative across dependencies. If auth is 99.99% and ledger is 99.99%, the serial path is roughly 99.98% — you lose nines by chaining. ShardPay parallelizes fraud and ledger where possible and caches auth tokens to protect the product-level SLO.

Error budgets make nines actionable. At 99.99%, ShardPay gets ~52 minutes of downtime per year. That budget covers planned deploys (~15 min/month), unplanned failovers, and dependency blips. When the budget burns too fast, feature deploys pause and reliability work takes priority.

SLO design at ShardPay

  • SLI (Service Level Indicator) — the measured signal, e.g., ratio of successful transfers under 200ms. ShardPay computes SLIs from distributed trace spans, not just load balancer logs.
  • SLO (Service Level Objective) — internal target, e.g., 99.99% monthly success rate. Breaching SLO consumes error budget; breaching SLA triggers customer credits.
  • SLA (Service Level Agreement) — contractual commitment to merchants, usually looser than internal SLO. ShardPay's merchant SLA is 99.95%; internal SLO is 99.99% to leave margin.
  • Error budget policy — when budget remains, ship features aggressively; when exhausted, freeze deploys and fix reliability. ShardPay's weekly SLO review ties directly to release train decisions.

Reliability as the composite goal

Reliability is what customers actually experience: correct transfers, on time, every month. A system can be highly available (fast 503s) and highly durable (no data loss) but unreliable if it frequently returns wrong balances. ShardPay's reliability SLO includes a correctness dimension: reconciliation drift must stay below $0.01 per million dollars transferred.

Designing for reliability means choosing the right priority per component, measuring all three dimensions, and using error budgets to balance velocity against risk. Module 11 covers the observability stack that makes these metrics visible in production.

Quick recall

Everything you need if you only revisit this box.

  1. Durability protects data; availability protects uptime; reliability combines both with correctness.
  2. Four nines ≈ 52 minutes downtime/year — budget it across deploys and incidents.
  3. ShardPay's ledger is CP: reject writes rather than lose committed transfers.

Test yourself

Answer these before moving on — recall is what makes it stick.