PrepZone Logo
PrepZone

Read Replicas and Replication Lag

Serve reads from followers but know your staleness budget — ShardPay balance queries.

Read these first

Why this matters

  • Serving reads from followers without a lag budget causes user-visible bugs — "I paid but balance didn't change."
  • ShardPay routes balance checks to the primary after user actions; history and analytics hit replicas with documented staleness.
  • Monitoring replication lag is mandatory — it is the leading indicator before replica-induced incidents.
  • Interviewers ask how you expose staleness to clients and when you force leader reads.
Write
LeaderLatest data
Follower 1Lag 50ms
Follower 2Lag 200ms
Followers trail the leader — reads may return stale data until replication catches up.

Staleness budget

ShardPay's transaction history API tolerates 30 seconds of replication lag on read replicas. Live balance checks after a transfer hit the primary or a session-routed replica caught up within 200ms. Merchants see a "syncing" badge when replicationLagMs exceeds the SLO — transparency beats silent wrong data.

A staleness budget is a product decision encoded in infrastructure: not every replica is wrong for every query.

Key concepts

  • Replication lag — Time or byte offset between leader's latest commit and a follower's applied position. ShardPay measures lag as leader_commit_index - replica_applied_index and as wall-clock now() - last_replicated_at.
  • Read routing — Direct queries to primary, replica, or hybrid based on consistency needs. ShardPay's API gateway reads X-Read-Preference: primary|replica|session and routes accordingly.
  • Lag SLO — Maximum acceptable lag before alerting or degrading UX. ShardPay alerts at lag > 5s (warning) and > 30s (page on-call); history API shows syncing badge at > 10s.
  • GTID / log offset — Global transaction ID or commit index tracking replication position. ShardPay clients receive minCommitIndex in transfer responses for session-aware replica selection.
  • Catch-up reads — Replica promoted to serve session only after applying through minCommitIndex. Avoids leader load for every post-action poll while guaranteeing read-your-writes.

Routing rules by endpoint

EndpointTargetMax lag
POST /transfersLeader onlyN/A
GET /balance (default)Replica pool60s acceptable
GET /balance?fresh=trueLeader0
GET /balance + session tokenReplica ≥ minCommitIndex0 effective
GET /transactionsReplica30s
Fraud cumulative daily totalLeader or quorum0
  • POST /transfers

    TargetLeader only
    Max lagN/A
  • GET /balance (default)

    TargetReplica pool
    Max lag60s acceptable
  • GET /balance?fresh=true

    TargetLeader
    Max lag0
  • GET /balance + session token

    TargetReplica ≥ minCommitIndex
    Max lag0 effective
  • GET /transactions

    TargetReplica
    Max lag30s
  • Fraud cumulative daily total

    TargetLeader or quorum
    Max lag0

ShardPay read routing by endpoint

Walkthrough: Payroll day replica stress

  1. Merchant runs payroll — 2,000 transfers in 10 minutes on one shard.
  2. Leader write throughput spikes; replication to followers falls behind.
  3. replication_lag_ms climbs to 8 seconds — still within history API budget.
  4. Dashboard default balance polls start showing stale values for merchants without session tokens.
  5. ShardPay auto-enables syncing badge globally when fleet median lag > 5s.
  6. On-call scales read replica IOPS; lag recovers. Post-incident: lower default poll rate on balance widget.
Java
// ShardPay replica selector with lag guard
public Replica selectReplica(ReadPreference pref, Optional<SessionToken> session) {
    if (pref == ReadPreference.PRIMARY) return leader;

    if (session.isPresent()) {
        return replicaPool.firstCaughtUp(session.get().minCommitIndex())
            .orElse(leader);
    }

    Replica replica = replicaPool.leastLoaded();
    if (replica.lagMs() > MAX_STALE_MS) {
        metrics.increment("replica_lag_fallback_to_leader");
        return leader;
    }
    return replica;
}

Operating replicas under load

Operations

  • Replica lag dashboards — Per-shard, per-replica panels: bytes behind, seconds behind, apply rate. ShardPay's Grafana row links lag spikes to leader CPU and batch job schedules.
  • Hot replica problem — All read traffic hits one follower. ShardPay runs three followers per shard and load-balances reads with weighted least-connections.
  • Delayed replicas and backups — Backups from lagging replicas miss recent data. ShardPay snapshots from a follower that is < 1s behind or from leader during maintenance window.
  • Cascading replica tiers — Follower-of-follower for analytics reduces leader fan-out. ShardPay's warehouse ingest reads from a dedicated cascading replica that may lag minutes — never serves user-facing APIs.
  • Failover and lag — Before promoting replica to leader, confirm it has applied all committed entries. ShardPay's runbook blocks promotion if lag_entries > 0 for any entry below leader's commit_index.

Example: Support ticket from stale balance

A merchant transferred $2,000 to a vendor, received 201 Created, then polled balance from a replica 45 seconds behind. Dashboard showed full balance; merchant initiated a duplicate payment. Root cause: mobile app polled without session token and ignored replicationLagMs in the response header. Fix: SDK always sends session token after writes; UI shows "confirming…" until lagMs < 200 or 3s timeout triggers fresh read.

Quick recall

Everything you need if you only revisit this box.

  1. Replicas trade freshness for scale — document max lag per endpoint.
  2. Route strong reads to the leader; use session tokens for read-your-writes on replicas.
  3. Monitor and alert on replication lag; expose staleness to clients via headers and UI badges.

Test yourself

Answer these before moving on — recall is what makes it stick.