PrepZone Logo
PrepZone

Backup and Disaster Recovery

RPO, RTO, point-in-time recovery, and cross-region failover for VaultCommerce data.

Why this matters

  • A regional AWS outage without a tested failover plan means VaultCommerce cannot process checkouts — revenue stops and carts abandon.
  • Interviewers ask you to define RPO/RTO, then walk through how Postgres PITR and Redis persistence map to each.
  • Backups you have never restored are Schrödinger's backups — VaultCommerce runs quarterly restore drills into an isolated VPC.
App writesINSERT / UPDATE
LeaderPrimary node
Replica 1Read traffic
Replica 2Read traffic
Writes go to the leader; replicas serve read traffic asynchronously.

RPO and RTO for VaultCommerce

Targets by data tier

  • Orders and payments (Postgres) — RPO 5 minutes, RTO 30 minutes. Achieved with continuous WAL archiving plus cross-region read replica promoted on failover.
  • Sessions and cart cache (Redis) — RPO 1 minute, RTO 10 minutes. AOF with appendfsync everysec; accept sub-second loss on catastrophic node failure.
  • Product search (Elasticsearch 8) — RPO 15 minutes, RTO 45 minutes. Snapshot to S3 every 15 minutes; search is rebuildable from Postgres if snapshots fail.
  • Analytics events (DynamoDB) — RPO 0 with PITR enabled; RTO 60 minutes for table restore into a new region.

RPO answers: "If the database vanishes right now, how many minutes of writes are gone?" RTO answers: "How long until checkout works again?"

Postgres 16 backup stack

VaultCommerce runs RDS PostgreSQL 16 with automated backups and WAL archiving to S3.

Java
-- Verify PITR window (RDS: retention set in console; self-managed: archive_command)
SELECT pg_switch_wal();
SELECT * FROM pg_stat_archiver;

Automated snapshots — daily full backup; retention 35 days.
Continuous archiving — WAL segments shipped to S3 every 60 seconds; enables point-in-time recovery to any second within the window.
Logical backups — weekly pg_dump of schema-only for migration testing; not used for production restore (too slow at 400 GB).

Java
# Restore to a specific timestamp (self-managed example)
pg_restore --dbname=vaultcommerce_recovery \
  --target-time="2026-09-15 14:32:00+00" backup_manifest

Cross-region failover

Primary: us-east-1. Standby: us-west-2 async replica.

StepActionOwner
1Detect primary unreachable (RDS event + synthetic checkout probe)PagerDuty → on-call
2Confirm replica lag < RPO budget (CloudWatch ReplicaLag)DBA
3Promote replica or trigger RDS cross-region failoverDBA + SRE
4Update Route 53 / app config to new writer endpointSRE
5Invalidate Redis sessions; warm critical cache keysApp team
6Post-incident: verify row counts, replay failed payments queueAll
  • 1

    ActionDetect primary unreachable (RDS event + synthetic checkout probe)
    OwnerPagerDuty → on-call
  • 2

    ActionConfirm replica lag < RPO budget (CloudWatch ReplicaLag)
    OwnerDBA
  • 3

    ActionPromote replica or trigger RDS cross-region failover
    OwnerDBA + SRE
  • 4

    ActionUpdate Route 53 / app config to new writer endpoint
    OwnerSRE
  • 5

    ActionInvalidate Redis sessions; warm critical cache keys
    OwnerApp team
  • 6

    ActionPost-incident: verify row counts, replay failed payments queue
    OwnerAll

Run this playbook in a game day every quarter. VaultCommerce's last drill caught a hardcoded writer hostname.

Redis 7 persistence for DR

Sessions and inventory locks live in Redis 7. VaultCommerce uses AOF (appendonly yes, appendfsync everysec) plus RDB snapshots every 6 hours to S3 via redis-cli --rdb.

Java
redis-cli INFO persistence
# aof_enabled:1
# aof_last_rewrite_time_sec:...
# rdb_last_save_time:...

For Redis Cluster, restore means spinning new nodes and loading RDB/AOF — plan for 10–15 minutes. Cart data is also in Postgres; Redis loss degrades UX but does not lose paid orders.

Elasticsearch snapshot recovery

Java
PUT _snapshot/vaultcommerce-s3/daily-2026-09-15
{
  "indices": "products-*",
  "ignore_unavailable": true,
  "include_global_state": false
}

If ES is down but Postgres is healthy, reindex from source tables — slower (2–4 hours) but always available as a last resort.

Quick recall

Everything you need if you only revisit this box.

  • RPO = acceptable data loss window; RTO = acceptable downtime window — define both per data tier.
  • Postgres 16: automated snapshots + continuous WAL archiving enable point-in-time recovery.
  • Cross-region failover requires lag checks, DNS/config updates, and cache warming — rehearse quarterly.
  • Redis 7: AOF everysec + periodic RDB; sessions are expendable relative to Postgres orders.
  • Elasticsearch 8: S3 snapshots for fast restore; Postgres reindex as fallback.
  • Untested backups are not backups — run restore drills into an isolated environment.

Test yourself

Answer these before moving on — recall is what makes it stick.