Why this matters
- A regional AWS outage without a tested failover plan means VaultCommerce cannot process checkouts — revenue stops and carts abandon.
- Interviewers ask you to define RPO/RTO, then walk through how Postgres PITR and Redis persistence map to each.
- Backups you have never restored are Schrödinger's backups — VaultCommerce runs quarterly restore drills into an isolated VPC.
RPO and RTO for VaultCommerce
Targets by data tier
- Orders and payments (Postgres) — RPO 5 minutes, RTO 30 minutes. Achieved with continuous WAL archiving plus cross-region read replica promoted on failover.
- Sessions and cart cache (Redis) — RPO 1 minute, RTO 10 minutes. AOF with
appendfsync everysec; accept sub-second loss on catastrophic node failure. - Product search (Elasticsearch 8) — RPO 15 minutes, RTO 45 minutes. Snapshot to S3 every 15 minutes; search is rebuildable from Postgres if snapshots fail.
- Analytics events (DynamoDB) — RPO 0 with PITR enabled; RTO 60 minutes for table restore into a new region.
RPO answers: "If the database vanishes right now, how many minutes of writes are gone?" RTO answers: "How long until checkout works again?"
Postgres 16 backup stack
VaultCommerce runs RDS PostgreSQL 16 with automated backups and WAL archiving to S3.
-- Verify PITR window (RDS: retention set in console; self-managed: archive_command)
SELECT pg_switch_wal();
SELECT * FROM pg_stat_archiver;
Automated snapshots — daily full backup; retention 35 days.
Continuous archiving — WAL segments shipped to S3 every 60 seconds; enables point-in-time recovery to any second within the window.
Logical backups — weekly pg_dump of schema-only for migration testing; not used for production restore (too slow at 400 GB).
# Restore to a specific timestamp (self-managed example)
pg_restore --dbname=vaultcommerce_recovery \
--target-time="2026-09-15 14:32:00+00" backup_manifest
Cross-region failover
Primary: us-east-1. Standby: us-west-2 async replica.
| Step | Action | Owner |
|---|---|---|
| 1 | Detect primary unreachable (RDS event + synthetic checkout probe) | PagerDuty → on-call |
| 2 | Confirm replica lag < RPO budget (CloudWatch ReplicaLag) | DBA |
| 3 | Promote replica or trigger RDS cross-region failover | DBA + SRE |
| 4 | Update Route 53 / app config to new writer endpoint | SRE |
| 5 | Invalidate Redis sessions; warm critical cache keys | App team |
| 6 | Post-incident: verify row counts, replay failed payments queue | All |
1
ActionDetect primary unreachable (RDS event + synthetic checkout probe)OwnerPagerDuty → on-call2
ActionConfirm replica lag < RPO budget (CloudWatch ReplicaLag)OwnerDBA3
ActionPromote replica or trigger RDS cross-region failoverOwnerDBA + SRE4
ActionUpdate Route 53 / app config to new writer endpointOwnerSRE5
ActionInvalidate Redis sessions; warm critical cache keysOwnerApp team6
ActionPost-incident: verify row counts, replay failed payments queueOwnerAll
Run this playbook in a game day every quarter. VaultCommerce's last drill caught a hardcoded writer hostname.
Redis 7 persistence for DR
Sessions and inventory locks live in Redis 7. VaultCommerce uses AOF (appendonly yes, appendfsync everysec) plus RDB snapshots every 6 hours to S3 via redis-cli --rdb.
redis-cli INFO persistence
# aof_enabled:1
# aof_last_rewrite_time_sec:...
# rdb_last_save_time:...
For Redis Cluster, restore means spinning new nodes and loading RDB/AOF — plan for 10–15 minutes. Cart data is also in Postgres; Redis loss degrades UX but does not lose paid orders.
Elasticsearch snapshot recovery
PUT _snapshot/vaultcommerce-s3/daily-2026-09-15
{
"indices": "products-*",
"ignore_unavailable": true,
"include_global_state": false
}
If ES is down but Postgres is healthy, reindex from source tables — slower (2–4 hours) but always available as a last resort.
Quick recall
Everything you need if you only revisit this box.
- RPO = acceptable data loss window; RTO = acceptable downtime window — define both per data tier.
- Postgres 16: automated snapshots + continuous WAL archiving enable point-in-time recovery.
- Cross-region failover requires lag checks, DNS/config updates, and cache warming — rehearse quarterly.
- Redis 7: AOF everysec + periodic RDB; sessions are expendable relative to Postgres orders.
- Elasticsearch 8: S3 snapshots for fast restore; Postgres reindex as fallback.
- Untested backups are not backups — run restore drills into an isolated environment.
Test yourself
Answer these before moving on — recall is what makes it stick.