Plan id: v3-dr-backup-plan.v1
Verifier: pnpm verify:v3 dr-backup (proposed)
Evidence artifacts: V3/operations/dr-backup-release-gate.json,
V3/operations/dr-restore-drill-<quarter>.json (proposed)
Owner: Capacity engineering (data on-call), joint with Lilith-Privacy lead for consent/audit classes and residency constraints
The POP-loss path is drilled (PIXEL_STREAMING_POP_CAPACITY_MANAGEMENT.md
multi-cloud failover), but compute failover is not data protection: nothing to
date sets recovery objectives for the data layer itself. This plan sets RPO/RTO
per data class, the backup cadence that achieves them, a quarterly restore drill
with pass criteria, and the region-loss runbook.
Residency boundary (binding): backups and replicas of residency-pinned data must
stay inside the residency zone signed off in V3/PER_REGION_RULES.md
(us-west, eu-de, uk, in, br). The pop-capacity residency drill already
establishes that India-to-eu-west-1 fallback requires explicit consent — the
same rule applies to data: we accept longer downtime over a residency
violation.
RPO/RTO Targets per Data Class#
| Data class | Store | RPO | RTO | Mechanism |
|---|---|---|---|---|
| World/persistence DBs (avatar history, attendance, holdings, tickets, commerce ledger) | Postgres | 5 min | 60 min | Continuous WAL archiving + nightly base backup (pgBackRest), PITR restore |
| Consent and audit ledgers (consent decisions, privacy-seat ledger, operator audit events, C2PA refs) | Postgres (dedicated DB) | 0 (committed txns) | 30 min | Synchronous in-zone standby (synchronous_commit=remote_apply), plus WAL archiving; daily immutable export to object-lock bucket |
| Session/queue state (gateway sessions, surge queue, rate counters) | Redis | 60 s | 15 min | AOF everysec + in-zone replica + 15-min RDB snapshots to object storage |
| Media/object storage (recordings, signed-edition masters, asset bundles, pak chunks) | MinIO / object storage | 15 min | 4 h (origin rebuild); 0 for already-cached CDN reads | Versioned bucket replication to a second in-zone site; CDN serves stale during origin loss |
| Live capture segments (mid-concert, pre-publication) | Capture pipeline + MinIO | one segment (≤ 10 s) | 3 min | Hot-standby capture worker (existing runbook) + segment-level replication on write |
Why these numbers: Redis session loss is survivable by design — the reconnect
gate (RECONNECT_VALIDATION.md) and login path rebuild session state, so 60 s
of session-state RPO costs reconnects, not data. The consent ledger is the
opposite extreme: a consent decision that exists in a user's memory but not in
our ledger is a legal exposure (RECORDING_CONSENT_FLOW.md publication gate
depends on it), so committed consent transactions get RPO 0 via synchronous
replication. Commerce/holdings sit at RPO 5 min because Stripe is an independent
source of truth for payments — a ≤5-min gap is reconcilable from provider
records via the existing refund/reversal machinery.
Backup Cadence and Retention#
| Action | Cadence | Retention |
|---|---|---|
Postgres base backup (pgbackrest backup --type=full, per stanza per zone) |
nightly 03:00 zone-local | 35 days PITR window |
Postgres WAL archiving (archive_command → pgBackRest repo) |
continuous | 35 days |
| Postgres monthly archival full | 1st of month | 12 months |
| Consent/audit immutable export (signed JSONL to object-lock bucket, compliance mode) | daily | 7 years |
| Redis RDB snapshot to object storage | every 15 min | 72 h |
| MinIO bucket replication (versioned, delete-marker replication off) | continuous (≤15 min lag alarmed) | versions kept 90 days |
| Off-provider copy (Postgres monthly full + consent exports → Azure in-zone region) | monthly | 12 months |
All backups are encrypted with per-zone KMS keys; backup-repo credentials are write-only from production (no production role can delete or overwrite the object-lock buckets — ransomware/insider containment). Backup success and replication lag are alarmed: a missed nightly base backup or WAL-archive lag
5 min (the RPO) pages the data on-call — an unalarmed backup failure silently converts RPO 5 min into RPO 24 h.
Quarterly Restore Drill#
Owner: data on-call engineer of the week; witness: one engineer not on the data
rota; evidence: V3/operations/dr-restore-drill-<quarter>.json, reviewed in the
weekly ops review.
Procedure (per zone, rotating so every zone is drilled at least annually;
eu-de and us-west every quarter):
- Pick the drill target time T = now − 24 h ± a random offset (the drill must not always restore to a tidy midnight boundary).
- Provision an isolated restore instance (no production network routes; separate security group) in the same zone.
- Postgres PITR:
pgbackrest --stanza=v3-world-<zone> --type=time --target="<T>" restorethen start and let recovery complete. Record wall-clock from command to accepting-connections. - Validation, all required:
pg_amcheckclean on restored indexes/heap;- row-count deltas vs production-at-T (from the hourly stats snapshot) within 0.1% per audited table;
- ledger coherence: consent-ledger entries for a sampled set of 100 recorded sessions are complete and ordered (the same invariant the publication gate checks);
- application smoke: a DSAR export (
DSAR_V1_PIPELINE.mdcategory set) builds successfully against the restored instance for a synthetic user.
- Redis: restore the latest RDB snapshot to an isolated instance; verify the surge-queue and session keyspaces deserialize and TTLs are sane.
- MinIO: restore 3 sampled objects per bucket class from the replica site, verify checksums against the C2PA/manifest hashes where applicable.
- Consent immutable export: fetch one daily export from the object-lock bucket, verify its signature chain.
- Write the drill evidence file: timings per step, validation results, restored-to point, deviations.
Pass criteria (all required, fail-closed):
- Postgres accepting-connections within 45 min and validation complete within 60 min (the RTO) — measured, not estimated;
- achieved restore point within 5 min of T (proves the RPO is real);
- consent-ledger DB restore within 30 min;
- zero validation failures in step 4;
- Redis and MinIO restores complete within their RTOs;
- evidence file written and reviewed within 7 days.
A failed drill is a GA-severity finding: the gate goes red, and the fix plus a re-drill must complete within 30 days.
Region-Loss Runbook#
- Owner: Capacity incident commander; severity SEV-1.
- Scenario: total loss of a primary cloud region hosting a residency zone's data plane (not just a POP — the POP-only case is already drilled).
What fails over:
- Pixel Streaming sessions: per the existing multi-cloud drills — secondary
cloud capacity in-zone (e.g., Azure
centralindiaforap-south-1), re-establishment ~30 s. - Edge/web/BFF: multi-region edge stays up; BFF traffic shifts to the surviving region serving non-residency-pinned routes.
- Global (non-pinned) Postgres data: promote the cross-region async replica; accept up to RPO 5 min of loss; reconcile commerce gaps against Stripe records before re-enabling payouts.
What is sacrificed (pre-decided, so nobody improvises under pressure):
- All in-flight sessions in the lost region: resume tokens die with the region;
users re-login (surge queue armed first — see
V3/launch/launch-surge-plan.md, the re-login stampede is a surge). - Up to 5 min of world/persistence writes (RPO), and up to 60 s of session state.
- Residency-pinned data does not fail over out of zone. If the
inzone is lost, India users see a regional outage page until in-zone restore completes (in-zone secondary site or cloud-secondary in-country region); we do not promote an out-of-country replica without the explicit-consent basis established in the residency drill. Same rule foreu-de,uk,br. - Free-tier capacity in surviving regions is shed first (85% backpressure ladder) to absorb displaced paid/booked users.
Order of operations: declare SEV-1 and assign commander → arm surge queue → fail over pxstream per capacity drills → promote non-pinned replicas → verify consent-ledger synchronous standby state before accepting new recordings anywhere → begin in-zone restore for pinned data from backups → status-page updates every 30 min → commerce reconciliation before payouts re-enable → post-incident review with measured RPO/RTO vs targets.
Fail-Closed Criteria#
The gate fails closed when any of the following holds:
- any data class above lacks a configured, alarmed backup mechanism matching its cadence row;
- WAL-archive or replication lag alarms are missing or routed to no pager;
- the most recent quarterly restore drill is missing, failed, or older than 100 days;
- backup repositories are writable-or-deletable by production roles, or object-lock is absent on consent/audit exports;
- any backup or replica of residency-pinned data exists outside its residency zone;
- the region-loss runbook has no named commander role or has not had a tabletop
walk-through before GA (evidence:
drill-region-loss-tabletop-ga).