Disaster Recovery Plan
This plan applies to stateful production services hosted on the current Contabo VPS:
Status
- Effective date: 2026-04-15
- Priority: governance-first baseline implementation
- Approval model: single-owner explicit approval or manual execution by owner
- Backup destinations: encrypted Google Drive for WordPress/databases; same-host RustFS remains interim coverage for other state
- Critical service recovery target: RTO 4 hours
- Critical service data-loss target: RPO 1 hour
Scope
This plan applies to stateful production services hosted on the current Contabo VPS:
- Databases: MSSQL, PostgreSQL, MySQL, MongoDB
- Registry data and metadata
- Object storage (RustFS)
- RabbitMQ definitions/messages where required by service recovery policy
Scheduled encrypted Google Drive delivery is active for the four scoped WordPress sites, weekly PostgreSQL/MySQL/MongoDB logical sets, and monthly PostgreSQL physical snapshots. RustFS remains same-host interim coverage for state outside this rollout, so those systems still require separately approved offsite coverage.
Repository support for OAuth-backed, client-side encrypted Google Drive delivery is implemented for WordPress, weekly logical databases, and monthly PostgreSQL physical snapshots. Production OAuth/crypt configuration, canaries, complete workflows, targeted logical restores, and the first physical snapshot/download/isolated restore are validated. All three backup timers are active.
Data Classes
| Class | Description | Example Systems | Target |
|---|---|---|---|
| Tier 1 | Critical state | PostgreSQL, MySQL, MSSQL, MongoDB, RustFS object data | RTO <= 4h, RPO <= 1h |
| Tier 2 | Important operational state | Registry metadata, RabbitMQ runtime data | Restore in same maintenance window |
| Tier 3 | Rebuildable state | Stateless containers and images | Rebuild from CI/registry |
Backup Policy Baseline
- Minimum cadence for Tier 1 data must satisfy RPO 1 hour.
- The current weekly database backup timer is a retention baseline and does not satisfy the one-hour Tier 1 RPO. This remains an explicit DR gap until a separate higher-frequency recovery mechanism is approved and validated.
- Every backup job must write integrity metadata (timestamp, source, checksum, size).
- Keep retention policy explicit in runbook execution logs.
- Backups must be encrypted at rest where the destination supports it.
- Google Drive delivery must use
rclone crypt, verify withrclone cryptcheck, and preserve local artifacts whenever upload or verification fails. - Weekly PostgreSQL coverage must include globals and every connectable non-template database as individual custom-format dumps.
- Monthly PostgreSQL container recovery must use online
pg_basebackupwith included WAL and a native SHA-256 manifest; do not archive a live PostgreSQL Docker volume. - Keep the newest 12 weekly remote/3 local sets, 4 monthly PostgreSQL physical remote/2 local snapshots, and 4 WordPress archives per site. Prune only after verification.
- Shared database sets must remain central-only; client Drives may receive only data explicitly mapped to that client.
- Any backup job or restore operation on production services requires owner approval.
Restore Validation Policy
- Run at least one restore validation drill per month for one Tier 1 system.
- Record measured recovery time and data recovery point in a dated incident or ops note.
- If measured results exceed targets, open corrective work in next-steps immediately.
Interim Risk Statement
- Same-host backups reduce accidental deletion risk but do not cover full host-loss scenarios.
- Scheduled Google Drive coverage now protects the scoped WordPress and database data from host loss. The weekly logical cadence still misses the one-hour Tier-1 RPO, and RustFS/registry/RabbitMQ or other service state remains outside this activation.
Execution Checklist
- Finalize RTO/RPO targets for critical services (4h/1h).
- Finalize interim backup destination (RustFS on same VPS).
- Implement automated backup jobs for Tier 1 systems.
- Validate first end-to-end restore test against RTO/RPO targets.
- Define and implement the encrypted Google Drive destination model in the repository.
- Implement guarded logical and physical PostgreSQL recovery tooling in the repository.
- Complete production deployment of the encrypted Google Drive destination for the scoped WordPress and PostgreSQL/MySQL/MongoDB workflows, including physical restore.
- Rehearse host-loss recovery path after offsite target is live.