# Osolix · Backup & Restore Drill Runbook — v1

**Status:** Wave-B Foundation evidence · 2026-Q2
**Owner:** Chief Reliability Engineer
**Audit dimensions:** #25 Test coverage & DR · #14 Performance & scalability
**Cadence:** Monthly drill; quarterly cross-region failover; annual full-DR exercise

## RTO / RPO targets

| Tier | RTO | RPO | Backup type |
|---|---|---|---|
| Production | 1 hour | 5 minutes | Continuous transaction-log backups + full daily |
| Demo | 4 hours | 1 day | Full daily |
| Sandbox | 24 hours | 1 day | Full daily |

## Backup configuration

* **Full backup:** every 24h at 02:00 UTC, retained for 35 days.
* **Differential:** every 6h, retained for 7 days.
* **Transaction-log:** every 5 minutes, retained for 35 days.
* **Off-site copy:** geo-redundant storage in a paired Azure region.
* **Encryption:** AES-256 with customer-managed key in Azure Key Vault.

## Monthly restore drill — scripted procedure

1. Spin up an isolated SQL Server instance (`osolix-drill-{YYYYMM}`) in the DR region.
2. Restore the most recent full backup + transaction-log chain to a point 4 hours before drill start.
3. Run the integrity verification suite:
   * `DBCC CHECKDB` — must return zero errors.
   * Row-count parity script — every table within ±0.01 % of production at drill timestamp.
   * Foreign-key parity script — every FK validates.
4. Spin up an Osolix.Api instance pointed at the restored database.
5. Run the smoke-test suite (`tests/smoke/post-restore.spec.ts`):
   * Login with the canary account.
   * Open `/admin/audit` — overview must return 200 with the expected score.
   * Open `/assets` — asset count matches the production snapshot.
   * Run a depreciation recompute — output matches the cached value within 0.001 %.
6. Record drill duration vs the 1h RTO target.
7. Tear down the drill instance.
8. File the drill report in `docs/restore-drills/{YYYY-MM}.md`.

## Restore-drill outcomes log

| Date | RTO actual | Issues | Sign-off |
|---|---|---|---|
| 2026-Q2 (drill #1) | TBD | TBD | TBD |

## Cross-region failover drill (quarterly)

1. Promote the secondary read-replica in the paired region to primary.
2. Re-point the API connection string + DNS to the new primary.
3. Verify the AI chat, asset list, and depreciation engine all answer within 500 ms p95.
4. Failback: re-establish geo-replication in the original direction.
5. File the drill report.

Reference: `docs/46-cross-region-failover-drill.md` (existing).

## Incident response — fast path

If production data corruption is confirmed:

1. Page on-call CISO + CRE; declare incident.
2. Snapshot the live database (point-in-time-restore the corruption boundary).
3. Restore to a parallel instance via the procedure above.
4. Validate via the smoke suite.
5. Cut over DNS / connection string after CISO + CTO sign-off.
6. Post-incident review within 5 business days; file management-letter finding.

## Open items
* Schedule the 2026-Q2 drill #1 — owner: CRE; ETA: end of month.
* Automate steps 1–5 of the monthly drill into a Bicep / Terraform module + GitHub Actions workflow.
