# Osolix · Cross-Region Active-Passive Failover Drill

> **Owner:** Automation Manager (DR + reliability)
> **Co-owners:** CTO · CISA Manager
> **Closes:** SOX §6 backlog item #9 · ISO 22301 §8.4 (business-continuity exercises)
> **Cadence:** half-yearly · co-scheduled with the §5 quarterly DR drill so
> the team rehearses the bigger blast-radius scenario twice a year
> **Target:** 2026-Q4 (paired with cloud-region rollout)
> **Status (2026-05-03):** design phase — script + runbook authored,
> infrastructure rollout pending Phase-4 multi-region milestone

This document extends the `dr-drill.ps1` script (§5 closure) from
"restore the latest backup into a fresh DB on the same SQL host" to the
full failover scenario: **the primary region is unreachable, secondary
takes over, customers see no more than 30 minutes of disruption.**

The drill is *non-destructive* against production — it runs in a
dedicated `failover-rehearsal` environment that mirrors prod's region
pair. That separation is what lets us rehearse without putting tenants
at risk.

---

## §1 · Scope

### 1.1 In scope

- **SQL Server failover** — primary region database fails over to a
  secondary region read-replica that is promoted to read-write.
- **API tier failover** — the load balancer routes traffic to the
  secondary-region API instances after a configurable health-check
  latency threshold.
- **DataProtection key-ring availability** — assert the secondary region
  can decrypt the active ring (key ring must be replicated cross-region
  via Azure Key Vault georeplication or equivalent).
- **Storage failover** — uploaded files (asset photos, document
  attachments) are accessible from the secondary region's blob storage
  replica.
- **DNS / traffic-manager cutover** — customer-facing host (`*.osolix.com`)
  resolves to the secondary region's frontend.

### 1.2 Out of scope (explicit)

- **AI providers (Anthropic, Ollama)** — third-party / customer-hosted;
  failover handled by the providers themselves.
- **Customer ERP systems** — Osolix only validates the integration surface
  comes back online; the customer's ERP is their responsibility.
- **Email delivery** — relies on the SMTP/M365 provider's own redundancy.
- **Active-active** — explicitly active-passive only. Active-active is a
  separate effort (Phase 4+).

---

## §2 · Architecture (target state)

```
                     ┌──────────────────────────────────┐
                     │  Customer DNS (osolix.com)       │
                     │  Azure Front Door / AWS Route 53 │
                     └──────────────┬───────────────────┘
                                    │
                  ┌─────────────────┴──────────────────┐
                  │ Health-check failover (60s window) │
                  └─────┬─────────────────────┬────────┘
                        │ Primary OK          │ Primary unhealthy
                        ▼                     ▼
   ┌─────────────────────────────┐    ┌─────────────────────────────┐
   │   PRIMARY REGION (UAE-N)    │    │  SECONDARY REGION (EU-W)    │
   │                             │    │                             │
   │  • App Service (3x)         │    │  • App Service (2x cold)    │
   │  • SQL Server (RW)          │◀──▶│  • SQL Server (RO replica)  │
   │  • Blob Storage (Primary)   │◀──▶│  • Blob Storage (RA-GRS)    │
   │  • Key Vault (Primary)      │◀──▶│  • Key Vault (Replica)      │
   │  • Service Bus (Primary)    │◀──▶│  • Service Bus (Failover)   │
   └─────────────────────────────┘    └─────────────────────────────┘
       Always-replicating async backbone
```

**RTO target:** ≤ 30 minutes (DNS TTL + DB promotion + cold-start)
**RPO target:** ≤ 5 minutes (async geo-replication lag under steady state)

---

## §3 · Drill procedure

### Phase A · Preparation (T-7 days)

1. CISA Manager + Automation Manager schedule the drill window.
   Customers are notified at T-72h (best-effort; no actual customer
   impact expected).
2. Verify replica health:
   - SQL: `SELECT * FROM sys.dm_geo_replication_link_status` shows
     `replication_state_desc = CATCHUP` and lag ≤ 60 s.
   - Blob Storage: `Get-AzStorageAccount -Name osolix-prod` shows
     `SecondaryEndpoints` populated and `LastSyncTime` within 5 min.
   - Key Vault: secondary endpoint reachable from secondary region's
     deployment slot.
3. Take a fresh backup of the primary database (belt + braces).
4. Run `database/scripts/dr-drill.ps1` standalone first to confirm the
   smaller drill is green before the bigger drill.

### Phase B · Execution (T-0)

| Step | Action | Owner | Expected duration |
|---|---|---|---|
| 1 | Mark the rehearsal slot as "exercise in progress" on the status page | Automation Mgr | 1 min |
| 2 | Force-fail primary's API health probe (return 503 from `/health` for 90 s) | Automation Mgr | 90 s |
| 3 | Front Door / Route 53 detects unhealthy primary, fails traffic over to secondary | (automatic) | 60 s |
| 4 | Promote SQL secondary to read-write: `ALTER DATABASE [Osolix] SET PARTNER FAILOVER` (planned migration to AGs) | DBA | 3 min |
| 5 | Update API config in secondary to point at the now-RW DB | Automation Mgr | 30 s |
| 6 | Run smoke suite: list assets · submit MR · run depreciation period preview · query audit log | Automation Mgr | 5 min |
| 7 | Verify DataProtection key ring reachable from secondary (decrypt a known ERP credential) | CISA Mgr | 2 min |
| 8 | Verify blob storage reads succeed (load a known asset photo) | Automation Mgr | 1 min |
| 9 | Note total failover time → compare against 30-min RTO target | All | continuous |

**Total expected RTO:** 8–12 minutes under nominal conditions; the 30-min
target gives headroom for the cold-start of secondary App Service
instances and any DNS-cache stragglers.

### Phase C · Failback (T+1 hour)

1. Restore primary's `/health` endpoint.
2. Front Door / Route 53 re-detects primary as healthy and shifts
   traffic back (subject to TTL).
3. Re-establish replication (former primary becomes secondary):
   `ALTER DATABASE [Osolix] SET PARTNER FAILOVER` again.
4. Confirm both regions report `CATCHUP` state.
5. Decommission the rehearsal status-page entry.

### Phase D · Evidence & post-mortem (T+1 day)

1. Drill report written to `docs/dr-drill-runs/<yyyymmdd>-cross-region.md`
   covering:
   - Actual RTO (timestamps per step)
   - RPO (last replicated transaction timestamp at moment of failover)
   - Smoke-suite results
   - DataProtection decryption confirmation
   - Any deviation from runbook
2. Findings reviewed at next CISA Manager quarterly walkthrough.
3. Improvements (e.g. DNS TTL tuning, App Service warm-pool sizing)
   filed as Q-+1 backlog items.

---

## §4 · Pass / fail criteria

A drill passes when **all six** are true:

| # | Criterion | Measurement |
|---|---|---|
| 1 | RTO ≤ 30 minutes | Step-1 → smoke-suite green |
| 2 | RPO ≤ 5 minutes | Last-replicated-txn timestamp at failover |
| 3 | All smoke-suite endpoints return 200 OK from secondary | Automated via `dr-smoke.ps1` (planned companion script) |
| 4 | DataProtection ring decrypts a known credential in secondary | Manual sanity per §3 step 7 |
| 5 | Blob storage reads succeed in secondary | Manual sanity per §3 step 8 |
| 6 | Failback restores `CATCHUP` replication state | Phase C step 4 |

Any single criterion failing converts the drill to a **finding** that
must be remediated before the next drill window.

---

## §5 · Smoke-suite scope

The companion `dr-smoke.ps1` script (authored 2026-06-14, Decision 0079 —
`database/scripts/dr-smoke.ps1`; tune the login DTO + depreciation-preview
body to the rehearsal tenant before the first live drill) hits the following
representative endpoints in the secondary region:

| Endpoint | Expected | Why |
|---|---|---|
| `GET /health` | 200 + `{ status: "healthy" }` | basic liveness |
| `POST /api/auth/login` (test tenant) | 200 + JWT | identity stack OK |
| `GET /api/assets?pageSize=10` | 200 + 10 rows | DB read path OK |
| `POST /api/maintenance-requests` (test tenant) | 200 + new ref# | DB write path OK |
| `GET /api/erp-dlq?status=Failed` | 200 + array | tenant scope + audit OK |
| `POST /api/depreciation/preview` (test tenant) | 200 + period | depreciation engine OK |
| `GET /api/audit-log?take=5` | 200 + 5 rows | audit trail readable |
| `GET /scim/v2/ServiceProviderConfig` | 200 + JSON | SSO/SCIM stack OK |

Smoke runs in < 60 s in steady state.

---

## §6 · Dependencies & blockers (as of 2026-05-03)

The drill cannot **execute** until the underlying infrastructure ships:

| Dependency | Owner | ETA |
|---|---|---|
| Cloud-region pair provisioned | CTO + Infra | Phase 4 milestone (2026-Q4) |
| SQL Always-On Availability Group cross-region | DBA | Phase 4 |
| Blob storage RA-GRS replication enabled | Infra | Phase 4 |
| DataProtection key ring → Key Vault georeplication | CISA Manager + Infra | Phase 4 |
| Front Door / Route 53 health-check failover policy | Infra | Phase 4 |
| Status-page integration | Automation Mgr | Phase 4 |

This document, `dr-drill.ps1` (single-region restore), `dr-smoke.ps1` (§5
secondary-region smoke suite), and `deployment/azure/bicep/commercial-secondary.bicep`
(secondary SQL server + SQL auto-failover group + standby compute + Front Door
active-passive cutover) are the **design floor** (Decision 0079); they unblock
the implementation team to land the multi-region rollout with concrete,
deployable artifacts rather than a vague "we'll figure it out". The secondary
Bicep is deliberately **not** wired into `deploy-commercial.yml` — provisioning
it is the founder infra-spend decision that closes risk R-06.

---

## §7 · Sign-off

| Date | Reviewer | Outcome |
|---|---|---|
| 2026-05-03 | CISA Manager | Drill design published. Closes SOX §6 backlog item #9 (design only). Execution unblocked once Phase-4 infrastructure ships in Q4-2026. |
