# Osolix · Production Deployment Runbook — v1

**Status:** Wave-B Foundation evidence · 2026-Q2
**Owner:** Chief Reliability Engineer + DevOps Lead
**Audit dimensions:** #25 Test/DR · #14 Performance · #18 Observability

This runbook is the canonical step-by-step for releasing a new Osolix
build to production. Every release goes through these checks; no
exceptions.

## 1 · Pre-flight (T-2 hours)

* [ ] All CI gates green: build · unit-tests · integration-tests · Playwright e2e · axe-core a11y · k6 perf-baseline · npm audit · dotnet vulnerable-package scan.
* [ ] Migration plan reviewed: any new EF migration is reversible OR has a rollback script committed.
* [ ] Feature-flag plan: any new feature is behind a flag; default OFF for the first 24 hours.
* [ ] Release notes drafted in the customer-facing change log.
* [ ] On-call rota confirmed: primary + secondary engineer for the next 24 hours.
* [ ] Stripe webhook signing keys rotated quarterly and verified.

## 2 · Deploy (T-0)

* [ ] Build the production image (Dockerfile + tag with git SHA).
* [ ] Push to the container registry.
* [ ] Apply EF migrations against staging FIRST; verify schema drift is zero.
* [ ] Apply EF migrations against production with the maintenance flag set (banner: "Brief maintenance window — ~5 minutes").
* [ ] Drain the active connections; rolling-update the API + watchdog cluster (1 instance at a time, 30s health-check between).
* [ ] Verify `/api/health` returns 200 from every node.
* [ ] Verify `/api/audit/overview` returns the expected score (regression detector — if pass-% drops > 5 points, abort + rollback).

## 3 · Post-deploy (T+30 min)

* [ ] Smoke test: the 12 critical Playwright paths run against production.
* [ ] AI Governance dashboard: latency-p95 unchanged or better.
* [ ] Status page: green across the board.
* [ ] Customer support tickets: monitor for any spike in the next 2 hours.

## 4 · Rollback (if anything goes wrong)

If any pre-flight or post-deploy check fails:

* [ ] Stop the deployment immediately.
* [ ] Revert the container image to the previous tag (registry has the last 10 builds).
* [ ] Run the migration rollback script if a schema change was applied.
* [ ] Verify the rollback restored full functionality.
* [ ] Open a P1 incident; post-mortem within 5 business days.

## 5 · DR drill (monthly)

* [ ] Spin up the DR-region replica + restore the latest backup.
* [ ] Run the smoke-test suite against the DR instance.
* [ ] Measure RTO + RPO actuals; file the drill report in `docs/restore-drills/{YYYY-MM}.md`.
* [ ] If actuals > target: file a P2 incident + remediation ticket.

## 6 · Quarterly cross-region failover

* [ ] Schedule the failover during a defined maintenance window (publish on `/system-status` 7 days ahead).
* [ ] Execute per `docs/46-cross-region-failover-drill.md`.
* [ ] Re-establish geo-replication after success.
* [ ] File the drill report.

## 7 · Annual audit cycle

* [ ] SOC 2 Type II field engagement: 12 months of evidence collection.
* [ ] ISO 27001 surveillance audit.
* [ ] Pen-test (external + internal) per `docs/34-pen-test-engagement-pack.md`.
* [ ] Re-attest every horizontal dimension on the platform-audit charter (security, compliance, performance, etc.).

## 8 · Open items
* Build the public release-notes page at `/welcome/release-notes`.
* Wire automatic incident creation on health-check failure → PagerDuty.
* Add the rollback-runbook trigger to the AI Governance dashboard so an on-call engineer can launch it in 1 click.
