# Osolix · Incident Response Playbook — v1

**Status:** Wave-B Foundation evidence · 2026-Q2
**Owner:** SRE Lead + CISO
**Audit dimensions:** #11 Security · #18 Observability · #25 Test/DR

This playbook is the operating manual for responding to any
production incident. Severity levels, on-call rota, communication
channels, and post-mortem cadence all defined.

## 1 · Severity definitions

| Severity | Definition | Target response | Target resolution |
|---|---|---|---|
| **P0 — Critical** | Multi-tenant data leak; full outage; data corruption | Pager fires within 5 min | 4 hours |
| **P1 — Major** | Single-tenant outage; AI engine wholesale failure; security probe successful | 30 min | 8 hours |
| **P2 — Minor** | Subset of users affected; degraded performance; non-critical feature broken | 2 hours | 24 hours |
| **P3 — Cosmetic** | UI bug; typo; non-blocking | Next business day | Within sprint |

## 2 · On-call rota

* **Primary on-call**: rotates weekly (engineering team).
* **Secondary on-call**: rotates weekly (different engineer).
* **Escalation**: SRE Lead → CTO → CEO (P0 only).
* **Customer-facing**: Customer Success Lead is paged on every P0 / P1.

## 3 · Communication channels

| Channel | Use |
|---|---|
| PagerDuty | Pager-trigger only |
| #incident-{id} (Slack) | Real-time engineering chat for the duration |
| Status Page (`/system-status`) | Customer-facing comms; updated every 30 min during P0 / P1 |
| Email (subscribed customers) | Initial alert + resolution comms |
| Twitter / LinkedIn | P0 only, after status-page update |

## 4 · Response steps (P0)

1. **Detect** — alert fires (PagerDuty / status-page health-check / customer report).
2. **Acknowledge** — primary on-call ack within 5 min.
3. **Triage** — confirm severity; spin up #incident-{id}.
4. **Mitigate** — restore service (rollback, toggle kill-switch, drain bad node).
5. **Communicate** — first status-page update within 15 min; cadence every 30 min.
6. **Investigate** — root-cause analysis after mitigation.
7. **Resolve** — status-page closure + customer email.
8. **Post-mortem** — within 5 business days; published in the customer's tenant audit log.

## 5 · Common runbooks

### 5.1 Database connection-pool exhaustion
* Symptom: 5xx spikes; "Cannot acquire connection" log entries.
* Mitigation: scale up DB connection pool; identify the long-running query.
* Permanent fix: add the query to the slow-query monitor; index missing.

### 5.2 Anthropic outage
* Symptom: `/api/ai/chat` returns Offline answers exclusively.
* Mitigation: confirm Anthropic status page; flip the `ai.online-killswitch` to OFF.
* Customer comms: optional banner (Online AI is disabled platform-wide; Offline still works).

### 5.3 Stripe webhook processing failure
* Symptom: subscription state drift between Osolix + Stripe.
* Mitigation: re-process webhooks from the dead-letter queue (`/admin/integrations` → DLQ tab).
* Permanent fix: investigate webhook-signature verification logic.

### 5.4 Cross-tenant probe succeeded (P0)
* Symptom: customer reports seeing another tenant's data.
* Mitigation: kill all sessions; audit query filter on the affected entity; restore from backup if data corrupted.
* CISO notified within 5 min; follow GDPR Art. 33 / 34 breach-notification clock.

### 5.5 Pen-test finding (post-engagement)
* Severity assigned; remediation timeline per CVSS:
  * Critical (9–10): 24h
  * High (7–8.9): 72h
  * Medium (4–6.9): 30 days
  * Low (0–3.9): next quarterly cycle.

## 6 · Post-mortem template

```
# Incident {ID} — {YYYY-MM-DD} — {Title}

**Severity:** P0 / P1 / P2 / P3
**Duration:** {start} → {end} ({total} mins of impact)
**Affected:** {tenants / regions / features}

## Timeline
- HH:MM — alert fired
- HH:MM — primary on-call ack
- HH:MM — mitigation begun (action: …)
- HH:MM — service restored

## Root cause
{Plain-English narrative}

## Contributing factors
- {Factor 1}
- {Factor 2}

## What went well
- {…}

## What didn't go well
- {…}

## Action items
| # | Action | Owner | ETA |
| 1 | {…} | {…} | {…} |
```

## 7 · Open items
* Wire status-page auto-update on PagerDuty incident state changes.
* Build the DLQ replay UI (currently CLI only).
* Quarterly chaos-monkey day to validate runbooks.
