Incident Response Checklist

Detect, triage, mitigate, resolve, postmortem — the steps that hold under pressure.

Checklist steps

  1. Declare incident

    Open the incident channel. Assign IC, Comms, Scribe.

  2. Triage severity

    Sev1: outage or data loss. Sev2: major degradation. Sev3: partial impact. Sev4: low impact.

  3. Update status page

    Within 5 minutes for Sev1/Sev2. Be honest, no jargon.

  4. Mitigate

    Stop the bleeding first. Root cause can wait.

  5. Resolve

    Confirm restoration with metrics and synthetic checks.

  6. Postmortem

    Blameless. Timeline, root cause, action items with owners.

Frequently asked questions

Who declares an incident?

Anyone can. Better to over-declare than under-declare.

Do all incidents need postmortems?

Sev1 and Sev2 always. Sev3/Sev4 are optional.