Data Incident Postmortem & Root Cause Runbook

A practical, structured postmortem runbook that combines clear guidance with a reusable interactive postmortem form. Capture incident intake, triage, root cause analysis (including guided 5 Whys), corrective actions with owners and SLAs, communications, and follow-up monitoring so teams recover faster and reliably learn from incidents.

Interactive Tool

Data Incident Postmortem & Root Cause Runbook

Purpose

This runbook helps teams capture a clear, actionable postmortem after a data, analytics, or model incident. Use the form to record what happened, how it was diagnosed and mitigated, the root cause using structured techniques, and the corrective actions with owners and timelines. Structured entries make follow-up easier and enable team learning.

Tips before you start

  • Prefer facts: record timestamps, who performed actions, and concrete evidence (logs, query results, dashboards).
  • If possible, link to the original incident/ticket and snapshots of relevant metrics or dashboards.
  • Keep communications factual and focused on impact, containment, and next steps.
Unique identifier (ticket, alert ID, or internal incident number).
Who first reported or observed the incident.
Best contact for clarifying details.
Earliest timestamp when the problem is known to have existed (approximate is OK).
When the problem was detected or alerted.
Elapsed minutes between earliest known time and detection.
Elapsed minutes from detection to restoration of acceptable service or data quality.
List affected services, pipelines, databases, models, dashboards, APIs, etc.
Concise description of business and user impact (who, what, and degree).
If known, provide counts (users, transactions, rows).
Use your team’s severity scale.
What was done to contain damage or mitigate impact (commands, rollbacks, feature toggles, temporary fixes). Include timestamps and actor names.
Quick check of whether a recent change may be related.
Confirm whether producer systems were healthy or failing.
Which tests ran and what they showed (schema checks, null rates, distribution changes).
If yes, describe what was rolled back and when.
URLs to dashboards, alert pages, log queries, artifacts, or saved outputs supporting the investigation.
One- to three-sentence summary of the root cause once determined. If unknown, state hypotheses.
Start with the immediate cause (what happened?).
Why did that happen?
Deeper causal factor.
Deeper causal factor.
The deeper organizational, process, or technical weakness uncovered.
Optional: sketch of a causal tree, contributing factors, timelines, or test results.
Describe recommended fixes to address root cause(s). Prefer specific, testable actions.
List owners and target completion dates or SLA windows. Use rows like: Team/Person — Action — Due (YYYY-MM-DD).
What to monitor, metrics/thresholds, and for how long (e.g., 7 days of elevated TTL checks).
How long enhanced monitoring should run.
Copy or summary of customer/internal communications: who, channel, message summary, timestamp.
List teams, managers, or external parties notified about the incident and updates.
Scheduled review to discuss the postmortem and confirm actions.
Short, actionable lessons the team should remember for future prevention or detection.
Describe other prevention ideas not listed above.
When this postmortem was finalized or published.
Who will ensure corrective actions are completed.
Set to yes when follow-ups are complete and monitoring window ended.
Any other context, references or next steps.
You can explore this tool now. Sign in or create an account to save your responses and return to them later.
Make this tool part of your work

Save a personal copy, bring it to your team, or tailor the questions and workflow to fit what you are hungry to improve.

Member customization and team collaboration are coming soon.

Discussion

Comments and conversation will live here.