Data Quality & Observability — Interactive Runbook Templates

Interactive runbook templates for common data quality failures. Capture detection details, follow guided remediation steps, assign ownership, record SLAs and communications, and save a complete incident record for tracking and post-incident review.

Interactive Tool

Data Quality Runbook — Incident Triage & Remediation

Use this interactive runbook when a data quality issue is detected. Capture triage details, follow remediation steps, assign owners, and record outcomes so incidents are resolved faster and learning is preserved.

This form provides guided templates for common failures (missing/late data, schema drift, unexpected volume changes) and a standardized record for SLAs, notifications, and post-incident actions.

Provide a unique identifier (or leave to auto-generate if integrated with an alert/ticket system).
e.g., 2026-08-26T14:32:00Z
Monitoring rule, alerting system, user, or model name.
Select the template that best matches the failure to surface recommended checks.
Severity guides escalation, communications, and SLA targets.
List identifiers and known owners; include lineage hints if available.
Which KPIs, dashboards, or downstream models are impacted?
Describe what the alert showed, which metric changed, or what data is missing.
Check actions already performed to avoid duplication.
Person or team responsible for remediation and communications.
Target hours to acknowledge the incident.
Target hours to resolve the incident.
Step-by-step actions taken, commands, queries, and checkpoints. Include links to logs, queries, or notebooks.
Describe fixes, backfills, patches, configuration changes, or other mitigations.
List Slack channels, mailing lists, product owners, and customers to contact.
Document where and when notifications were posted and link to messages/tickets.
JIRA, ServiceNow, PagerDuty, or other ticket URL if used.
When the incident was resolved or reverted.
What fixed the issue? Include root cause hypothesis and key evidence.
List permanent fixes, tests, monitoring improvements, and the owners responsible with due dates.
Suggestions to avoid recurrence and improvements to runbook or monitoring.
Typical post-incident items to close the loop.
Set to yes when all corrective actions are assigned or completed.
You can explore this tool now. Sign in or create an account to save your responses and return to them later.
Make this tool part of your work

Save a personal copy, bring it to your team, or tailor the questions and workflow to fit what you are hungry to improve.

Member customization and team collaboration are coming soon.

Discussion

Comments and conversation will live here.