Data Quality Incident Report & Triage Form

A structured, savable incident form to capture data quality defects, assess impact, guide rapid triage, assign ownership and SLAs, track remediation steps, and collect post-incident review actions.

Interactive Tool

Data Quality Incident Report & Triage Form

This form helps teams capture data quality incidents in a consistent way so detection-to-resolution time shrinks, ownership is clear, and corrective actions are tracked. Save the form to record the incident; your submissions can later be reviewed, triaged and reported. Provide as much context as possible — example records, affected datasets/pipelines, and links to alerts or dashboards materially speed investigation.

Who is reporting this? Include name and best contact (email or Slack).
Use ISO 8601 if possible (e.g., 2026-08-26T15:04:00Z).
List dataset and pipeline names, unique IDs, and approximate locations (table, topic, bucket, schema). Include links to catalog entries if available.
Describe what you observed (e.g., missing rows, unexpected nulls, schema mismatch), and paste 1–5 example records or query results that illustrate the issue.
Which systems, reports, models, customers, or business processes are affected? Quantify where possible (e.g., number of reports, customers, percent of rows).
Choose the appropriate severity to drive SLA and escalation.
Quickly classify the likely problem area to route to the right team.
E.g., paused downstream jobs, switched to cached reports, disabled alerts, or manual workaround steps. Include who implemented the mitigation.
Assign an owner and provide contact. This owner coordinates the fix and updates status.
Enter expected time-to-fix in hours based on severity and business needs.
Update as the incident progresses.
State the most likely cause based on initial investigation. E.g., bad upstream file, schema change, transformation bug.
List the queries, checks, or tools you’ll run (or already ran) to confirm root cause — e.g., replay jobs, inspect source files, review change logs, check lineage.
Choose the actions you will or have taken to limit impact.
Include links to dashboards, alert IDs, monitoring graphs, catalog entries, or external incident tickets.
Helps measure observability coverage; if no, note where lineage is missing.
If yes, include which snapshot/time and who performed it.
Describe proposed preventative changes (tests, data contracts, monitoring, owner changes) to avoid recurrence. To be completed during post-incident review.
Record what messages were sent, to whom, and when. Useful for audits and stakeholder confidence.
Summarize the confirmed root cause, steps taken to remediate, and verification checks used to confirm resolution.
Consent helps teams reuse lessons learned; sensitive incidents may be limited.
You can explore this tool now. Sign in or create an account to save your responses and return to them later.
Make this tool part of your work

Save a personal copy, bring it to your team, or tailor the questions and workflow to fit what you are hungry to improve.

Member customization and team collaboration are coming soon.

Discussion

Comments and conversation will live here.