Data Engineering › Data Quality & Observability
Data Incident
A data quality failure that reaches users, and how to respond.
Also known as: data quality incident, data outage, bad data incident, data downtime incident
A data incident is a data problem that affects the people or systems relying on the data: a dashboard showing wrong revenue, a model trained on corrupted input, a report that’s a day late, customers emailed with the wrong numbers. The pipeline might be “green”, but the data is wrong, missing or stale.
Common forms
- Wrong data: duplicates, a bad join, a changed definition, a unit mix-up, a source bug.
- Missing or incomplete data: a failed load, a partial day, a dropped column.
- Stale data: a table didn’t update (data freshness).
- Schema or meaning change upstream that silently altered results (schema drift).
- Exposed sensitive data (a privacy incident, handled by security and compliance processes).
Responding, step by step
- Detect. Via a test or monitor (ideally), or a stakeholder noticing (data observability, data tests).
- Triage the impact. Which tables are affected? Use lineage to find downstream dashboards, models and exports, and who uses them (data lineage). How wrong, and since when?
- Communicate early. Tell affected consumers what’s known, what to avoid using and when you’ll update. A short, honest message beats silence. Mark affected data or dashboards as suspect.
- Stop the bleeding. Pause the pipeline, block publication of bad data (circuit breaker for data), or revert to the last good version (table time travel and snapshots help).
- Find the cause (debugging a failed pipeline, root cause analysis).
- Fix and recover. Correct the logic or data, then reprocess the affected range with idempotent reruns (backfill, reruns), and verify the numbers against an independent source.
- Confirm and close: tell stakeholders the data is correct again and what changed.
- Learn. Hold a blameless postmortem, and add the test, check or contract that would have caught it.
Preventing and shrinking them
- Test at critical points, and block bad data before it reaches consumers (write-audit-publish).
- Monitor freshness, volume and distributions, and alert on anomalies (data anomaly detection).
- Know ownership of each table, so someone responds (data ownership).
- Agree on contracts with upstream producers (data contracts).
- Measure how long data was bad and how long detection and recovery took (data downtime).
The most damaging data incidents are the ones nobody notices. A wrong number used in decisions for weeks costs far more than a loud failure.