Contents

Data Engineering › Working as a Data Engineer

Debugging a Failed Pipeline

Finding which task broke, why, and what data it affected.

Also known as: pipeline failure debugging, troubleshooting data pipelines, pipeline broke, debugging data jobs

An alert says a pipeline failed, or someone says a dashboard looks wrong. A systematic approach beats guessing.

1. Find out what broke, and since when

  • Which task failed, in which run, for which data interval? The orchestrator’s graph shows it.
  • When did it last succeed? What changed since? (a code deploy, a config or credential change, an upstream release, a schema change).
  • Is it new, or has it been failing quietly?

2. Read the error

Open the task logs and read the first real error, not the cascade after it (reading error messages, reading production logs). Then classify:

TypeTypical signs
InfrastructureTimeout, out of memory, worker lost, disk full, no capacity
AccessPermission denied, expired token or password, rotated key
Source problemSource down, API error or rate limit, file missing or empty
Schema / data changeColumn not found, type mismatch, new value that breaks a parse (schema changes)
Code bugException from your logic, a bad join or an unhandled null
Data volume / skewThe job suddenly runs far longer or out of memory after a spike

3. Judge the impact

  • What data is affected, and what is downstream of it? Use lineage to find consumers.
  • Is bad data already published? A failure that stops the pipeline is annoying. Wrong numbers served silently are worse.
  • Tell the affected people early, with what you know and when you’ll update them (data incidents).

4. Fix and recover

  • Fix the cause, not just retry. If it was transient, retry and note why.
  • Rerun safely. This requires idempotent tasks, so that rerunning doesn’t duplicate.
  • Backfill any missed intervals (reprocessing history).
  • Verify the result: row counts, spot checks, the original failing check now passing.

5. Prevent a repeat

  • Add a test or alert that would have caught it earlier.
  • Update the runbook, and fix the root cause (a contract with the source team, better retry settings).
  • For serious incidents, write a short blameless review.

Habits that help: keep logs and run history accessible, include run IDs and row counts in log messages, and write down what you tried. When on call, see data on-call.