Contents

Data Engineering › Orchestration & Pipelines

Retries and Failure Handling

Retrying transient failures and alerting on real ones.

Also known as: retries, failure handling in pipelines, task retries, retry policy, pipeline failure alerts

Pipelines fail. Networks hiccup, a source API returns a 503, a warehouse is briefly overloaded, a cloud worker is reclaimed mid-task. Good failure handling separates problems that fix themselves if you try again from problems that need a person.

FailureExampleWhat to do
TransientTimeout, rate limit, brief outageRetry automatically
PermanentBug in the code, missing permission, a column that no longer existsDon’t retry endlessly. Alert and fix
Data problemSource delivered garbage or nothingStop downstream tasks, quarantine, alert
# Airflow-style task settings
default_args = {
    "retries": 3,
    "retry_delay": timedelta(minutes=5),
    "retry_exponential_backoff": True,     # wait longer after each failure
    "execution_timeout": timedelta(hours=1),
}

Rules

  • Retries require idempotency. If a task half-finishes and runs again, the result must be the same as one clean run, with no duplicate rows (idempotent pipelines). Otherwise retries make things worse.
  • Limit retries and use backoff, as with any retry (retry with backoff). Hammering a struggling source doesn’t help it recover.
  • Set timeouts so a hung task fails and can retry, instead of blocking forever.
  • Alert on final failure, not on every retry (that creates noise, see alerting). Include what failed, which run and a link to the logs.
  • Alert on missing runs and staleness too: no failure message because nothing ran is still a failure (data freshness).
  • Don’t retry what can’t succeed: an authentication error or an invalid query won’t fix itself.
  • Stop downstream work when upstream failed or data checks failed, so you don’t publish wrong numbers (circuit breaker for data).
  • Partial failures: make sure a failed run can be resumed or fully redone without leftover half-written output (write to a temporary location, then swap in).

When a person gets involved

Failures that survive retries need a person: look at logs, find the cause and rerun (debugging a failed pipeline). Keep a runbook for the common ones.

The goal isn’t “never fail”; it’s “fail loudly, safely and recoverably”.