Contents

Data Engineering › Orchestration & Pipelines

Reruns and Catch-Up

Re-running past intervals after a failure or code change.

Also known as: catch-up, catchup, backfills in orchestrators, rerunning past runs, clearing task runs, Airflow catchup

Pipelines don’t always run when they should. The orchestrator was down, a source was late, or a bug in the logic means last month’s results were wrong. Catch-up and reruns are how you get the right data for past intervals.

  • Catch-up: automatically running the intervals that were missed between a start date (or the last run) and now.
  • Rerun: deliberately running intervals again, after a failure or a code change.

Catch-up behavior

Many orchestrators have a setting for what to do after a pipeline is deployed or resumed after downtime. In Airflow, for example, catchup=True creates runs for every missed interval since the start date, and catchup=False only runs the latest one.

with DAG("orders_daily", start_date=datetime(2024, 1, 1), schedule="@daily", catchup=False):
    ...

Neither is always right. With catchup=True on a new pipeline with an old start date, you suddenly launch hundreds of runs at once. With False, missed days stay missing unless you backfill them.

Rerunning

  • Retry a failed task (and what depends on it) after fixing the cause.
  • Clear or re-trigger specific runs for a date range, from the UI or command line.
  • Backfill a range after a logic change or a source correction (backfill).

What makes it safe

  • Idempotent, partitioned runs. Each rerun for a date replaces that date’s output, and nothing duplicates (idempotent pipelines, partitioned runs). Without this, reruns are dangerous.
  • Use the logical date, not the current time.
  • Be careful with dependencies. Rerunning an upstream interval means downstream ones are now stale, so rerun them too (and in order).
  • Limit concurrency. Running 365 days in parallel can overload the source or warehouse. Cap parallel runs, and throttle.
  • Mind side effects: reruns that send emails, post to APIs or trigger alerts. Skip or protect those steps.
  • Mind cost: a large backfill can be expensive. Estimate before launching.

Good practice

  • Decide the catch-up policy deliberately for each pipeline.
  • Tell downstream users when historical data is being reprocessed, since numbers will change (data incidents).
  • Verify after reruns: row counts, totals and tests on the affected partitions.
  • Keep a record of what was rerun and why.
  • Ensure the inputs still exist. If you’ve deleted raw data, or the source can’t replay old states, you can’t reproduce history (landing zone).