Contents

Data Engineering › Orchestration & Pipelines

Task Dependencies

Which steps must finish before others can start.

Also known as: task ordering, upstream and downstream, DAG dependencies

Task dependencies say which steps of a pipeline must finish before others may start. If task B needs the output of task A, then B depends on A: A is upstream, B is downstream.

extract_orders ──┐
                 ├──▶ join_orders_customers ──▶ build_revenue_table ──▶ refresh_dashboard
extract_customers┘

The two extracts have no dependency on each other, so an orchestrator can run them in parallel. The join waits for both. This graph of tasks and dependencies is a DAG.

Orchestrators let you declare them, for example in Apache Airflow:

# e.g. Airflow-style syntax
[extract_orders, extract_customers] >> join_orders_customers >> build_revenue_table

Why get them right

  • Correctness. Without a dependency, a downstream step may read half-written or yesterday’s data.
  • Speed. Independent tasks run in parallel; missing this makes pipelines slow for no reason.
  • Failure handling. If extract_orders fails, downstream tasks are not run on bad data, and when you fix and re-run it they continue from there.

Classic mistakes

  • Using the clock instead of a dependency, like “job B starts at 2:00 because job A usually finishes by then”. The first slow day, B reads incomplete data. Make B depend on A (or on a sensor for the data it needs).
  • Hidden dependencies, such as a task reading a table that another pipeline writes, which the orchestrator can’t see. Make these explicit.
  • Over-linking. Chaining everything in a line when steps are independent makes runs slow and failures cascade needlessly.
  • Circular dependencies aren’t allowed; a DAG has no loops.

Dependencies decide the order; also make each task safe to re-run. See idempotent pipeline.