Contents

Data Engineering › Orchestration & Pipelines

Sensors

Tasks that wait for a condition, like a file arriving.

Also known as: sensor task, file sensor, waiting task, polling task

A sensor is a task in an orchestrator that waits for a condition to become true before letting downstream tasks run. Typical conditions: a file has landed, a partition exists, or an upstream table has new rows. The orchestrator checks (polls) the condition repeatedly, and proceeds once it is met or gives up after a timeout.

It replaces the fragile pattern “run at 02:00 because the upstream file usually arrives by then”. The first day the upstream is late, your job runs against incomplete data and nobody notices. A sensor makes the dependency explicit: the job starts when the data is actually there.

# illustrative, Airflow-style
wait_for_orders = FileSensor(
    task_id="wait_for_orders",
    filepath="/landing/orders/2024-06-01.csv",
    poke_interval=60,   # seconds between checks
    timeout=3600,       # give up after an hour
)
wait_for_orders >> transform >> load

Watch what “ready” means. A file can exist before it is fully written. Wait for a completion marker (often a _SUCCESS file) or a manifest, not just for the path to appear. An upstream table can likewise have a partition that exists but is still being filled.

Sensors cost resources while they wait. A polling sensor can hold a worker slot for as long as it waits, so dozens of them can starve the tasks that do real work. Many orchestrators offer a mode that frees the slot between checks or a deferrable/event-driven variant; prefer those for long waits. Always set a timeout, and make a timeout fail the run (or raise an alert) rather than silently succeeding and letting downstream jobs read missing data.

When not to use one: if your orchestrator supports data-aware or asset-based triggers, those express “run when this dataset is updated” without polling (data-aware scheduling, asset-based orchestration). A sensor is the right tool when the condition is something the orchestrator cannot track by itself, such as an external file drop.

Related: pipeline scheduling, task dependencies, idempotent pipelines, data freshness.