Data Engineering › Orchestration & Pipelines
Pipeline Scheduling
Running pipelines on a time schedule or when upstream data arrives.
Also known as: pipeline scheduling, cron for pipelines, scheduled runs, schedule interval, triggering pipelines
Pipelines run either on a time schedule (“every day at 02:00”) or when something happens (a file arrives, an upstream table updates).
Time-based schedules
Most orchestrators accept cron expressions:
0 2 * * * every day at 02:00
*/15 * * * * every 15 minutes
0 6 * * 1 Mondays at 06:00
(Fields: minute, hour, day of month, month, day of week.)
The idea people trip over: data interval vs run time
A daily run at 02:00 on June 2nd usually processes June 1st’s data. In some orchestrators (Airflow is the well-known example), a run is triggered at the end of the interval it covers. So it’s important to separate:
- Logical date / data interval: which data this run is responsible for (June 1st).
- Actual start time: when it really ran (June 2nd, 02:00).
Use the logical date inside your job, not “now”, so that reruns and backfills process the right day (partitioned runs).
Event-based and data-aware triggers
Instead of guessing a time, wait for the thing you depend on:
- Sensors that poll for a file or a table update (sensors).
- Data-aware scheduling: run when upstream datasets have been updated (data-aware scheduling).
- A message or API call that starts the run.
This avoids the fragile pattern “the upstream job usually finishes by 01:30, so run at 02:00”, which breaks the first time it runs late and your job processes incomplete data.
Practical guidance
- Schedule in UTC and avoid daylight-saving surprises (time zones).
- Don’t pile everything at midnight. Spread schedules out so you don’t create a thundering herd on sources and the warehouse.
- Avoid overlapping runs of the same pipeline unless it’s designed for it. Set limits on concurrent runs.
- Decide what happens to missed runs after downtime: skip them or catch up (catch-up and reruns).
- Set deadlines. If the report must be ready by 08:00, alert when it isn’t, instead of discovering it at 09:00 (data freshness).
- Give every schedule an owner, so that someone is responsible when it fails.
Application-side scheduled work follows similar rules (scheduled jobs).