Data Engineering › Working as a Data Engineer
Data On-Call
Being responsible for pipeline failures and late data.
Also known as: data on call, pipeline on-call, on-call for data, data engineer on-call
Data on-call is being the person responsible for the data pipelines when they fail or run late: broken tasks, missing partitions, data that arrives hours after it should, or numbers that look wrong. It is the data version of on-call for software services.
The classic mistake is alerting on everything. If every task failure pages someone, the rotation learns to ignore pages, and a real outage gets lost in the noise. A flaky source that retries and succeeds is not an incident. Alerts should fire on things a person must act on now: data that has not arrived by its promised time, a table that failed its tests, or a downstream dashboard reading stale values.
How data on-call differs from software on-call
- Failure is often silent. A service that is down is obvious; a pipeline that loads yesterday’s numbers unchanged may page no one while a decision is made on wrong data. Watch freshness, volume and schema, not just job success.
- “Late” is a first-class failure. Many issues are about timeliness against a data SLA, not a crash (data downtime).
- You often reprocess rather than roll back. Fixing the cause and rerunning the affected partitions is normal (reprocessing history).
- The cause may be upstream. A schema change in a source system, or a partner’s API, can break you through no fault of your own.
Making it humane and useful
- Alert on symptoms that matter. Tie alerts to consumer promises and severity, and route non-urgent ones to a queue (alerting).
- Write runbooks. The first five minutes of debugging a failed pipeline should not require the original author.
- Retry transient failures automatically. Distinguish retryable errors from real ones (pipeline retries).
- Have clear ownership and escalation, so it is obvious who fixes a broken source.
- Follow up with blameless postmortems and fix causes, not people.
Rotations should be shared across a team, with realistic expectations about out-of-hours coverage. Being paged at 3am for a non-critical report is a sign the alerting needs work, not that the engineer is unlucky.