Contents

AI & Data › Data Engineering Basics · also in Orchestration & Pipelines

Data Pipeline

A sequence of steps that moves and transforms data.

Also known as: data pipelines, pipeline, ETL pipeline, data flow pipeline

A data pipeline is an automated sequence of steps that moves data from where it’s produced to where it’s used, transforming it on the way: extract from sources, clean and combine, load into a warehouse, publish for reports.

sources ──► ingest ──► store raw ──► clean & transform ──► serve (dashboards, ML, apps)

Pipelines can run in batches on a schedule or process data continuously (batch vs stream).

What turns a script into a pipeline

A one-off script copies data once. A pipeline is dependable, repeated and observable:

  • Scheduled or triggered automatically (scheduling).
  • Ordered with dependencies between steps (pipeline DAG).
  • Idempotent: rerunning a step gives the same result without duplicates (idempotent pipelines).
  • Handles failure: retries, alerts, partial recovery (retries).
  • Tested and monitored: data checks, freshness, row counts (data tests).
  • Versioned and deployed like code.
  • Documented, with an owner.

Typical problems

  • A source changes its schema and the pipeline breaks (schema drift).
  • Late or missing data produces wrong numbers silently.
  • One slow step delays everything downstream.
  • Nobody knows what depends on a table, so changes break things.
  • Costs creep up as data grows.

Design for failure and change from the start, as those are the main reasons pipelines need human attention.