Contents

Data Engineering › Data Engineering Foundations

Data Engineering

Building the systems that collect, move, store and prepare data for analysis and ML.

Also known as: DE, data engineer, data platform engineering

Data engineering is building and running the systems that collect, move, store and prepare data so that analysts, data scientists, ML models and applications can use it reliably. Data engineers rarely produce the final chart or model; they make sure the right data arrives, in the right shape, on time, and can be trusted.

A typical example: orders live in a production database, ad clicks in a third-party API, and app events in log files. A data engineer builds pipelines that copy these into a warehouse, clean and combine them, and keep that running every day, so an analyst can ask “revenue by region last week” with one query.

What the work looks like

  • Ingest data from databases, APIs, files and event streams. See data ingestion.
  • Store it where it can be queried cheaply: a warehouse, a data lake, or both.
  • Transform it from raw and messy into clean, documented tables. See data transformation.
  • Schedule and monitor the pipelines, and fix them when they break at 3 a.m.
  • Protect and document the data: access, quality checks, ownership.

Most of it is plain software engineering: SQL, Python, version control, testing, and operating systems in production.

How it differs from nearby roles

Software engineers build applications; data engineers build the plumbing that moves the data those applications produce. Analysts and data scientists consume the output. See data roles.

What goes wrong

The classic mistake is jumping to dashboards or machine learning on top of unreliable, inconsistent data. The model is only as good as what feeds it. That is the idea behind the hierarchy of needs: reliable collection and storage come first. Another common problem is building a pipeline that works once but can’t be re-run safely, which turns each failure into manual cleanup.