Contents

Data Engineering › DataOps & Platform

DataOps

Applying DevOps practices like CI, testing and monitoring to data.

Also known as: Data Ops, data operations, DataOps practices, data engineering operations

DataOps is applying the ideas of DevOps (automation, version control, testing, continuous delivery, monitoring and collaboration) to the creation and operation of data pipelines and analytics. The goal: deliver trustworthy data faster and more reliably, with fewer firefights.

It’s more a way of working than a tool, and one of the “undercurrents” of the data engineering lifecycle (data engineering lifecycle).

The practices

PracticeWhat it means for data
Version controlSQL, pipeline code, tests, schemas and docs all live in Git, with code review
Automated testingData tests, unit tests for logic, schema checks, run on every change (data tests)
CI/CDChanges are built, tested and deployed automatically (CI/CD for data)
EnvironmentsDevelopment and production are separate, so experiments never touch real consumers (dev and prod environments)
Monitoring and observabilityFreshness, volume, quality and costs are watched, with alerts (data observability)
ReproducibilityAnything can be rebuilt from code and raw data. Pipelines are idempotent (idempotent pipelines)
Infrastructure as codeWarehouses, pipelines and permissions are defined in code (infrastructure as code)
Collaboration and feedbackShort cycles with the people who use the data, and clear ownership (data ownership)
Incident practiceOn-call, runbooks and blameless reviews for data failures (data incidents)
Cost awarenessQuery and warehouse spend is monitored and managed (warehouse cost management)

The problems it responds to

  • Pipelines changed by hand in production, with no history of who changed what.
  • “It works on my laptop” analytics with no tests.
  • Breakages found by angry stakeholders.
  • Slow, risky releases, so change is avoided and the data goes stale.
  • Nobody sure who owns a table or whether it’s still used.

Starting small

You don’t adopt it all at once:

  1. Put everything in version control and require reviews.
  2. Add a few critical tests, and run them automatically.
  3. Separate dev and prod.
  4. Automate deploys.
  5. Add monitoring and alerts for the important tables.
  6. Review failures, and fix causes.

The measure of success is outcomes: fewer data incidents, faster safe changes, and more trust in the numbers. Compare with how software teams measure delivery performance (DORA metrics).