Contents

Data Engineering › DataOps & Platform

CI/CD for Data Pipelines

Testing and deploying pipeline and model changes automatically.

Also known as: CI/CD for data, data pipeline CI, dbt CI, testing data pipelines in CI, analytics CI/CD

CI/CD for data applies continuous integration and delivery practices to data code: SQL models, pipeline definitions, schemas and transformations. Changes are tested automatically before merging, and deployed in a controlled, repeatable way, instead of someone running scripts by hand against production.

What CI usually does on a pull request

  1. Lint and format SQL and Python.
  2. Build the changed models in an isolated environment, such as a dedicated schema or a clone of production, using development or sample data (dev and prod environments).
  3. Run data tests on the results: uniqueness, not-null, accepted values, relationships, row-count sanity (data tests).
  4. Run unit tests on transformation logic with small fixed inputs.
  5. Compare outputs to production (a “data diff”): did row counts or values change unexpectedly because of this change?
  6. Check impact: which downstream tables and dashboards depend on what changed (lineage).
  7. Validate configuration, orchestrator definitions (DAGs import cleanly) and schema migrations.
# illustrative CI job
steps:
  - run: sqlfluff lint models/
  - run: dbt build --select state:modified+ --target ci     # build and test only changed models and their dependents
  - run: pytest tests/                                      # unit tests for Python code

What CD does

  • Deploys merged changes to production automatically: updates models and pipeline definitions.
  • Promotes through environments (dev → staging → production) where relevant.
  • Runs migrations for schemas and applies infrastructure changes through code (infrastructure as code).
  • Has a rollback or revert path. Reverting the code, then rerunning or restoring (rollback).

What’s different from application CI/CD

  • Data is part of the system. Code can be correct and the result wrong because of the data. You need data-level tests, not just code tests.
  • State and cost. Building models on large data in CI can be slow and expensive. Use sampled or limited data, clone techniques, and only build what changed.
  • Production data in CI is sensitive. Use masked or synthetic data in lower environments (data masking).
  • Backfills: a logic change may need historical data reprocessed, which isn’t a normal “deploy” step (backfill).
  • Consumers downstream (dashboards, ML) need warning about breaking changes (data contracts).

Habits

  • Everything in version control: models, tests, orchestration code, documentation.
  • Require passing CI to merge.
  • Keep deployments small and frequent.
  • Monitor after deploy: freshness, volume and test results (data observability).
  • Treat it as part of DataOps.