Contents

Data Engineering › DataOps & Platform

Development and Production Data Environments

Separate schemas or warehouses so development never touches production data.

Also known as: dev and prod data environments, data environments, separate dev and prod, development warehouse, staging data

Software teams keep development separate from production (environments). Data teams need the same separation, so that someone testing a new model doesn’t overwrite a table that the CFO’s dashboard reads.

Common ways to separate

ApproachHowNotes
Separate schemas in one warehouseanalytics_dev_ana vs analyticsCheap and common. Each developer gets a personal schema
Separate databases or projectsdev vs prodStronger isolation and permissions
Separate warehouses/accountsA fully separate environmentStrongest isolation, more setup and cost
Zero-copy clones or branchesInstantly clone production data for a testCheap realism, where the platform supports it

A typical flow: develop in your own schema against sample or cloned data, open a pull request, run CI in an isolated schema, then deploy to production through automation.

The data question

Realistic testing needs realistic data, which creates tension:

  • Production data is sensitive. Don’t copy raw personal data into development, where more people have access and controls are weaker. Use masked or synthetic data (data masking), or restrict access tightly.
  • Full production volume is expensive. Use samples or limited date ranges for development (sampling data), and full-scale runs only where needed (staging, performance tests).
  • Small samples miss real edge cases (skew, nulls, odd values). Include known problem cases, and test on realistic volume before release.
  • Cloned or read-only access to production sources can let dev models read real source data safely, while writing only to dev schemas.

Rules that prevent accidents

  • Developers write only to dev. Production write permissions belong to the deployment process, not to individuals (least privilege).
  • Make the target environment obvious (names, prompts, banners). Running a destructive command on the wrong environment is a classic mistake.
  • Parametrize the environment in code (the schema name comes from configuration), with no hard-coded production names (configuration).
  • Clean up old development schemas and clones, which accumulate cost and clutter.
  • Keep dev and prod structure aligned: same tools, same code, same versions.
  • Control costs: dev warehouses can be small and auto-suspend (warehouse cost management).
  • Don’t let consumers depend on dev data.

Environments are part of DataOps and make safe, frequent change possible.