Contents

Data Engineering › Data Engineering Foundations

Source Systems

Where data originates: application databases, APIs, logs, files, SaaS tools and devices.

Also known as: data sources, upstream systems, operational systems, source system

Source systems are where data originates: the systems that generate data as a side effect of doing something else. The data engineer usually doesn’t own them, but everything downstream depends on them.

Common types

SourceExamplesTypical access
Application databasesPostgres, MySQL behind the productQuery a replica, or read the change log (log-based CDC)
APIs / SaaS toolsPayments, CRM, support, ad platformsPull via REST APIs (API ingestion), often rate-limited
FilesCSV, JSON and Parquet drops from partners or exportsPick up from storage (file ingestion)
Event streams and logsClickstream, application logs, message queuesConsume continuously
Devices and sensors (IoT)Machines, vehicles, wearablesStreaming, often noisy and intermittent (IoT data)
Third-party dataPurchased or public datasetsFiles or APIs (external data)

What to find out about each one

  • Who owns it, and how do you reach them when it changes?
  • What does it contain: schema, keys, meaning of columns, which fields are reliable?
  • How does it change: does a row get updated in place (you lose history), or only appended?
  • How often does it change, and how fresh do you need the copy to be?
  • How can you read it safely: load limits, rate limits, a read replica, a maintenance window?
  • Is it reliable: duplicates, late or missing data, outages, clock problems?
  • What’s sensitive in it (personal or regulated data), and what are you allowed to copy?

Pitfalls

  • Querying the production database directly for big extracts can slow down the application. Use a replica, an export or change data capture.
  • Schema changes upstream can silently break your pipeline (schema drift). Agree on notifications or contracts with the owning team.
  • Source data quality is the ceiling for everything downstream. Fix problems at the source when you can.
  • Business logic lives in the source. What does “active customer” mean in the app? Understand it before reinterpreting it.
  • Operational databases are built for fast transactions, not analysis, which is why data gets copied elsewhere.

Treat the relationship with source system owners as a core part of the job. Many pipeline failures are communication failures.