Contents

Data Engineering › Ingestion

Batch vs Streaming Ingestion

Loading data on a schedule vs continuously as it's produced.

Also known as: batch ingestion, streaming ingestion, micro-batch ingestion, real-time ingestion

Ingestion can bring data in on a schedule (batch) or continuously as it’s produced (streaming). The choice sets freshness, cost and complexity.

Batch ingestionStreaming ingestion
HowRun a job every N minutes or hours; load what’s newConsume events continuously from a stream or change log
FreshnessMinutes to a daySeconds
ComplexityLower: simple jobs, easy rerunsHigher: ordering, duplicates, late data, state
CostPay when runningAlways-on infrastructure
Typical sourcesDatabase extracts, file drops, API pullsEvent streams, message queues, change data capture
Batch:     source ─(every hour: SELECT ... WHERE updated_at > :last_run)─► landing zone
Streaming: source ─(each change/event)─► stream ─► consumer ─► storage within seconds

Batch in practice

Typically: remember where you stopped (a high-water mark), pull everything newer, write it, and advance the mark. Or load files that landed since the last run (full vs incremental load). Failures are easy to handle: rerun the batch. It’s the right default.

Streaming in practice

Changes arrive as a flow: from an event bus, or from a database’s change log (log-based CDC), which captures every insert, update and delete without hammering the source with queries. You have to handle:

  • Duplicates: at-least-once delivery is common, so deduplicate (deduplication).
  • Ordering: events can arrive out of order.
  • Back-pressure and recovery: replaying from a saved position after a failure (messages vs streams).

Choosing

Ask: how fresh does the consumer really need it? “The dashboard is read each morning” means batch is enough. “Fraud must be flagged before shipping” means stream. Streaming costs more to build and operate, so adopt it for use cases that need it.

There’s a middle ground, micro-batching: tiny batches every few seconds or minutes, which is often good enough for “near real time”. Many pipelines also mix the two: stream in for freshness, batch to reconcile and reprocess. See batch vs stream processing for the processing side of the same trade-off.