Contents

Architecture & System Design › Events & Integration · also in Batch & Distributed Processing

Batch vs Stream Processing

Processing data in periodic chunks vs continuously.

Also known as: batch vs streaming, batch processing, stream processing vs batch, real-time vs batch

Two ways to process data:

  • Batch: collect data over a period, then process it all together on a schedule (hourly, nightly). The input is a bounded chunk, with a start and an end.
  • Stream: process each event as it arrives, continuously. The input is unbounded, with no end.
BatchStream
LatencyMinutes to hoursSeconds or less
InputA finite file, table or partitionA never-ending flow of events
Typical useNightly reports, billing runs, model trainingFraud alerts, live dashboards, notifications
ComplexitySimplerHarder: ordering, late data, state, recovery
Failure handlingRerun the jobNeeds checkpoints and careful replay
Cost patternCompute only while runningAlways running
Batch:   [ file/day ] ──► job runs at 02:00 ──► report
Stream:  event ► event ► event ► ... ──► processor ──► results continuously

Why batch is still everywhere

Batch is simpler to build, test, debug and rerun: a failed run is just run again on the same input. Many questions don’t need fresh-to-the-second answers. A daily revenue report is fine as a batch.

When streaming is worth it

When the value of the result drops quickly with time: blocking fraudulent payments, alerting on outages, updating a live leaderboard, reacting to a user’s action while they’re still on the page.

Things streaming makes you handle

  • Event time vs processing time: the order events arrive in isn’t the order they happened in.
  • Late and out-of-order data, so you use windows and decide how long to wait.
  • State (counts, sessions) that must survive restarts.
  • Delivery guarantees and duplicates (delivery guarantees).

In-between

Micro-batch processes small batches every few seconds, giving near-real-time results using batch ideas. Some systems unify the two, treating a batch as a bounded stream. A common architecture is a stream for fresh, approximate results plus batch to correct and reprocess history.

Start with the latency you actually need. Don’t build a streaming system for a report read once a day. See stream processing and ETL.