Data Engineering › Batch & Distributed Processing
Single-Node Engines (DuckDB, Polars)
Fast local processing that often makes a cluster unnecessary.
Also known as: DuckDB, Polars, single-machine data processing, local analytics engines, in-process analytics
Single-node engines process data on one machine, but do it fast enough that many jobs people assume need a cluster don’t. Two popular examples:
- DuckDB: an in-process analytical SQL database. You query files directly. No server to run.
- Polars: a DataFrame library written in Rust, with a lazy, parallel query engine.
import duckdb
duckdb.sql("""
SELECT country, SUM(amount) AS revenue
FROM 'orders/*.parquet'
GROUP BY country
ORDER BY revenue DESC
""").show()
import polars as pl
(pl.scan_parquet("orders/*.parquet")
.filter(pl.col("status") == "paid")
.group_by("country").agg(pl.col("amount").sum())
.collect())
Why they’re fast
- Columnar and vectorized: they read only the needed columns and process data in batches, using all CPU cores (row vs columnar).
- They query files in place (Parquet, CSV, JSON) without a load step.
- Out-of-core execution: many can spill to disk when data exceeds memory (spill to disk), so “bigger than RAM” doesn’t automatically mean “needs Spark”.
- No cluster overhead: no scheduling, network shuffles or serialization between machines.
Why this matters
Modern machines have many cores, lots of memory and fast disks. A great deal of analytical data, often hundreds of gigabytes, can be processed on a single powerful machine, which is far simpler to run, debug and pay for than a cluster. A cluster brings operational complexity: configuration, failures, data movement and cost (Apache Spark).
When to reach for a cluster instead
- Data or processing truly exceeds a single machine’s memory, disk or time budget.
- You need to scale out continuously as data grows, or share one engine across many concurrent users and jobs (distributed query engines).
- Your platform already standardizes on one.
Practical advice
- Start on one machine. Measure. Add distribution when you have evidence you need it.
- Compare options on your data and queries. Benchmarks vary widely with workload.
- They’re also excellent for local development, testing pipelines and ad-hoc analysis (DataFrames).
- They’re not a replacement for a shared warehouse with concurrency, access control and governance. They complement it.