Batch & Distributed Processing
Processing large datasets across many machines.
Backend Engineer track
Junior
Write correct code, ship small changes safely, ask good questions.
Nothing here yet.
Mid-level
Own a feature end to end without hand-holding.
Nothing here yet.
Senior
Own a system, its failure modes, and its trade-offs.
- Batch vs Stream ProcessingProcessing data in periodic chunks vs continuously.
- Data LocalityKeeping data close to where it's processed.
- Distributed Batch ProcessingEngines like Spark that split big jobs across many machines.
- MapReduceProcessing huge datasets by mapping in parallel and then reducing the results.
Staff
Shape how many teams build, across systems.
Nothing here yet.
Principal
Set technical direction for the organization.
Nothing here yet.
Data Analyst track
Junior
Write correct SQL, build trusted dashboards, ask good questions.
- Notebooks (Jupyter)Interactive documents mixing code, output and notes.
Mid-level
Own an analysis end to end, from vague question to recommendation.
- DataFrameA table-like data structure with named columns, as in pandas, Polars and Spark.
- Distributed SQL Query Engines (Trino, Presto)Querying data where it lives, across lakes and databases.
- pandasPython's standard DataFrame library for data analysis.
- Single-Node Engines (DuckDB, Polars)Fast local processing that often makes a cluster unnecessary.
Senior
Own experimentation and metrics design; call out bad numbers.
- Apache SparkThe most widely used engine for distributed batch and streaming processing.
Staff
Shape how the organization measures and decides.
Nothing here yet.
Principal
Set measurement strategy across the company.
Nothing here yet.
Data Engineer track
Junior
Build and fix pipelines from clear specs; write correct SQL.
Core: start here
- Batch ProcessingProcessing a bounded chunk of data in one run.
- Batch vs Stream ProcessingProcessing data in periodic chunks vs continuously.
- DataFrameA table-like data structure with named columns, as in pandas, Polars and Spark.
- Single-Node Engines (DuckDB, Polars)Fast local processing that often makes a cluster unnecessary.
2 more junior concepts
- Notebooks (Jupyter)Interactive documents mixing code, output and notes.
- pandasPython's standard DataFrame library for data analysis.
Mid-level
Own pipelines and models end to end, including their quality.
Core: start here
- Apache SparkThe most widely used engine for distributed batch and streaming processing.
- Distributed SQL Query Engines (Trino, Presto)Querying data where it lives, across lakes and databases.
- ShuffleRedistributing data across machines by key, often the slowest step.
- Transformations vs ActionsLazy steps that build a plan vs commands that trigger execution.
7 more mid-level concepts
- Data LocalityKeeping data close to where it's processed.
- Distributed Batch ProcessingEngines like Spark that split big jobs across many machines.
- Distributed ComputingSplitting work across many machines that coordinate over a network.
- Driver and ExecutorsThe process that plans a job and the workers that run its tasks.
- MapReduceProcessing huge datasets by mapping in parallel and then reducing the results.
- Partitions in Distributed ProcessingThe chunks of data that tasks process in parallel.
- User-Defined Function (UDF)Custom code called from SQL or DataFrame operations, and its performance cost.
Senior
Design the platform's storage, processing and modeling choices.
Core: start here
- Broadcast JoinSending a small table to every worker to avoid shuffling the big one.
- Data SkewA few keys holding most of the data, so one task runs forever.
3 more senior concepts
- Cluster Resource Manager (YARN, Kubernetes)Allocating CPU and memory to jobs on a shared cluster.
- Query FederationOne query reading from several different systems.
- Spill to DiskRunning out of memory and writing intermediate data to disk.
Staff
Shape how the whole organization produces and uses data.
Nothing here yet.
Principal
Set data strategy and architecture across the company.
Nothing here yet.