Contents

Data Engineering

Batch & Distributed Processing

Processing large datasets across many machines.

Backend Engineer track

Junior

Write correct code, ship small changes safely, ask good questions.

Nothing here yet.

Mid-level

Own a feature end to end without hand-holding.

Nothing here yet.

Senior

Own a system, its failure modes, and its trade-offs.

Staff

Shape how many teams build, across systems.

Nothing here yet.

Principal

Set technical direction for the organization.

Nothing here yet.

Data Analyst track

Junior

Write correct SQL, build trusted dashboards, ask good questions.

Mid-level

Own an analysis end to end, from vague question to recommendation.

Senior

Own experimentation and metrics design; call out bad numbers.

  • Apache SparkThe most widely used engine for distributed batch and streaming processing.

Staff

Shape how the organization measures and decides.

Nothing here yet.

Principal

Set measurement strategy across the company.

Nothing here yet.

Data Engineer track

Junior

Build and fix pipelines from clear specs; write correct SQL.

Core: start here

2 more junior concepts
  • Notebooks (Jupyter)Interactive documents mixing code, output and notes.
  • pandasPython's standard DataFrame library for data analysis.

Mid-level

Own pipelines and models end to end, including their quality.

Core: start here

7 more mid-level concepts

Senior

Design the platform's storage, processing and modeling choices.

Core: start here

  • Broadcast JoinSending a small table to every worker to avoid shuffling the big one.
  • Data SkewA few keys holding most of the data, so one task runs forever.
3 more senior concepts

Staff

Shape how the whole organization produces and uses data.

Nothing here yet.

Principal

Set data strategy and architecture across the company.

Nothing here yet.