Storage, Formats & Lakehouse
Warehouses, lakes, file formats and table formats.
Backend Engineer track
Junior
Write correct code, ship small changes safely, ask good questions.
- CSVThe simplest tabular format, and its quoting, encoding and type pitfalls.
Mid-level
Own a feature end to end without hand-holding.
- Data WarehouseA database optimized for analytics, like BigQuery or Snowflake.
Senior
Own a system, its failure modes, and its trade-offs.
- AvroA binary serialization format whose schemas are designed to evolve.
- Columnar StorageStoring data by column for fast analytics.
- Data LakeCheap storage for raw data in any format.
- ParquetA columnar file format for analytics.
- PartitioningDividing data into parts, within one machine or across many.
- Table PartitioningSplitting one huge table into smaller physical pieces, e.g. by month.
Staff
Shape how many teams build, across systems.
Nothing here yet.
Principal
Set technical direction for the organization.
Nothing here yet.
Data Analyst track
Junior
Write correct SQL, build trusted dashboards, ask good questions.
- Data LakeCheap storage for raw data in any format.
- Data WarehouseA database optimized for analytics, like BigQuery or Snowflake.
Mid-level
Own an analysis end to end, from vague question to recommendation.
Nothing here yet.
Senior
Own experimentation and metrics design; call out bad numbers.
Nothing here yet.
Staff
Shape how the organization measures and decides.
Nothing here yet.
Principal
Set measurement strategy across the company.
Nothing here yet.
Data Engineer track
Junior
Build and fix pipelines from clear specs; write correct SQL.
Core: start here
- Columnar StorageStoring data by column for fast analytics.
- CSVThe simplest tabular format, and its quoting, encoding and type pitfalls.
- Data LakeCheap storage for raw data in any format.
- Data WarehouseA database optimized for analytics, like BigQuery or Snowflake.
- Medallion Architecture (Bronze, Silver, Gold)Layering data from raw to cleaned to business-ready.
- ParquetA columnar file format for analytics.
- Partitioned Tables (Hive-Style)Organizing files into folders like date=2024-06-01 so queries can skip data.
- PartitioningDividing data into parts, within one machine or across many.
- Row vs Columnar File FormatsCSV and Avro store rows together; Parquet and ORC store columns together.
1 more junior concepts
- JSON LinesOne JSON object per line, easy to stream and append.
Mid-level
Own pipelines and models end to end, including their quality.
Core: start here
- CompactionMerging small files into fewer, larger ones.
- Compression Codecs (Snappy, Zstd, Gzip)Trading CPU for smaller files, and which codecs are splittable.
- LakehouseWarehouse-style tables and transactions on top of cheap lake storage.
- Open Table Formats (Iceberg, Delta, Hudi)Metadata layers that give files in a lake ACID transactions, schemas and time travel.
- Partition PruningThe engine skipping partitions a query doesn't need.
- Predicate PushdownFiltering inside the storage layer before data is read.
- Separation of Storage and ComputeScaling query engines independently from where data lives.
- Small Files ProblemToo many tiny files slowing queries and overloading metadata.
- Table Catalog / MetastoreThe registry of table names, schemas and file locations, like Hive Metastore or Glue.
7 more mid-level concepts
- AvroA binary serialization format whose schemas are designed to evolve.
- Data MartA subset of the warehouse focused on one team or subject.
- Data Retention and TieringExpiring or moving old data to cheaper storage on purpose.
- Distributed File System (HDFS)Storing huge files across many machines with replication.
- MPP Data WarehouseWarehouses that split one query across many nodes in parallel.
- Table PartitioningSplitting one huge table into smaller physical pieces, e.g. by month.
- Time TravelQuerying a table as it was at an earlier point in time.
Senior
Design the platform's storage, processing and modeling choices.
Core: start here
- Clustering and Z-OrderingSorting data inside files so related rows sit together for faster scans.
5 more senior concepts
- Apache ArrowAn in-memory columnar format for moving data between tools without conversion.
- BucketingHashing rows into a fixed number of files by key to speed joins.
- Data SharingGiving another team or company live access to data without copying it.
- ORCA columnar format common in the Hadoop ecosystem.
- Zero-Copy CloneCopying a table instantly by sharing its underlying files.
Staff
Shape how the whole organization produces and uses data.
Nothing here yet.
Principal
Set data strategy and architecture across the company.
Nothing here yet.