Data Engineering › Storage, Formats & Lakehouse
Lakehouse
Warehouse-style tables and transactions on top of cheap lake storage.
Also known as: data lakehouse, lakehouse architecture, lake house
A lakehouse combines the two older ideas: the cheap, open, flexible storage of a data lake with the reliability and performance features of a data warehouse: transactions, schemas, fast SQL. The term was popularized by Databricks, and the approach is built into several platforms.
files in object storage (Parquet)
▲
open table format (metadata, transactions, schema, time travel)
▲
query engines: SQL, Spark, ML, BI tools, all reading the same tables
The problem it addresses
Teams often ran both: a lake for raw data and ML, and a warehouse for BI. That meant two copies of data, pipelines to move between them, two sets of access rules and arguments about which numbers were right. A lakehouse aims to keep one copy that serves both.
What makes it possible
- Open file formats (mostly Parquet) on cheap object storage.
- An open table format (Apache Iceberg, Delta Lake, Apache Hudi) on top of those files, giving ACID transactions, schema enforcement and evolution, time travel and row-level updates and deletes.
- A catalog that registers tables and their locations (catalog and metastore).
- Separation of storage and compute: engines scale independently from the data (storage and compute separation).
- Fast query engines with the usual optimizations: columnar reads, pruning, caching (partition pruning, predicate pushdown).
Data is commonly organized in layers from raw to curated (medallion architecture).
Benefits
- Less duplication and fewer movements of data.
- Many tools (SQL, Spark, Python, ML) can work on the same tables.
- Openness: your data isn’t locked inside one vendor’s proprietary storage.
- Handles structured and semi-structured data, and supports streaming and batch.
Things to be realistic about
- It’s an architecture and a marketing term, and different vendors implement it differently, with varying maturity. Check specifics for the features you need (concurrent writes, governance, performance).
- You take on table maintenance (compaction, snapshot cleanup) (compaction, small files).
- A good warehouse may still outperform for pure BI workloads. The lakehouse’s strength is flexibility.
- Governance and quality don’t come automatically. You still need them (data governance).
Choose based on your workloads, your team’s skills and your tolerance for operational work.