Contents

Data Engineering › Storage, Formats & Lakehouse

ORC

A columnar format common in the Hadoop ecosystem.

Also known as: Optimized Row Columnar, ORC file format, ORC format

ORC (Optimized Row Columnar) is a columnar file format used mostly in the Hadoop ecosystem. It stores each column separately, so a query that reads two of twenty columns only reads those two, and it keeps lightweight indexes (such as min/max per group of rows) so an engine can skip data that cannot match a filter.

ORC came out of Hive and is well supported by Hive, Spark, and other JVM-based engines. Parquet is the other common columnar format, and the two are often compared.

The classic mistake is picking a format by habit and then hitting an ecosystem wall. If your processing tools only read Parquet, an ORC dataset forces an extra conversion step; if your cluster is Hive-centric and already writes ORC, switching for no reason adds work. Both formats give the main columnar benefits — predicate pushdown, compression, and reading only the needed columns — so the choice is usually about ecosystem fit, not raw speed.

ORC and Parquet

  • ORC originated in Hive; Parquet came from the wider Hadoop/analytics community. Parquet has broader language support outside the JVM.
  • Both support schema evolution, nested types and column-level compression; the specifics and quality of support vary by reader.
  • ORC has tighter integration with Hive’s transactional tables in some versions. That is an engine-specific feature, not a general property of the format.
  • Do not assume one is always faster. Benchmarks flip depending on data, engine and version.

Trade-offs

Columnar formats are built for scans and analytics, not for row-level updates or tiny files. A folder of thousands of small ORC files hurts query planning and metadata; use compaction and partitioning. For durable analytics storage in a data lake, ORC or Parquet is a good default; for interchange between arbitrary tools, Parquet is the safer bet.