Contents

Data Engineering › Storage, Formats & Lakehouse

Compression Codecs (Snappy, Zstd, Gzip)

Trading CPU for smaller files, and which codecs are splittable.

Also known as: Snappy, Zstd, gzip, compression codec, LZ4, zstandard, file compression

A compression codec is the algorithm used to shrink data. In data engineering, you pick one for files (Parquet, CSV, JSON lines), messages and backups. The trade-off is CPU time and speed vs storage size.

CodecCompression ratioSpeedNotes
GzipGoodSlowishUbiquitous and widely compatible
SnappyModerateVery fastLong-time default in many Parquet and Hadoop-era tools. Optimized for speed
LZ4ModerateExtremely fastCommon where throughput matters most
Zstd (Zstandard)Very good, tunableFastA strong modern choice: good ratios with high speed, with adjustable levels
Bzip2HighSlowRarely worth it today

Exact numbers depend on your data. Test on your own files rather than trusting general rankings.

Splittability

Distributed engines read different parts of a big file in parallel. This works only if the format can be split:

  • A compressed text file (orders.csv.gz) generally can’t be split: one task must read it from the start, so a huge gzip file becomes a bottleneck. Use many moderately sized files instead.
  • Columnar formats like Parquet and ORC compress inside (per column chunk) and carry their own structure, so the file stays splittable regardless of the codec. This is a major reason to convert to Parquet (Parquet).

How to choose

  • For analytical files (Parquet), Zstd or Snappy are the usual choices: Snappy for the fastest reads and writes, Zstd for smaller files at similar speed. Gzip when you need broad compatibility.
  • For cold storage and archives, favor higher compression levels (smaller files; slower to write).
  • For hot paths (streaming messages, shuffles, temporary data), favor speed (LZ4 or Snappy).
  • For text transfer (CSV or JSON files shared with partners), gzip is the lingua franca.
df.to_parquet("orders.parquet", compression="zstd")

Things to remember

  • Columnar data compresses better than row data, since similar values are stored together (row vs columnar).
  • Compression saves storage and I/O, which often makes queries faster even though decompression costs CPU, since disk and network are slower than CPUs.
  • Already-compressed data (images, video, encrypted data) doesn’t compress further.
  • Both writers and readers need support for the codec you choose, so check your engines’ compatibility.
  • Don’t compress tiny files into smaller tiny files. The real fix is fewer, larger files (file compaction).