Data Engineering › Storage, Formats & Lakehouse
Small Files Problem
Too many tiny files slowing queries and overloading metadata.
Also known as: small file problem, tiny files, too many small files, many small files
The small files problem is when a table or dataset consists of huge numbers of tiny files (kilobytes to a few megabytes each), which makes queries slow and puts heavy load on storage metadata. It’s one of the most common performance issues in data lakes.
Why it happens
- Streaming or micro-batch writes that create a new file every few seconds or minutes.
- Over-partitioning: partitioning by a high-cardinality column (such as user ID) or by hour or minute on small data, so each partition holds only a bit of data.
- Many parallel tasks, each writing its own small output file (for example, 2,000 Spark tasks writing 2,000 files for a few megabytes of data).
- Frequent small incremental loads appending a file each time.
- Data sources that deliver lots of small files (logs, IoT).
Why it hurts
- Per-file overhead dominates: listing, opening, reading metadata and scheduling a task for each file. Reading 100,000 tiny files can take far longer than reading 100 large ones holding the same data.
- Metadata pressure: object storage listings are slow, and catalogs and the engine’s planner struggle with huge file lists.
- Poor compression and column statistics: small files compress less well, and statistics span too little data to skip effectively (predicate pushdown).
- Slow planning before any real work begins.
- Higher cost: more requests to storage, and more tasks.
How to fix and prevent it
- Compact: periodically rewrite small files into larger ones (compaction), and many open table formats have commands for it.
- Write larger files: buffer more before flushing, repartition or coalesce output so each write task produces a reasonably sized file, and tune writer settings.
- Partition less aggressively. Choose a column and granularity that give sizable partitions (Hive partitioning). Aim for files of hundreds of megabytes where the engine prefers it.
- Batch ingestion into fewer, larger writes instead of per-event files.
- Use a table format that handles file sizing and maintenance.
Detecting it
Look at the number of files and their average size per table or partition, and at query planning time. A table with thousands of files per partition averaging a few MB is a warning sign. Check this regularly, especially for streaming outputs.
Rule of thumb: fewer, larger files (but not enormous ones, which limit parallelism).