Contents

Data Engineering › Working as a Data Engineer

Tuning a Slow Pipeline

Finding the slow step and fixing it: partitions, joins, skew, file sizes.

Also known as: pipeline tuning, slow pipeline, pipeline optimization, speeding up pipelines

Tuning a slow pipeline means finding which step is slow and fixing that step, instead of guessing or throwing more machines at it. A pipeline is a chain; its runtime is set by the slowest stage, the bottleneck, so work on that.

The classic mistake is optimizing before measuring. People rewrite a transformation that is fast while the real cost is a shuffle in a join, or data skew where one task processes most of the rows. Read the execution plan first: for SQL, look at the query plan (EXPLAIN); for Spark, the stage and task view shows where time and data go. Then fix the actual bottleneck.

Common causes and fixes

  • Reading too much data. Filter and select columns early so the engine reads less. Partitioned tables let it skip whole partitions (partition pruning); columnar formats read only needed columns (row vs columnar).
  • Expensive joins. A large-large join shuffles both sides. If one side is small, a broadcast join avoids the shuffle. Bucketing both sides on the join key can help too.
  • Data skew. If a few keys hold most of the rows, one task runs far longer than the rest. Salting hot keys or handling them separately spreads the load (data skew).
  • Too many small files. Thousands of tiny files add metadata and open overhead; compaction merges them (small files problem).
  • Spilling to disk. When a stage needs more memory than it has, it writes to disk and slows down (spill to disk). Reduce data per task, or adjust partitioning, rather than only adding memory.
  • Recomputing everything. If a job rebuilds a full table daily, an incremental model that processes only new or changed partitions is usually much cheaper.

Keep it in proportion

Tuning has a cost: more complex code, harder maintenance, and sometimes less readable SQL. Optimize the pipelines that actually hurt, usually the ones on the critical path or those that dominate cost. Do not micro-optimize a job that runs in two minutes once a day. And measure again after each change, since moving the bottleneck just reveals the next one.