Contents

Backend Development › Database Internals

Write Amplification

One logical write turning into many physical ones.

Also known as: write amplification, write amplification factor, WA

Write amplification is the ratio of bytes physically written to storage for each logical write the application performs. A single UPDATE can cause many physical writes — the WAL record, the data page, index updates, and (in an LSM tree) repeated compaction rewrites — so writing one logical byte may cost several physical ones.

logical write: 1 row updated
physical writes: WAL + data page + index pages + (LSM) compaction rewrites
write amplification = physical bytes / logical bytes

Why it matters: storage has finite write throughput and, on SSDs, finite write endurance (cells wear out after so many writes). High write amplification means slower writes, more I/O contention, and shorter SSD life. It’s a central performance characteristic of storage engines.

The main sources:

  • WAL — every change is written to the log before the data (see write-ahead log).
  • Page rewrites — a B-tree update rewrites a whole page, even for a small change.
  • Compaction — an LSM tree rewrites data multiple times as it merges levels (see SSTables and compaction).
  • Index maintenance — updating every index a changed column belongs to.
  • Replication — writing to replicas multiplies physical writes.

The classic mistakes:

  • Assuming a write costs one write. The WAL, indexes, page rewrites and compaction all multiply it. Understanding this explains why bulk writes and many small updates feel different.
  • Ignoring compaction’s amplification. LSM engines can amplify writes several-fold through compaction; this wears SSDs and consumes bandwidth, especially under heavy write load (see SSTables and compaction).
  • Too many indexes. Each index adds write cost on every affected change; a table with many indexes pays for them on every write. Index deliberately (see indexes).
  • Small random writes instead of batching. Many tiny writes each pay the WAL and page costs; batching amortises them (a reason bulk loads are efficient).
  • Forgetting the SSD endurance angle. High write amplification shortens flash life; for write-heavy workloads, engine and index choices affect hardware longevity.
  • Only optimising reads. A system tuned purely for read speed (many indexes, read-optimised structures) can have terrible write amplification.

How to manage it: understand where your writes multiply (WAL, indexes, compaction), batch where possible, avoid unnecessary indexes, and choose an engine matching your read/write mix. It’s the hidden cost behind storage-engine design — the reason “just write the data” is never quite one write. See storage engine and B-tree vs LSM.