Contents

Backend Development › Database Internals

Checkpoint

Flushing in-memory changes to disk so recovery doesn't replay the whole log.

Also known as: checkpoint, database checkpoint, checkpointing

A checkpoint is the process of flushing the database’s in-memory (dirty) data pages to their data files on disk, so that the write-ahead log up to that point is no longer needed for recovery and can be truncated. It periodically reconciles the fast, append-only WAL with the actual data files.

writes → WAL (durable, append) + dirty pages in buffer pool
checkpoint → write dirty pages to data files → WAL before this point can be discarded

Why it matters: without checkpoints, crash recovery would have to replay the entire WAL from the beginning, taking longer and longer as the database runs. After a checkpoint, recovery only needs to replay from the last checkpoint forward — bounded recovery time. Checkpoints also let the WAL (and its disk usage) be recycled.

The classic mistakes:

  • No checkpoints → endless WAL growth and recovery. Without checkpointing, the log grows without bound and recovery replays everything. The database effectively never forgets. Checkpoints bound both.
  • Checkpoints too aggressive. Frequent checkpoints cause bursts of I/O as many dirty pages flush at once, harming performance. They should be spread out gradually.
  • Checkpoints too infrequent. Rare checkpoints mean a long recovery time and a large WAL, and more data to redo after a crash. Balance I/O impact against recovery time.
  • Confusing checkpoints with commits. A commit means the WAL is durable; it doesn’t mean data files are updated. Checkpoints handle the latter. Durability lives in the WAL, not the data files.
  • Forgetting the impact of write spikes. A burst of writes creates many dirty pages, and a checkpoint must flush them; on slow storage this can stall.
  • Ignoring checkpoint behavior on replicas. Standbys redo the primary’s WAL; checkpoint timing affects their lag and recovery too.

How to think about it: the WAL makes commits durable; checkpoints make the data files catch up so recovery stays bounded and disk doesn’t fill with log. Tuning the frequency is a trade between I/O spikes and recovery time. It’s a core part of the storage engine’s durability machinery, alongside the buffer pool it flushes (see write-ahead log).