Contents

AI & Data › LLM & AI Engineering

Chunking

Splitting documents into pieces for embedding and retrieval.

Also known as: chunking, text chunking, document splitting

Chunking splits source documents into smaller passages before they are embedded and indexed for retrieval. A good chunk is a self-contained unit of meaning, sized so it can be matched to a question and still fit comfortably in a model’s input alongside other material.

document → split at structure (headings, paragraphs) → chunks with overlap and metadata → index

Chunking decides what a retriever can find. Chunks that are too large blur several topics into one vector; chunks that are too small lose the context needed to answer. The right size depends on how users phrase questions and how documents are organised.

The classic mistakes:

  • Fixed character windows. Splitting every N characters cuts sentences and tables mid-thought. Split on structural boundaries where possible.
  • No overlap or context. A key fact at a chunk boundary can be separated from the term it refers to. Add modest overlap or carry the section heading into each chunk.
  • Losing metadata. Without the source title, section and date, retrieved passages cannot be cited or filtered. Attach metadata to every chunk.
  • Ignoring tables and code. Naive splitting destroys table structure and code blocks. Treat them as units.
  • Never evaluating chunk size. Test retrieval with a few chunk sizes on real questions before settling.

Practice: split on structure, keep metadata, add small overlap, and choose the size by measuring retrieval quality.

Look at a sample of chunks by hand before indexing everything. Broken tables, orphaned headings and chunks that start mid-sentence are obvious to a reader and easy to miss in aggregate metrics, but they quietly degrade every answer built on them.