Chunking is the step in a retrieval-augmented generation (RAG) pipeline where a document gets split into smaller pieces before each piece is turned into a vector embedding and stored in an index. It sounds like a minor implementation detail, but it is one of the highest-leverage decisions in a RAG system: get chunk size and overlap wrong, and even a well-chosen embedding model and a well-tuned language model will keep giving wrong or incomplete answers from documents that clearly contain the right information.

Why chunking matters this much

RAG works by embedding a user's question into a vector, searching an index of pre-embedded document chunks for the ones whose vectors are closest, and giving those chunks to a language model as context for its answer. Every step downstream of chunking depends on the chunks themselves being reasonably self-contained and topically coherent:

This is one of the most common root causes behind a RAG system that seems to have the right documents indexed but keeps giving vague, wrong, or incomplete answers; see our related piece on why RAG gives wrong answers from the right documents for the broader diagnosis.

Two chunking strategies, compared

Naive fixed-size chunking

The simplest approach: split the text every N characters or every N words, regardless of where sentences, paragraphs or headings fall. It is fast, predictable, and trivial to implement, which is why it is often the first approach anyone tries. Its weakness is exactly its simplicity: a fixed-size split has no idea where a paragraph ends, so it will sometimes cut a sentence in half, separate a heading from its content, or split a table or a list mid-way, producing chunks that read as incoherent fragments even though the split logic itself worked correctly.

Structure-aware chunking

A more deliberate approach: split on natural document boundaries, typically paragraph breaks, and pack whole paragraphs into each chunk up to a target size, only falling back to a fixed-size split for the rare paragraph that is larger than the target size on its own. This tends to produce chunks that read coherently in isolation, which helps both the embedding step (a chunk about one coherent idea embeds more precisely) and the final answer the language model writes from that chunk. The tradeoff is a small amount of extra complexity and slightly less predictable chunk sizes, since a structure-aware splitter respects paragraph boundaries rather than hitting an exact character count every time.

More sophisticated variants exist beyond these two, such as splitting on semantic similarity between sentences or using a document's actual heading hierarchy (H1, H2, H3) to chunk section by section, but fixed-size and paragraph-boundary-aware chunking are the two most common starting points and the ones worth understanding first.

Choosing a chunk size

There is no single correct chunk size; it depends on the kind of document and the kind of question being asked. A few practical starting points:

These are starting points, not fixed rules. The only reliable way to pick a chunk size is to test retrieval against real questions your system actually needs to answer and adjust from there.

What overlap actually buys you
Overlap does not fix a fundamentally bad chunk boundary; it just softens the damage. If a sentence spans a boundary, a 10 to 20 percent overlap increases the odds that sentence appears whole in at least one of the two adjacent chunks, rather than being cut in half in both. It is a mitigation, not a replacement for choosing sensible boundaries in the first place.

Seeing it for yourself

Our free RAG chunking visualizer lets you paste a document, pick a chunk size and overlap in either characters or words, and see exactly where the boundaries fall, with alternating backgrounds marking each chunk and a highlighted strip showing the overlap with the previous one. It supports both strategies described above side by side, along with basic stats: number of chunks, average size, and the smallest and largest chunk, so you can compare before committing to an approach. No embedding or real retrieval happens in the tool; it is purely a visualization of the splitting step, the part of a RAG pipeline that is hardest to reason about without actually seeing it.