Chunking is the step in a retrieval-augmented generation (RAG) pipeline where a document gets split into smaller pieces before each piece is turned into a vector embedding and stored in an index. It sounds like a minor implementation detail, but it is one of the highest-leverage decisions in a RAG system: get chunk size and overlap wrong, and even a well-chosen embedding model and a well-tuned language model will keep giving wrong or incomplete answers from documents that clearly contain the right information.
Why chunking matters this much
RAG works by embedding a user's question into a vector, searching an index of pre-embedded document chunks for the ones whose vectors are closest, and giving those chunks to a language model as context for its answer. Every step downstream of chunking depends on the chunks themselves being reasonably self-contained and topically coherent:
- Chunks that are too large mix several topics or claims into one embedding. A search for one specific fact returns a chunk whose overall vector is a blend of that fact and several unrelated ones, diluting how well it matches the query and making it more likely a more relevant, smaller chunk elsewhere in the index gets missed or outranked.
- Chunks that are too small can match a query on surface similarity, a shared keyword or phrase, while lacking the surrounding context actually needed to answer the question. A chunk containing only the sentence "the deadline was extended" is a poor answer without the sentence before it saying which deadline.
- Chunks that ignore document structure can split a heading away from the paragraph it introduces, or cut a numbered list in half, producing retrieved context that reads as a confusing fragment even when it is topically the right passage.
This is one of the most common root causes behind a RAG system that seems to have the right documents indexed but keeps giving vague, wrong, or incomplete answers; see our related piece on why RAG gives wrong answers from the right documents for the broader diagnosis.
Two chunking strategies, compared
Naive fixed-size chunking
The simplest approach: split the text every N characters or every N words, regardless of where sentences, paragraphs or headings fall. It is fast, predictable, and trivial to implement, which is why it is often the first approach anyone tries. Its weakness is exactly its simplicity: a fixed-size split has no idea where a paragraph ends, so it will sometimes cut a sentence in half, separate a heading from its content, or split a table or a list mid-way, producing chunks that read as incoherent fragments even though the split logic itself worked correctly.
Structure-aware chunking
A more deliberate approach: split on natural document boundaries, typically paragraph breaks, and pack whole paragraphs into each chunk up to a target size, only falling back to a fixed-size split for the rare paragraph that is larger than the target size on its own. This tends to produce chunks that read coherently in isolation, which helps both the embedding step (a chunk about one coherent idea embeds more precisely) and the final answer the language model writes from that chunk. The tradeoff is a small amount of extra complexity and slightly less predictable chunk sizes, since a structure-aware splitter respects paragraph boundaries rather than hitting an exact character count every time.
More sophisticated variants exist beyond these two, such as splitting on semantic similarity between sentences or using a document's actual heading hierarchy (H1, H2, H3) to chunk section by section, but fixed-size and paragraph-boundary-aware chunking are the two most common starting points and the ones worth understanding first.
Choosing a chunk size
There is no single correct chunk size; it depends on the kind of document and the kind of question being asked. A few practical starting points:
- Shorter chunks (a few hundred words) tend to work well for dense reference material, like API documentation or policy text, where a question usually has one precise answer in one place.
- Longer chunks (closer to a thousand words) can work better for narrative or explanatory content, where the answer to a question depends on surrounding context that a very short chunk would cut away.
- Overlap of roughly 10 to 20 percent of the chunk size is a common starting point, enough to preserve context across a boundary without duplicating a large fraction of the index.
These are starting points, not fixed rules. The only reliable way to pick a chunk size is to test retrieval against real questions your system actually needs to answer and adjust from there.
Seeing it for yourself
Our free RAG chunking visualizer lets you paste a document, pick a chunk size and overlap in either characters or words, and see exactly where the boundaries fall, with alternating backgrounds marking each chunk and a highlighted strip showing the overlap with the previous one. It supports both strategies described above side by side, along with basic stats: number of chunks, average size, and the smallest and largest chunk, so you can compare before committing to an approach. No embedding or real retrieval happens in the tool; it is purely a visualization of the splitting step, the part of a RAG pipeline that is hardest to reason about without actually seeing it.