Paste a document, pick a chunk size and overlap, and see exactly where the boundaries fall. Compares naive fixed-size chunking against a structure-aware option that respects paragraph breaks.
Chunking is splitting a document into smaller pieces before turning each piece into a vector embedding and storing it in an index. Retrieval-augmented generation searches those chunks for the ones most relevant to a question, then gives them to a language model as context.
Fixed-size (naive) chunking splits purely by character or word count, ignoring where sentences or paragraphs fall, so it can cut a paragraph or a heading in half. Structure-aware chunking instead tries to keep whole paragraphs together, splitting on blank lines, and only breaks a single paragraph mid-way if that one paragraph is larger than the target size.
Without overlap, a sentence or idea that spans a chunk boundary can be cut in half, losing meaning in both halves. A small overlap repeats the tail of one chunk at the start of the next, so context near a boundary is not lost entirely.
No. It only visualizes the splitting step: where the boundaries would fall and how big each resulting chunk is. No embedding model runs and no vector index is created.