GreyScript AI

RAG Chunking Visualizer

Paste a document, pick a chunk size and overlap, and see exactly where the boundaries fall. Compares naive fixed-size chunking against a structure-aware option that respects paragraph breaks.

Your document
Chunking strategy
How much of the previous chunk repeats at the start of the next one.
Stats
Chunks
Alternating backgrounds mark chunk boundaries. A darker leading strip on a chunk shows the part that overlaps with the previous chunk.
Everything is split in your browser. No embedding, indexing or real retrieval happens here; this only visualizes where the splitting step would place chunk boundaries. Nothing you paste is sent anywhere or saved.

Questions and answers

What is chunking in a RAG pipeline?

Chunking is splitting a document into smaller pieces before turning each piece into a vector embedding and storing it in an index. Retrieval-augmented generation searches those chunks for the ones most relevant to a question, then gives them to a language model as context.

What is the difference between fixed-size and structure-aware chunking?

Fixed-size (naive) chunking splits purely by character or word count, ignoring where sentences or paragraphs fall, so it can cut a paragraph or a heading in half. Structure-aware chunking instead tries to keep whole paragraphs together, splitting on blank lines, and only breaks a single paragraph mid-way if that one paragraph is larger than the target size.

Why does overlap between chunks matter?

Without overlap, a sentence or idea that spans a chunk boundary can be cut in half, losing meaning in both halves. A small overlap repeats the tail of one chunk at the start of the next, so context near a boundary is not lost entirely.

Does this tool actually create embeddings or search anything?

No. It only visualizes the splitting step: where the boundaries would fall and how big each resulting chunk is. No embedding model runs and no vector index is created.