Chunking is the least glamorous decision in a RAG pipeline and one of the most consequential. Policy documents are full of exceptions that sit one paragraph away from the rule they modify. Split them apart and the generator sees half the policy.
Measure chunk-boundary loss
Chunk-boundary loss is the share of gold answers whose supporting sentence spans two chunks. It is easy to compute once gold passages are mapped to character offsets, and it predicts a specific kind of failure: answers that state a rule and miss its exception.
Fixed windows versus semantic chunks
Fixed 512-token windows with 64 tokens of overlap are a reasonable default. Semantic chunking splits on headings and paragraph boundaries and caps chunks at a smaller size, around 320 tokens.
On a telecom help center, semantic chunks reduced boundary loss from 7.2% to 3.1% and raised context precision, because each chunk carries one idea with its heading.
- Keep the section heading in every chunk's text.
- Never split a table row from its header.
- Store the article version and effective date as metadata.
Watch the cost side
Smaller chunks mean more of them in context for the same recall. Measure input tokens per request alongside recall; on that help center input tokens fell because fewer irrelevant paragraphs rode along with each relevant one.