Skip to content

Chunking choices that move recall

On a telecom help center we evaluated, moving from 512-token fixed windows to 320-token semantic chunks cut chunk-boundary loss from 7.2% to 3.1%. Here is what changed.

Relevant teamApplied research

7 min read

On this page
  1. Measure chunk-boundary loss
  2. Fixed windows versus semantic chunks
  3. Watch the cost side

Chunking is the least glamorous decision in a RAG pipeline and one of the most consequential. Policy documents are full of exceptions that sit one paragraph away from the rule they modify. Split them apart and the generator sees half the policy.

Measure chunk-boundary loss

Chunk-boundary loss is the share of gold answers whose supporting sentence spans two chunks. It is easy to compute once gold passages are mapped to character offsets, and it predicts a specific kind of failure: answers that state a rule and miss its exception.

Fixed windows versus semantic chunks

Fixed 512-token windows with 64 tokens of overlap are a reasonable default. Semantic chunking splits on headings and paragraph boundaries and caps chunks at a smaller size, around 320 tokens.

On a telecom help center, semantic chunks reduced boundary loss from 7.2% to 3.1% and raised context precision, because each chunk carries one idea with its heading.

  • Keep the section heading in every chunk's text.
  • Never split a table row from its header.
  • Store the article version and effective date as metadata.

Watch the cost side

Smaller chunks mean more of them in context for the same recall. Measure input tokens per request alongside recall; on that help center input tokens fell because fewer irrelevant paragraphs rode along with each relevant one.

Share
XLinkedIn
All articles
    • Evaluation
    • Retrieval

    Why aggregate scores hide retrieval failures

    A single accuracy number can rise while the answers your customers care about get worse. Here is how to break the number apart.

    9 min read

    • Model comparison
    • Evaluation

    Choosing a generator with evidence

    Comparing models on the same retrieval, the same prompt and the same questions, then reading the per-category differences.

    7 min read

    • Failure analysis
    • Process

    A failure taxonomy for support assistants

    Six categories, clear detection signals and one owner per category: a failure taxonomy your support team will actually use.

    8 min read

See why your retrieval fails

Run an evaluation on your own dataset and read every failure with its retrieved chunks, flagged claims and suggested fix.

Free plan includes 2,000 evaluated samples a month