Skip to content

Why aggregate scores hide retrieval failures

A single accuracy number can rise while the answers your customers care about get worse. Here is how to break the number apart.

Relevant teamApplied research

9 min read

On this page
  1. One number, four failure modes
  2. Separate retrieval from generation first
  3. Attribute every failure to a chunk
  4. Track the mix, not just the total

Most teams start evaluating a RAG system with one number: the share of answers a judge marks as correct. It is a reasonable start and a poor place to stay. Aggregate scores average over failure modes that have different causes, different owners and different fixes.

One number, four failure modes

When an answer is wrong, the cause usually sits in one of four places: the retriever never surfaced the right passage, it surfaced a stale or wrong passage, the generator ignored or contradicted a good passage, or the generator followed the passage but broke a format or policy rule.

A pass rate that moves from 81% to 84% tells you nothing about which of these moved. It is common to see retrieval improve while generation regresses, with the two effects canceling in the aggregate.

  • Missing context: the gold passage is not in the top k.
  • Retrieval mismatch: a confident but wrong or outdated passage ranks first.
  • Unsupported or incorrect claims: the context was fine, the answer was not.
  • Instruction-following failures: correct facts, wrong language, length or tone.

Separate retrieval from generation first

The cheapest diagnostic is to score retrieval on its own before judging answers. With a gold passage mapped to each question you can compute Recall@k and MRR without calling a generator at all.

If Recall@5 is 0.88 and your pass rate is 0.87, generation is not your bottleneck. If Recall@5 is 0.95 and the pass rate is 0.80, look at the prompt and the model before re-indexing anything.

python
from relevant import Relevantclient = Relevant()  # reads RELEVANT_API_KEYrun = client.evaluate(    project="support-assistant",    dataset="support-golden-v4",    config={"retriever": "bge-m3-hybrid"},  # retrieval only, no generator    judges=["retrieval"],)print(run.summary)

Attribute every failure to a chunk

Claim-level judging splits an answer into sentences and checks each against the retrieved chunks. A sentence with no supporting chunk is an unsupported claim; a sentence that contradicts a chunk is an incorrect fact.

The useful output is not the label but the pointer: which chunk was missing, which chunk was stale, which sentence went beyond the context. That pointer is what an engineer can act on.

Track the mix, not just the total

Once failures are categorized, plot the category mix over time. A release that cuts retrieval mismatch from 24% to 18% of failures while unsupported claims rise is a different story from one where everything drops evenly.

Share
XLinkedIn
All articles
    • Model comparison
    • Evaluation

    Choosing a generator with evidence

    Comparing models on the same retrieval, the same prompt and the same questions, then reading the per-category differences.

    7 min read

    • CI
    • Evaluation

    Regression gates for RAG in CI

    Block merges that make answers worse, without blocking every merge. Thresholds, sample sizes and what to do when a gate fails.

    10 min read

    • Retrieval
    • Chunking

    Chunking choices that move recall

    On a telecom help center we evaluated, moving from 512-token fixed windows to 320-token semantic chunks cut chunk-boundary loss from 7.2% to 3.1%. Here is what changed.

    7 min read

See why your retrieval fails

Run an evaluation on your own dataset and read every failure with its retrieved chunks, flagged claims and suggested fix.

Free plan includes 2,000 evaluated samples a month