Most teams start evaluating a RAG system with one number: the share of answers a judge marks as correct. It is a reasonable start and a poor place to stay. Aggregate scores average over failure modes that have different causes, different owners and different fixes.
One number, four failure modes
When an answer is wrong, the cause usually sits in one of four places: the retriever never surfaced the right passage, it surfaced a stale or wrong passage, the generator ignored or contradicted a good passage, or the generator followed the passage but broke a format or policy rule.
A pass rate that moves from 81% to 84% tells you nothing about which of these moved. It is common to see retrieval improve while generation regresses, with the two effects canceling in the aggregate.
- Missing context: the gold passage is not in the top k.
- Retrieval mismatch: a confident but wrong or outdated passage ranks first.
- Unsupported or incorrect claims: the context was fine, the answer was not.
- Instruction-following failures: correct facts, wrong language, length or tone.
Separate retrieval from generation first
The cheapest diagnostic is to score retrieval on its own before judging answers. With a gold passage mapped to each question you can compute Recall@k and MRR without calling a generator at all.
If Recall@5 is 0.88 and your pass rate is 0.87, generation is not your bottleneck. If Recall@5 is 0.95 and the pass rate is 0.80, look at the prompt and the model before re-indexing anything.
from relevant import Relevantclient = Relevant() # reads RELEVANT_API_KEYrun = client.evaluate( project="support-assistant", dataset="support-golden-v4", config={"retriever": "bge-m3-hybrid"}, # retrieval only, no generator judges=["retrieval"],)print(run.summary)Attribute every failure to a chunk
Claim-level judging splits an answer into sentences and checks each against the retrieved chunks. A sentence with no supporting chunk is an unsupported claim; a sentence that contradicts a chunk is an incorrect fact.
The useful output is not the label but the pointer: which chunk was missing, which chunk was stale, which sentence went beyond the context. That pointer is what an engineer can act on.
Track the mix, not just the total
Once failures are categorized, plot the category mix over time. A release that cuts retrieval mismatch from 24% to 18% of failures while unsupported claims rise is a different story from one where everything drops evenly.