Skip to content

Evaluation & failure analysis for LLM and retrieval systems

Understand why your retrieval fails.

Relevant traces every bad answer back to its cause: the passage that was never retrieved, the claim the context did not support, the instruction the model ignored. Then it shows you which fix moves the number.

Free plan includes 2,000 evaluated samples a month · No credit card

Product preview

Open overview in the dashboard

Overview

Support Assistant · 14,820 samples evaluated in the last 30 days

Pass rate

86.9%

+5.7 pts(improvement)vs previous 30 days

Open failures

1,944

−18%(improvement)vs previous 30 days

Pass rate

Last 30 days vs previous 30 days

  • Last 30 days
  • Previous 30 days

Open failures by category

Last 30 days

  • Unsupported525 · 27%
  • Incorrect408 · 21%
  • Missing369 · 19%
  • Mismatch350 · 18%
  • Instructions214 · 11%
  • Other78 · 4%

Recent failures

Newest first

Five most recent open or triaged failures
IDQuestionCategorySeverity
F-2412How much data can I use in Germany on Unlimited Plus before I get charged?MismatchHigh
F-2411Can I get a refund for the roaming pack I didn't use?UnsupportedCritical
F-2410My account is billed in pounds. What is the late fee if I pay after the due date?InstructionsHigh
F-2409I upgraded from Essential to Unlimited on the 18th. Why is my bill €41.20 instead of €34.90?UnsupportedCritical
F-2408My eSIM stopped working after I reset my phone. Do I need to pay for a new one?UnsupportedHigh

The problem

Aggregate scores tell you that quality dropped. They never tell you why.

A single pass rate hides four different failures: a missing passage, a stale article, an invented fee, an ignored instruction. Each needs a different fix, and guessing wrong costs a release cycle.

  • 01

    Accuracy without explanation

    Your dashboard says 81.2%. It cannot say which questions failed, whether the retriever or the generator was at fault, or what changed.

  • 02

    Retrieval and generation are entangled

    A wrong number can come from a stale article ranked first or from a model that ignored the right one. Without chunk-level attribution, both look the same.

  • 03

    Manual review doesn’t scale

    At 2.7 minutes per answer, reading 14,820 answers a month is 667 reviewer-hours. Most teams read a few percent and hope those were the ones that mattered.

  • 04

    Regressions ship silently

    A prompt tweak fixes billing and quietly breaks roaming. Latency creeps up half a second. Nobody notices until customers do.

How it works

One loop from trace to fix

Relevant fits around the pipeline you already have. Send traces or upload a dataset, run judges you can audit, read failures by cause, and gate the next release on the result.

  1. 01

    Ingest

    Log production traces with the SDK or upload CSV and JSONL datasets with gold answers and source passages.

  2. 02

    Evaluate

    Run rubric-based LLM judges and retrieval metrics on every sample, with rationales you can read and calibrate.

  3. 03

    Diagnose

    Every failure gets a category, a root-cause cluster and chunk-level attribution, so you know where to look first.

  4. 04

    Improve

    Compare fixes on the same questions, promote the winner and block regressions in CI with release gates.

Capabilities

Everything between a failing answer and a shipped fix

Built for the engineers who own retrieval quality: precise metrics, readable evidence and workflows that end in a decision.

  • Evaluation runs

    Rubric-based LLM judges and human review on every sample. Versioned configs, nightly series and release candidates side by side.

  • Retrieval diagnostics

    Recall@k, MRR and NDCG per question, chunk-level attribution, and detection of stale versions and wrong sections.

  • Failure taxonomy and clustering

    Six categories with explicit signals, plus root-cause clusters that group the same mistake across hundreds of questions.

  • Ground-truth review

    A keyboard-first queue to approve, edit or reject gold answers, with the source passage beside every item.

  • Model and retriever comparison

    Hold retrieval fixed and swap generators, or the reverse. Per-metric, per-category and per-question deltas.

  • Regression gates for CI

    Run a gate dataset on every pull request. Absolute floors, deltas against a pinned baseline, and a clear pass or block.

  • Reports and exports

    Weekly quality, release readiness and retrieval health reports. Print to PDF or export any table as CSV.

  • Privacy controls

    Data processed in the region you choose, role-based access for every project, and retention you configure per workspace.

Failure taxonomy

Every failure gets a cause, not just a score

Relevant labels each failed answer with one of six categories using explicit signals: claim-level grounding, gold-passage rank, article version and instruction checks. You fix the biggest category first.

Share of open failures · Support Assistant, last 30 days1,944 total
  • Unsupported claims

    The answer states something that is not present in the retrieved context.

    27%525 open failures

  • Incorrect factual answers

    The answer contradicts the knowledge base, usually by misreading or conflating two policies.

    21%408 open failures

  • Missing context

    A relevant article exists but was not retrieved, or was only partially retrieved.

    19%369 open failures

  • Retrieval mismatch

    The retriever returned the wrong article, the wrong section or a stale version.

    18%350 open failures

  • Instruction-following failures

    The answer ignored format, tone, policy or language constraints from the system prompt.

    11%214 open failures

  • Other

    Over-refusal, wrong language without a prompt constraint, or truncated answers.

    4%78 open failures

Explore failure analysis

Retrieval was correct here: both relevant chunks ranked first and second. The generator subtracted two amounts from the question and presented the difference as a fee.

F-2409Unsupported claimsCriticalsupport-assistant v3.4 · billing-edge-cases

Question

I upgraded from Essential to Unlimited on the 18th. Why is my bill €41.20 instead of €34.90?

Model answer · Claude Sonnet 5.5

Your bill is higher because upgrades are charged a one-off €6.30 upgrade fee (unsupported by retrieved context) kb-plans-008.

Retrieved context

gold rank 1

  1. #1
    Changing your plan· Upgrades

    Upgrades take effect immediately. Your next bill shows a pro-rated charge for the new plan and a credit for the old plan.

    0.84relevant
  2. #2
    Understanding your monthly bill· Pro-rated charges

    Pro-rated charges cover the days between a plan change and the end of the cycle.

    0.80relevant
  3. #3
    Understanding your monthly bill· One-off charges

    One-off charges include roaming packs, device payments and late fees.

    0.66off-topic
Judge · score 0.06
The answer invents an upgrade fee by subtracting the two amounts in the question. The retrieved context explains proration and contains no fee.
Suggested fix
Add a rule to explain amounts only through retrieved bill components; add proration worked examples to the few-shot set.

Use cases

Wherever answers depend on what you retrieve

The same loop works for any system that retrieves before it responds. What changes is which failures cost the most.

  • Customer support assistants

    Answers about fees, plans and policies where a single invented number becomes a complaint.

    What you debug

    • Unsupported fees and refund promises
    • Stale policy versions outranking current ones
    • Tone, language and escalation rules
    • Faithfulness
    • Pass rate
    • Instruction adherence
  • Internal knowledge search

    Employees searching wikis, tickets and runbooks spread across systems with uneven freshness.

    What you debug

    • Duplicate and near-duplicate pages
    • Permission-filtered recall gaps
    • Queries phrased in team jargon
    • Recall@k
    • MRR
    • Context precision
  • Documentation Q&A

    Developer docs with versioned APIs, where the right answer depends on the version in use.

    What you debug

    • Answers from the wrong API version
    • Code samples that never appeared in the docs
    • Chunks that split a parameter from its default
    • NDCG@10
    • Answer correctness
    • Faithfulness
  • Legal and finance research

    Assistants that summarize contracts, filings and regulation, where every claim needs a source.

    What you debug

    • Claims without a supporting clause
    • Exceptions retrieved without their rule
    • Dates and thresholds paraphrased incorrectly
    • Claim-level grounding
    • Context recall
    • Human review
  • E-commerce search relevance

    Product search and shopping assistants ranking thousands of similar items for intent-heavy queries.

    What you debug

    • Attribute mismatches (size, color, compatibility)
    • Reranker regressions on long-tail queries
    • Out-of-stock items ranked first
    • NDCG@10
    • MRR
    • Recall@k

For developers

Evaluations that live in your code

Run the same evaluation from a notebook, a script or a pull request. Configs are versioned, results are reproducible, and thresholds fail the build when quality drops.

  • Typed SDKs for Python and TypeScript, plus a CLI for CI pipelines
  • Log traces from LangChain, LlamaIndex or your own retrieval code
  • Gate pull requests on the same thresholds you use for releases
  • Every result lands in the dashboard with chunks, verdicts and rationales

scripts/evaluate.ts

import { Relevant } from "@relevant/sdk";const relevant = new Relevant({ apiKey: process.env.RELEVANT_API_KEY });const run = await relevant.evaluate({  project: "support-assistant",  dataset: "support-golden-v4",  config: {    generator: "claude-sonnet-5-5",    retriever: "bge-m3-hybrid",    prompt: "support-prompt@v7",  },  judges: ["llm:rubric-v3", "retrieval"],  gate: { passRate: 0.85, faithfulness: 0.9, latencyP95Ms: 4000 },});console.log(run.summary);// pass_rate 0.869 · faithfulness 0.91 · recall@5 0.88for (const failure of run.failures({ category: "unsupported_claims" })) {  console.log(failure.id, failure.rootCause);}

Start in minutes

Find out why your last bad answer happened.

Upload a dataset or connect your traces, run your first evaluation and read every failure by cause. The free plan includes 2,000 evaluated samples a month.

No credit card
Free plan, one seat, three datasets
Your region
Data processed in the region you choose
Talk to us
Team and Enterprise onboarding on request

Questions about pricing or deployment? Contact the team or compare plans.