Skip to content

Product

Evaluation, diagnosis and review in one loop

Run evaluations on your own datasets, see which chunk caused each failure, and confirm every fix before it ships.

01 · Evaluation runs

Score every answer against a rubric you can read

Run a dataset through your pipeline and let LLM judges, rule checks and human reviewers score each answer. Every run records the dataset version, retriever, prompt and judge, so any result can be reproduced.

  • Rubric-based LLM judges

    Claim-level faithfulness, answer correctness and instruction adherence, each with a written rationale.

  • Human review alongside

    Route a share of answers to double-blind human scoring and adjudicate disagreements.

  • Reproducible by default

    Dataset version, retriever, chunking, prompt and judge are stored with the run and compared with a baseline.

How evaluation works

run_2026-10-06_v3-4-rc

support-assistant v3.4 · release candidate

Completed
Claude Sonnet 5.5bge-m3 · hybridsupport-prompt v7support-golden-v4Opus 5.5 · rubric v3
Pass rate
86.9%+5.7 pts(improvement)
Faithfulness
0.91+0.04(improvement)
Answer correctness
0.88+0.04(improvement)
Context recall
0.84+0.05(improvement)
Latency p95
3.8s+0.1s(regression)
Cost per request
$0.0042−$0.0004(improvement)

Deltas are against support-assistant v3.3 on the same dataset.

Recent runs

  • v3.4 · synthetic-hard-negatives v1.3

    synthetic-hard-negatives

    Running
    340 / 540
  • v3.4 · Claude Haiku 5.5 comparison

    support-golden-v4

    Failed
    418 / 1,240
  • v3.4 · Gemini 2.5 Pro comparison

    support-golden-v4

    Completed
    83.8%
  • v3.4 · GPT-5 comparison

    support-golden-v4

    Completed
    85.3%
  • support-assistant v3.4 · release candidate

    support-golden-v4

    Completed
    86.9%
run_2026-10-06_v3-4-rc compared with run_2026-08-03_v3-3-release

02 · Retrieval diagnostics

See which chunk caused the failure

Every failure carries the ranked chunks your retriever returned, with scores, article versions and a relevance mark against the gold passage. Retrieval problems stop hiding inside generation errors.

  • Chunk-level attribution

    Know whether the gold passage was missing, outranked or split across chunks.

  • Ranking metrics

    Recall@k, MRR and NDCG@10 for every run, dataset and domain.

  • Mismatch detection

    Stale article versions and wrong-section hits are flagged on the chunk, not averaged away.

How diagnosis works
F-2412Retrieval mismatchHigh

How much data can I use in Germany on Unlimited Plus before I get charged?

Retrieved chunks

  1. 1

    Roaming fair-use (2025 edition) · Fair-use allowance

    Unlimited Plus customers can use up to 25 GB per month in Zone 1 before fair-use charges apply.

    v2025.11Superseded version
    Score 0.83Not relevant
  2. 2

    Roaming fair-use · Fair-use allowance

    From 1 September 2026, Unlimited Plus includes 35 GB of Zone 1 roaming data per month.

    v2026.09Gold passage
    Score 0.81Relevant
  3. 3

    Roaming zones and countries · Zone 1: EU and EEA

    Zone 1 covers the 27 EU member states plus Iceland, Liechtenstein and Norway.

    v2026.09
    Score 0.74Not relevant

Model answer

Unlimited Plus includes Flagged by judge: 25 GB of roaming data per month in the EU kb-roaming-011. After that, data is charged at €0.0024 per MB.

Root cause: The 2025 edition of the fair-use article is still indexed and outranks the current version by 0.02.

Recall@5
0.88
MRR
0.74
NDCG@10
0.81
Stale-version rate
1.6%

Retrieval metrics for support-assistant v3.4 on support-golden-v4. Stale-version rate covers the last 30 days.

Every retrieved chunk is scored, versioned and marked relevant or not against the gold passage.

03 · Failure taxonomy and clustering

Six failure categories, grouped by root cause

Every failed answer lands in one category with the signals that put it there, so 'accuracy dropped' becomes 'unsupported claims on refund questions'. Related failures are clustered, so one fix closes many cases.

  • Consistent categories

    Unsupported claims, incorrect facts, missing context, retrieval mismatch, instruction-following and other, defined the same way on every run.

  • Root-cause clusters

    Failures that share a cause, such as a superseded article, are grouped around a representative question.

  • A first fix for each

    Each category carries a playbook, so triage ends with an action rather than a label.

How categories are assigned

1,944 open failures

Unsupported claims

The answer states something that is not present in the retrieved context.

Detection signals

  • Claim-level faithfulness below 0.85
  • Numbers, fees or durations absent from every retrieved chunk
  • Citations that do not contain the cited sentence

First fix: Require a citation after each factual sentence and refuse when none applies

1,944 open failures, last 30 days. Select a category to see how it is detected.

04 · Ground-truth review

Keep gold answers correct as your knowledge base changes

A wrong gold answer produces a false failure. Failures, judge disagreements and knowledge-base updates feed a review queue where reviewers approve, edit, reject or flag proposed answers in seconds.

  • Keyboard-first queue

    Approve, edit, reject or flag with A, E, R and F. Median review time is 2.7 minutes per item.

  • Source beside every proposal

    Each card shows the article, section and excerpt the proposed answer is based on.

  • Versioned datasets

    Approved and edited answers feed the next version of the dataset, and approved items export as CSV.

Open the review queue

412 awaiting review

1 of 3
gt_1042roaming-and-travelConfidence 0.94Pending

How much roaming data does Unlimited Plus include in Zone 1?

Current gold answer

Unlimited Plus includes 25 GB of roaming data per month in Zone 1.

Proposed gold answer

Unlimited Plus includes 35 GB of roaming data per month in Zone 1 (EU and EEA). Above 35 GB, data costs €0.0024 per MB until the end of the billing cycle.

kb-roaming-011 · Fair-use allowance

“From 1 September 2026, Unlimited Plus includes 35 GB of Zone 1 roaming data per month. Usage above the allowance is charged at €0.0024/MB.”

Roaming fair-use · submitted by Daniel Okafor

Review cards are keyboard-first. Click the card and press A, E, R or F; J and K move between items.

05 · Model and retriever comparison

Compare generators and retrievers on the same questions

Change one thing at a time. Run configurations over the same dataset with the same judge, then read deltas against a pinned baseline for quality, latency and cost.

  • Baseline-pinned deltas

    Challengers show signed deltas, and regressions are marked as clearly as improvements.

  • By dataset and category

    See where a configuration wins or regresses across datasets, domains and failure categories.

  • Side-by-side answers

    Open any question and read each configuration's answer next to the gold answer.

Open model comparison
  • Claude Sonnet 5.5

    Baseline
    Pass rate
    86.9%
    Faithfulness
    0.91
    Latency p95
    3.8s
    Cost per request
    $0.0042
  • GPT-5

    Pass rate
    85.3%−1.6 pts(regression)
    Faithfulness
    0.90−0.01(regression)
    Latency p95
    5.2s+1.4s(regression)
    Cost per request
    $0.0061+$0.0019(regression)
  • Gemini 2.5 Pro

    Pass rate
    83.8%−3.1 pts(regression)
    Faithfulness
    0.88−0.03(regression)
    Latency p95
    4.6s+0.8s(regression)
    Cost per request
    $0.0048+$0.0006(regression)

On support-golden-v4, Claude Sonnet 5.5 has the highest pass rate.

Same dataset, same judge, same rubric. Challengers are scored against the pinned baseline.

06 · Regression gates and reports

Block regressions in CI and report on every release

Set thresholds for quality, latency and critical failures. Relevant runs the gate on each pull request and fails the build when a threshold is missed. Weekly and release reports export as CSV or print as PDF.

  • Gates from your repository

    Thresholds live in a config file next to your code and run in GitHub Actions or any CI system.

  • Release readiness

    One report with gate checks, metric deltas against the last release and known issues.

  • Scheduled and exportable

    Weekly quality reports on a schedule, CSV exports and print-ready layouts.

See the CI example

Regression gate · support-assistant v3.4

run_2026-10-06_v3-4-rc vs production baseline

7 of 7 checks passed
  • Passed: Pass rate ≥ 85.0%86.9%
  • Passed: Faithfulness ≥ 0.900.91
  • Passed: Answer correctness ≥ 0.850.88
  • Passed: Context recall ≥ 0.800.84
  • Passed: Instruction adherence ≥ 0.920.93
  • Passed: p95 latency ≤ 4.0s3.8s
  • Passed: No new critical failures in account-security0 new

Latest reports

  • Failure trends · Jul 11 – Oct 8

    On demand · generated Oct 7 by Owen Reid

    CSV
  • Release readiness · v3.4

    Per release · generated Oct 6 by Maya Collins

    PDF
  • Weekly quality · Sep 28 – Oct 4

    Weekly · generated Oct 5 by Daniel Okafor

    PDF
Gate thresholds live in your repository; reports export as CSV and print as PDF.

Integrations

Connect the stack you already run

Relevant reads from your retriever and calls the models you choose. There is nothing to replace in your pipeline.

Frameworks and tracing

Capture retrievals and answers where they happen.

  • LangChain

    Trace callbacks and dataset import

  • LlamaIndex

    Retriever and query-engine traces

  • Python SDK

    log_trace from any pipeline

  • TypeScript SDK

    logTrace from Node and edge runtimes

  • relevant CLI

    datasets push, evaluate and gate in CI

Running something else? Import a JSONL export of your traces. See how ingestion works

Models and judges

Use any provider as a generator or as a judge.

  • OpenAI

    Generators and judges

  • Anthropic

    Generators and judges

  • Vertex AI

    Gemini generators in your project

  • Bedrock

    Models served from your AWS account

  • Cohere

    Embeddings, rerank and generation

Vector stores and search

Read ranked chunks and scores from your index.

  • Pinecone

    Namespaces and metadata filters

  • Weaviate

    Hybrid and vector queries

  • pgvector

    Postgres similarity search

  • Elasticsearch

    BM25, kNN and hybrid

  • Qdrant

    Payload filters and scores

Security and privacy

Your evaluation data stays yours

Evaluation data often contains real customer questions. Relevant keeps it where you put it and uses it only for your evaluations.

Read the security overview
  • Processed in the region you choose

    Pick a processing region for each workspace. Enterprise adds VPC and on-premises connectors so traces never leave your network.

  • Your data is not training data

    Datasets, traces and reviews are used to run your evaluations and for nothing else. Judge models run only on the questions and answers you send.

  • Access control and audit trails

    Role-based access on Team and above. SSO with SAML or OIDC, SCIM provisioning and audit logs on Enterprise.

  • Retention and deletion you control

    Retention is set per plan from 7 to 90 days, or custom on Enterprise. Data is encrypted in transit and at rest, and you can delete datasets and runs at any time.

See your own failures, not an average.

Create a free workspace, upload a dataset and get a failure breakdown from your first run.

Free plan includes 2,000 evaluated samples a month. No credit card.

In your first session

  • Upload a dataset or log traces from your pipeline
  • Run an evaluation with LLM judges and rule checks
  • Read a failure breakdown by category and cause