Product
Evaluation, diagnosis and review in one loop
Run evaluations on your own datasets, see which chunk caused each failure, and confirm every fix before it ships.
01 · Evaluation runs
Score every answer against a rubric you can read
Run a dataset through your pipeline and let LLM judges, rule checks and human reviewers score each answer. Every run records the dataset version, retriever, prompt and judge, so any result can be reproduced.
Rubric-based LLM judges
Claim-level faithfulness, answer correctness and instruction adherence, each with a written rationale.
Human review alongside
Route a share of answers to double-blind human scoring and adjudicate disagreements.
Reproducible by default
Dataset version, retriever, chunking, prompt and judge are stored with the run and compared with a baseline.
run_2026-10-06_v3-4-rc
support-assistant v3.4 · release candidate
- Pass rate
- 86.9%+5.7 pts(improvement)
- Faithfulness
- 0.91+0.04(improvement)
- Answer correctness
- 0.88+0.04(improvement)
- Context recall
- 0.84+0.05(improvement)
- Latency p95
- 3.8s+0.1s(regression)
- Cost per request
- $0.0042−$0.0004(improvement)
Deltas are against support-assistant v3.3 on the same dataset.
Recent runs
- Running
v3.4 · synthetic-hard-negatives v1.3
synthetic-hard-negatives
340 / 540 - Failed
v3.4 · Claude Haiku 5.5 comparison
support-golden-v4
418 / 1,240 - Completed
v3.4 · Gemini 2.5 Pro comparison
support-golden-v4
83.8% - Completed
v3.4 · GPT-5 comparison
support-golden-v4
85.3% - Completed
support-assistant v3.4 · release candidate
support-golden-v4
86.9%
02 · Retrieval diagnostics
See which chunk caused the failure
Every failure carries the ranked chunks your retriever returned, with scores, article versions and a relevance mark against the gold passage. Retrieval problems stop hiding inside generation errors.
Chunk-level attribution
Know whether the gold passage was missing, outranked or split across chunks.
Ranking metrics
Recall@k, MRR and NDCG@10 for every run, dataset and domain.
Mismatch detection
Stale article versions and wrong-section hits are flagged on the chunk, not averaged away.
How much data can I use in Germany on Unlimited Plus before I get charged?
Retrieved chunks
- 1
Roaming fair-use (2025 edition) · Fair-use allowance
Unlimited Plus customers can use up to 25 GB per month in Zone 1 before fair-use charges apply.
v2025.11Superseded versionScore 0.83Not relevant - 2
Roaming fair-use · Fair-use allowance
From 1 September 2026, Unlimited Plus includes 35 GB of Zone 1 roaming data per month.
v2026.09Gold passageScore 0.81Relevant - 3
Roaming zones and countries · Zone 1: EU and EEA
Zone 1 covers the 27 EU member states plus Iceland, Liechtenstein and Norway.
v2026.09Score 0.74Not relevant
Model answer
Unlimited Plus includes Flagged by judge: 25 GB of roaming data per month in the EU kb-roaming-011. After that, data is charged at €0.0024 per MB.
Root cause: The 2025 edition of the fair-use article is still indexed and outranks the current version by 0.02.
- Recall@5
- 0.88
- MRR
- 0.74
- NDCG@10
- 0.81
- Stale-version rate
- 1.6%
Retrieval metrics for support-assistant v3.4 on support-golden-v4. Stale-version rate covers the last 30 days.
03 · Failure taxonomy and clustering
Six failure categories, grouped by root cause
Every failed answer lands in one category with the signals that put it there, so 'accuracy dropped' becomes 'unsupported claims on refund questions'. Related failures are clustered, so one fix closes many cases.
Consistent categories
Unsupported claims, incorrect facts, missing context, retrieval mismatch, instruction-following and other, defined the same way on every run.
Root-cause clusters
Failures that share a cause, such as a superseded article, are grouped around a representative question.
A first fix for each
Each category carries a playbook, so triage ends with an action rather than a label.
1,944 open failures
Unsupported claims
The answer states something that is not present in the retrieved context.
Detection signals
- Claim-level faithfulness below 0.85
- Numbers, fees or durations absent from every retrieved chunk
- Citations that do not contain the cited sentence
First fix: Require a citation after each factual sentence and refuse when none applies
04 · Ground-truth review
Keep gold answers correct as your knowledge base changes
A wrong gold answer produces a false failure. Failures, judge disagreements and knowledge-base updates feed a review queue where reviewers approve, edit, reject or flag proposed answers in seconds.
Keyboard-first queue
Approve, edit, reject or flag with A, E, R and F. Median review time is 2.7 minutes per item.
Source beside every proposal
Each card shows the article, section and excerpt the proposed answer is based on.
Versioned datasets
Approved and edited answers feed the next version of the dataset, and approved items export as CSV.
412 awaiting review
How much roaming data does Unlimited Plus include in Zone 1?
Current gold answer
Unlimited Plus includes 25 GB of roaming data per month in Zone 1.
Proposed gold answer
Unlimited Plus includes 35 GB of roaming data per month in Zone 1 (EU and EEA). Above 35 GB, data costs €0.0024 per MB until the end of the billing cycle.
kb-roaming-011 · Fair-use allowance
“From 1 September 2026, Unlimited Plus includes 35 GB of Zone 1 roaming data per month. Usage above the allowance is charged at €0.0024/MB.”
Roaming fair-use · submitted by Daniel Okafor
05 · Model and retriever comparison
Compare generators and retrievers on the same questions
Change one thing at a time. Run configurations over the same dataset with the same judge, then read deltas against a pinned baseline for quality, latency and cost.
Baseline-pinned deltas
Challengers show signed deltas, and regressions are marked as clearly as improvements.
By dataset and category
See where a configuration wins or regresses across datasets, domains and failure categories.
Side-by-side answers
Open any question and read each configuration's answer next to the gold answer.
Claude Sonnet 5.5
Baseline- Pass rate
- 86.9%
- Faithfulness
- 0.91
- Latency p95
- 3.8s
- Cost per request
- $0.0042
GPT-5
- Pass rate
- 85.3%−1.6 pts(regression)
- Faithfulness
- 0.90−0.01(regression)
- Latency p95
- 5.2s+1.4s(regression)
- Cost per request
- $0.0061+$0.0019(regression)
Gemini 2.5 Pro
- Pass rate
- 83.8%−3.1 pts(regression)
- Faithfulness
- 0.88−0.03(regression)
- Latency p95
- 4.6s+0.8s(regression)
- Cost per request
- $0.0048+$0.0006(regression)
On support-golden-v4, Claude Sonnet 5.5 has the highest pass rate.
06 · Regression gates and reports
Block regressions in CI and report on every release
Set thresholds for quality, latency and critical failures. Relevant runs the gate on each pull request and fails the build when a threshold is missed. Weekly and release reports export as CSV or print as PDF.
Gates from your repository
Thresholds live in a config file next to your code and run in GitHub Actions or any CI system.
Release readiness
One report with gate checks, metric deltas against the last release and known issues.
Scheduled and exportable
Weekly quality reports on a schedule, CSV exports and print-ready layouts.
Regression gate · support-assistant v3.4
run_2026-10-06_v3-4-rc vs production baseline
- Passed: Pass rate ≥ 85.0%86.9%
- Passed: Faithfulness ≥ 0.900.91
- Passed: Answer correctness ≥ 0.850.88
- Passed: Context recall ≥ 0.800.84
- Passed: Instruction adherence ≥ 0.920.93
- Passed: p95 latency ≤ 4.0s3.8s
- Passed: No new critical failures in account-security0 new
Latest reports
- CSV
Failure trends · Jul 11 – Oct 8
On demand · generated Oct 7 by Owen Reid
Release readiness · v3.4
Per release · generated Oct 6 by Maya Collins
Weekly quality · Sep 28 – Oct 4
Weekly · generated Oct 5 by Daniel Okafor
Integrations
Connect the stack you already run
Relevant reads from your retriever and calls the models you choose. There is nothing to replace in your pipeline.
Frameworks and tracing
Capture retrievals and answers where they happen.
LangChain
Trace callbacks and dataset import
LlamaIndex
Retriever and query-engine traces
Python SDK
log_trace from any pipeline
TypeScript SDK
logTrace from Node and edge runtimes
relevant CLI
datasets push, evaluate and gate in CI
Running something else? Import a JSONL export of your traces. See how ingestion works
Models and judges
Use any provider as a generator or as a judge.
OpenAI
Generators and judges
Anthropic
Generators and judges
Vertex AI
Gemini generators in your project
Bedrock
Models served from your AWS account
Cohere
Embeddings, rerank and generation
Vector stores and search
Read ranked chunks and scores from your index.
Pinecone
Namespaces and metadata filters
Weaviate
Hybrid and vector queries
pgvector
Postgres similarity search
Elasticsearch
BM25, kNN and hybrid
Qdrant
Payload filters and scores
Security and privacy
Your evaluation data stays yours
Evaluation data often contains real customer questions. Relevant keeps it where you put it and uses it only for your evaluations.
Processed in the region you choose
Pick a processing region for each workspace. Enterprise adds VPC and on-premises connectors so traces never leave your network.
Your data is not training data
Datasets, traces and reviews are used to run your evaluations and for nothing else. Judge models run only on the questions and answers you send.
Access control and audit trails
Role-based access on Team and above. SSO with SAML or OIDC, SCIM provisioning and audit logs on Enterprise.
Retention and deletion you control
Retention is set per plan from 7 to 90 days, or custom on Enterprise. Data is encrypted in transit and at rest, and you can delete datasets and runs at any time.
See your own failures, not an average.
Create a free workspace, upload a dataset and get a failure breakdown from your first run.
Free plan includes 2,000 evaluated samples a month. No credit card.
In your first session
- Upload a dataset or log traces from your pipeline
- Run an evaluation with LLM judges and rule checks
- Read a failure breakdown by category and cause