How it works
From raw traces to a release you can defend
Relevant runs one loop. Ingest your questions and traces, evaluate every answer, diagnose each failure to its cause, and improve with evidence that holds up in review.
The data model
Five objects, traced end to end
Every result in Relevant resolves to the same chain, so a failing verdict always leads back to the chunk and the dataset item behind it.
- 01
Dataset
A versioned set of questions with gold answers and source passages.
support-golden-v4 · v4.6 · 1,240 items
- 02
Query
One question, replayed through your pipeline with a fixed configuration.
"How much data can I use in Germany on…"
- 03
Retrieved set
Ranked chunks with scores, article versions and a relevance mark.
Gold passage at rank 2
- 04
Answer
The generated response, split into claims and citations.
2 claims · 1 unsupported
- 05
Verdict
Judge scores, pass or fail, and one failure category.
Fail · Retrieval mismatch
Trace F-2412, run_2026-10-08_nightly.
01 · Ingest
Bring the questions and traces you already have
Start from a dataset of questions with gold answers, from logged production traces, or both. Relevant validates the data on the way in, so the first run starts from clean input.
Log traces from your pipeline
Send one trace per request, after the retrieval and generation calls you already make. Each trace stores the ranked chunks, scores, article versions and the final answer, which is what chunk-level diagnosis needs.
# pip install relevant
from relevant import Relevant
client = Relevant() # reads RELEVANT_API_KEY (keys start with rlv_)
chunks = retriever.search(question, k=8)
answer = generator.answer(question, chunks)
client.log_trace(
project="support-assistant",
dataset="support-golden-v4",
question=question,
retrieved=[
{
"id": c.id,
"article": c.article_id,
"section": c.section,
"score": c.score,
"version": c.version,
}
for c in chunks
],
answer={
"text": answer.text,
"model": "claude-sonnet-5-5",
"prompt": "support-prompt@v7",
},
)
Upload a dataset
Use CSV or JSONL with one row per question. Only question and gold_answer are required. Adding source_article and source_section turns on chunk-level attribution for that row.
id,question,gold_answer,domain,language,source_article,source_section,difficulty
gt_1008,How long does it take to port my number to Harbor?,"Once we receive a valid porting code, number transfers complete within 1 business day. Keep your old SIM active until your Harbor SIM shows signal; your old provider closes the line automatically.",porting,en,kb-porting-002,Timeline,easy
gt_1005,Is SIM swap protection turned on by default?,Yes. SIM swap protection is on for every line. A SIM or eSIM swap requires a one-time code sent to the account holder's verified email and photo ID in store. You can add a 24-hour hold on swaps in My Harbor under Security.,account,en,kb-account-005,How SIM swap protection works,medium
gt_1002,"If I upgrade my plan mid-month, how is my next bill calculated?","Upgrades take effect immediately. Your next bill includes a pro-rated charge for the new plan from the upgrade date to the end of the cycle and a credit for the unused days of the old plan, followed by the full monthly price of the new plan.",plans,en,kb-plans-008,Upgrades,medium
UTF-8, one row per question. Rows in different languages can share a file.
| Field | Type | Required | Notes |
|---|---|---|---|
id | string | Optional | Stable row identifier. Generated when omitted. |
question | string | Required | The customer question as it was asked. |
gold_answer | string | Required | The verified reference answer a reviewer would approve. |
domain | string | Optional | Topic label for per-domain breakdowns, for example roaming or billing. |
language | string | Optional | Language code of the question, for example en. Detected from the question when omitted. |
source_article | string | Optional | Knowledge-base article id. Enables chunk-level attribution. |
source_section | string | Optional | Heading of the section that supports the answer. |
difficulty | easy | medium | hard | Optional | Used to filter results and weight reports. |
relevant datasets push ./support-golden-v4.jsonl \
--name support-golden-v4 \
--version v4.6 \
--validate
What happens on import
Every file is checked before anything is written, and the results are shown before you confirm the import.
Schema and encoding
Rows with a missing question or gold answer are reported by line number before anything is imported.
Near-duplicate detection
Questions that differ only in wording are flagged, so one case is not counted twice.
Source mapping coverage
The share of rows linked to an article and section is shown, because attribution depends on it.
02 · Evaluate
Score every answer with judges you can audit
A run replays each dataset question through a pinned configuration and scores the result with the judges you choose. The configuration is a file in your repository, so a run can always be reproduced.
Describe the run in one file
The file pins the dataset version, generator, retriever, chunking, prompt and judges. Change one line, run again and compare.
# relevant.yaml
project: support-assistant
dataset: support-golden-v4@v4.6
pipeline:
generator: claude-sonnet-5-5
retriever: bge-m3-hybrid # BM25 + dense, reciprocal rank fusion
chunking: semantic-320
prompt: support-prompt@v7
top_k: 8
judges:
- type: llm
model: claude-opus-5-5
rubric: rubric-v3
temperature: 0
- type: rules
checks: [language_match, max_words_120, citation_ids_exist]
- type: human
share: 0.10 # double-blind, adjudicated on disagreement
metrics:
- pass_rate
- faithfulness
- answer_correctness
- context_precision
- context_recall
- recall@5
- mrr
- ndcg@10
relevant evaluate --config relevant.yaml --name "v3.4 release candidate"
A rubric you can read
Rubrics are versioned and stored with the run. Each criterion is a plain-language question, a scale and a threshold, so a judge's verdict can be argued with.
Support answer rubric
rubric-v3 · judge: Claude Opus 5.5 · temperature 0
| Criterion | Question put to the judge | Scale | Answer fails when |
|---|---|---|---|
| Faithfulnessgrounded | Is every factual claim supported by the retrieved context? | Pass or fail, from a 0 to 1 claim-level score | Below 0.85 |
| Answer correctnesscorrect | Do the facts (plan, fee, date, zone) match the gold answer? | Pass or fail, from a 0 to 1 score | Below 0.70 |
| Completenesscomplete | Does the answer cover every point in the gold answer? | 1 to 5 | Below 4 |
| Instruction adherencepolicy | Does the answer respect tone, language, length and escalation rules? | Pass or fail, rule checks plus judge | Any hard rule fails |
An answer passes only when every criterion passes. The judge returns a written rationale with each verdict, so every failure can be audited.
Choose your judges
Combine judges on the same run. Their verdicts are shown side by side for every answer.
- Default
LLM judge
Claude Opus 5.5 scores each answer against rubric v3 at temperature 0, decomposes answers into claims, and stores a written rationale with every verdict.
- Team and Enterprise
Human review
Your reviewers score a double-blind share of answers next to the LLM judge. Disagreements go to adjudication and feed the ground-truth review queue.
- Deterministic
Rule checks
Language match, length limits, required phrases, citation ids that exist and PII patterns. Cheap, fast and identical on every run.
- Pro and above
Your own judge
Point Relevant at any judge model endpoint you control and attach your own rubric. Judge model usage is billed by your provider.
Metrics every run reports
Retrieval and generation are measured separately, so a drop in one is never blamed on the other.
Outcome
- Pass ratehigher is better
- Share of samples the judge marked as passing all rubric checks.
Retrieval
- Context precisionhigher is better
- Rank-weighted precision of relevant chunks passed to the generator.
- Context recallhigher is better
- Share of gold-answer statements supported by the retrieved context.
- Recall@5higher is better
- Share of gold passages found in the top 5 retrieved chunks.
- MRRhigher is better
- Mean reciprocal rank of the first relevant chunk.
- NDCG@10higher is better
- Graded ranking quality of the top 10, normalized by the ideal order.
Generation
- Faithfulnesshigher is better
- Share of answer claims supported by the retrieved context.
- Answer correctnesshigher is better
- Agreement with the gold answer on facts, not style.
- Instruction adherencehigher is better
- Respect for tone, format, length, language and policy constraints.
Operational
- Latency p50lower is better
- Median end-to-end time to final token.
- Latency p95lower is better
- 95th percentile end-to-end time to final token.
- Cost per requestlower is better
- Model, embedding and reranking cost per answered request (USD).
03 · Diagnose
Trace every failure to the chunk that caused it
Relevant does not stop at a score. Each failed answer is labeled with one category and linked to the retrieved chunks, so a regression points to a cause instead of an average.
Walk through a failure
Pick a failure to see the evidence the judges had and the staged checks that assigned its category.
Choose a failure
Showing F-2412: Superseded article ranks first. Assigned category: Retrieval mismatch.
How much data can I use in Germany on Unlimited Plus before I get charged?
Retrieved chunks
- 1Roaming fair-use (2025 edition)Fair-use allowance · v2025.11 · superseded0.83Not relevant
- 2Roaming fair-useFair-use allowance · v2026.090.81Relevant
- 3Roaming zones and countriesZone 1: EU and EEA · v2026.090.74Not relevant
Model answer
Unlimited Plus includes Flagged by judge: 25 GB of roaming data per month in the EU kb-roaming-011. After that, data is charged at €0.0024 per MB.
Staged checks
Stages run in order. The first failing stage decides the category.
1 RetrievalFailed
Was the gold passage retrieved and ranked first among the relevant chunks?
The gold passage is at rank 2. A superseded 2025 edition of the article ranks first, and the generator used its 25 GB figure.
2 GroundingNot needed
Is every claim in the answer supported by the retrieved context?
3 CorrectnessNot needed
Does the answer agree with the gold answer on every key fact?
4 InstructionsNot needed
Were the language, format and policy rules followed?
Assigned category
Retrieval mismatch
Root cause: The 2025 edition of the fair-use article is still indexed and outranks the current version by 0.02.
Suggested fix: Filter retrieval to effective-dated articles (version >= v2026.09) and remove kb-roaming-011-2025 from the index.
How categories are assigned
The signals below are computed for every failed answer. The first stage that fails decides the category; the rest are recorded but do not change it.
| Category | Detection signals | First fix | Open |
|---|---|---|---|
| Unsupported claims |
| Require a citation after each factual sentence and refuse when none applies | 52527% |
| Incorrect factual answers |
| Check for a stale article version before changing the prompt | 40821% |
| Missing context |
| Add customer-phrased synonyms and regional terms to query expansion | 36919% |
| Retrieval mismatch |
| Filter or boost by effective date and locale metadata | 35018% |
| Instruction-following failures |
| Move hard constraints into automated rule checks | 21411% |
| Other |
| Relax refusal thresholds for intents with strong retrieval | 784% |
- Unsupported claims525 open · 27%
Detection signals
- Claim-level faithfulness below 0.85
- Numbers, fees or durations absent from every retrieved chunk
- Citations that do not contain the cited sentence
First fix: Require a citation after each factual sentence and refuse when none applies
- Incorrect factual answers408 open · 21%
Detection signals
- Answer correctness below 0.7 while faithfulness stays high
- Disagreement with gold on a key entity: plan, fee, date or zone
- Reviewer edits the fact but keeps the structure
First fix: Check for a stale article version before changing the prompt
- Missing context369 open · 19%
Detection signals
- Context recall below 0.7
- Gold passage outside the top 10
- Generic or hedged answer where gold is specific
First fix: Add customer-phrased synonyms and regional terms to query expansion
- Retrieval mismatch350 open · 18%
Detection signals
- High similarity on the wrong article title
- Stale-version flag on top-ranked chunks
- Exception clause split from its rule across chunk boundaries
First fix: Filter or boost by effective date and locale metadata
- Instruction-following failures214 open · 11%
Detection signals
- Rule check fails: length, language, required closing line
- Promises that policy reserves for agents, such as guaranteed refunds
- Judge flags tone or format
First fix: Move hard constraints into automated rule checks
- Other78 open · 4%
Detection signals
- Refusal phrases on answerable items
- Language detector mismatch
- Output stops mid-sentence
First fix: Relax refusal thresholds for intents with strong retrieval
Root-cause clusters
Failures that share a cause are grouped around a representative question, so one fix can close many cases.
Refund promises
4 failures“Can I get a refund for the roaming pack I didn't use?”
UnsupportedInstructionsPlan change proration
3 failures“I upgraded from Essential to Unlimited on the 18th. Why is my bill €41.20 instead of €34.90?”
UnsupportedIncorrectMismatchPorting direction confusion
3 failures“How long does it take to port my number from another operator?”
IncorrectMismatchSIM swap and fraud
3 failures“Someone swapped my SIM without my permission. How fast can you block it?”
IncorrectUnsupported
04 · Improve
Compare, fix the ground truth, and block regressions
Once you know why answers fail, change one thing, re-run the same dataset and read the difference. Wrong gold answers are corrected in review, and thresholds in CI keep fixed problems from coming back.
Compare before you ship
Every change is scored against the last release on the same dataset with the same judge. Gains and regressions are marked the same way.
support-assistant v3.3 to v3.4
support-golden-v4 · 320-token semantic chunks · support-prompt v7
| Metric | v3.3 | v3.4 | Change |
|---|---|---|---|
| Pass rate | 81.2% | 86.9% | +5.7 pts(improvement) |
| Faithfulness | 0.87 | 0.91 | +0.04(improvement) |
| Answer correctness | 0.84 | 0.88 | +0.04(improvement) |
| Context recall | 0.79 | 0.84 | +0.05(improvement) |
| MRR | 0.69 | 0.74 | +0.05(improvement) |
| Latency p95 | 3.7s | 3.8s | +0.1s(regression) |
| Cost per request | $0.0046 | $0.0042 | −$0.0004(improvement) |
Fix the ground truth
Some failures are not the model's fault: the gold answer is wrong or out of date. Those go to the review queue, where approved and edited answers feed the next version of the dataset.
- Awaiting review
- 412
- Median review time
- 2.7 min
Gate releases in CI
Thresholds live next to your code. The workflow runs your configuration on every pull request that touches prompts, retrieval or config, and fails the build when a threshold is missed. The CLI installs with pip install relevant and also runs as npx @relevant/cli in Node projects, with the same commands.
# .github/workflows/relevant-gate.yml
name: relevant-gate
on:
pull_request:
paths: ["prompts/**", "retrieval/**", "relevant.yaml"]
jobs:
evaluate:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.12"
- run: pip install relevant
- name: Run regression gate
env:
RELEVANT_API_KEY: ${{ secrets.RELEVANT_API_KEY }}
run: >
relevant gate
--config relevant.yaml
--gates relevant.gates.yaml
--baseline release/v3.3
--fail-on-regression
--comment
Regression gate · support-assistant v3.4
run_2026-10-06_v3-4-rc vs production baseline
- Passed: Pass rate ≥ 85.0%86.9%
- Passed: Faithfulness ≥ 0.900.91
- Passed: Answer correctness ≥ 0.850.88
- Passed: Context recall ≥ 0.800.84
- Passed: Instruction adherence ≥ 0.920.93
- Passed: p95 latency ≤ 4.0s3.8s
- Passed: No new critical failures in account-security0 new
FAQ
Questions about the workflow
Run the loop on your own data.
Upload a dataset or log a few traces, and see your first failure breakdown in the same session.
Free plan includes 2,000 evaluated samples a month. No credit card.
The loop, end to end
- Log traces or upload CSV or JSONL
- Evaluate with a rubric and judges
- Diagnose each failure to its chunk
- Gate your first release in CI