Skip to content

How it works

From raw traces to a release you can defend

Relevant runs one loop. Ingest your questions and traces, evaluate every answer, diagnose each failure to its cause, and improve with evidence that holds up in review.

The data model

Five objects, traced end to end

Every result in Relevant resolves to the same chain, so a failing verdict always leads back to the chunk and the dataset item behind it.

  1. 01

    Dataset

    A versioned set of questions with gold answers and source passages.

    support-golden-v4 · v4.6 · 1,240 items

  2. 02

    Query

    One question, replayed through your pipeline with a fixed configuration.

    "How much data can I use in Germany on…"

  3. 03

    Retrieved set

    Ranked chunks with scores, article versions and a relevance mark.

    Gold passage at rank 2

  4. 04

    Answer

    The generated response, split into claims and citations.

    2 claims · 1 unsupported

  5. 05

    Verdict

    Judge scores, pass or fail, and one failure category.

    Fail · Retrieval mismatch

Trace F-2412, run_2026-10-08_nightly.

01 · Ingest

Bring the questions and traces you already have

Start from a dataset of questions with gold answers, from logged production traces, or both. Relevant validates the data on the way in, so the first run starts from clean input.

Log traces from your pipeline

Send one trace per request, after the retrieval and generation calls you already make. Each trace stores the ranked chunks, scores, article versions and the final answer, which is what chunk-level diagnosis needs.

# pip install relevant
from relevant import Relevant

client = Relevant()  # reads RELEVANT_API_KEY (keys start with rlv_)

chunks = retriever.search(question, k=8)
answer = generator.answer(question, chunks)

client.log_trace(
    project="support-assistant",
    dataset="support-golden-v4",
    question=question,
    retrieved=[
        {
            "id": c.id,
            "article": c.article_id,
            "section": c.section,
            "score": c.score,
            "version": c.version,
        }
        for c in chunks
    ],
    answer={
        "text": answer.text,
        "model": "claude-sonnet-5-5",
        "prompt": "support-prompt@v7",
    },
)

Upload a dataset

Use CSV or JSONL with one row per question. Only question and gold_answer are required. Adding source_article and source_section turns on chunk-level attribution for that row.

id,question,gold_answer,domain,language,source_article,source_section,difficulty
gt_1008,How long does it take to port my number to Harbor?,"Once we receive a valid porting code, number transfers complete within 1 business day. Keep your old SIM active until your Harbor SIM shows signal; your old provider closes the line automatically.",porting,en,kb-porting-002,Timeline,easy
gt_1005,Is SIM swap protection turned on by default?,Yes. SIM swap protection is on for every line. A SIM or eSIM swap requires a one-time code sent to the account holder's verified email and photo ID in store. You can add a 24-hour hold on swaps in My Harbor under Security.,account,en,kb-account-005,How SIM swap protection works,medium
gt_1002,"If I upgrade my plan mid-month, how is my next bill calculated?","Upgrades take effect immediately. Your next bill includes a pro-rated charge for the new plan from the upgrade date to the end of the cycle and a credit for the unused days of the old plan, followed by the full monthly price of the new plan.",plans,en,kb-plans-008,Upgrades,medium

UTF-8, one row per question. Rows in different languages can share a file.

Dataset file schema
FieldTypeRequiredNotes
idstringOptionalStable row identifier. Generated when omitted.
questionstringRequiredThe customer question as it was asked.
gold_answerstringRequiredThe verified reference answer a reviewer would approve.
domainstringOptionalTopic label for per-domain breakdowns, for example roaming or billing.
languagestringOptionalLanguage code of the question, for example en. Detected from the question when omitted.
source_articlestringOptionalKnowledge-base article id. Enables chunk-level attribution.
source_sectionstringOptionalHeading of the section that supports the answer.
difficultyeasy | medium | hardOptionalUsed to filter results and weight reports.
Push from the command lineShell
relevant datasets push ./support-golden-v4.jsonl \
  --name support-golden-v4 \
  --version v4.6 \
  --validate

What happens on import

Every file is checked before anything is written, and the results are shown before you confirm the import.

  • Schema and encoding

    Rows with a missing question or gold answer are reported by line number before anything is imported.

  • Near-duplicate detection

    Questions that differ only in wording are flagged, so one case is not counted twice.

  • Source mapping coverage

    The share of rows linked to an article and section is shown, because attribution depends on it.

02 · Evaluate

Score every answer with judges you can audit

A run replays each dataset question through a pinned configuration and scores the result with the judges you choose. The configuration is a file in your repository, so a run can always be reproduced.

Describe the run in one file

The file pins the dataset version, generator, retriever, chunking, prompt and judges. Change one line, run again and compare.

relevant.yamlYAML
# relevant.yaml
project: support-assistant
dataset: support-golden-v4@v4.6

pipeline:
  generator: claude-sonnet-5-5
  retriever: bge-m3-hybrid      # BM25 + dense, reciprocal rank fusion
  chunking: semantic-320
  prompt: support-prompt@v7
  top_k: 8

judges:
  - type: llm
    model: claude-opus-5-5
    rubric: rubric-v3
    temperature: 0
  - type: rules
    checks: [language_match, max_words_120, citation_ids_exist]
  - type: human
    share: 0.10                 # double-blind, adjudicated on disagreement

metrics:
  - pass_rate
  - faithfulness
  - answer_correctness
  - context_precision
  - context_recall
  - recall@5
  - mrr
  - ndcg@10
Run itShell
relevant evaluate --config relevant.yaml --name "v3.4 release candidate"

A rubric you can read

Rubrics are versioned and stored with the run. Each criterion is a plain-language question, a scale and a threshold, so a judge's verdict can be argued with.

Support answer rubric

rubric-v3 · judge: Claude Opus 5.5 · temperature 0

Rubric criteria, scale and failure thresholds
CriterionQuestion put to the judgeScaleAnswer fails when
FaithfulnessgroundedIs every factual claim supported by the retrieved context?Pass or fail, from a 0 to 1 claim-level scoreBelow 0.85
Answer correctnesscorrectDo the facts (plan, fee, date, zone) match the gold answer?Pass or fail, from a 0 to 1 scoreBelow 0.70
CompletenesscompleteDoes the answer cover every point in the gold answer?1 to 5Below 4
Instruction adherencepolicyDoes the answer respect tone, language, length and escalation rules?Pass or fail, rule checks plus judgeAny hard rule fails

An answer passes only when every criterion passes. The judge returns a written rationale with each verdict, so every failure can be audited.

Choose your judges

Combine judges on the same run. Their verdicts are shown side by side for every answer.

  • Default

    LLM judge

    Claude Opus 5.5 scores each answer against rubric v3 at temperature 0, decomposes answers into claims, and stores a written rationale with every verdict.

  • Team and Enterprise

    Human review

    Your reviewers score a double-blind share of answers next to the LLM judge. Disagreements go to adjudication and feed the ground-truth review queue.

  • Deterministic

    Rule checks

    Language match, length limits, required phrases, citation ids that exist and PII patterns. Cheap, fast and identical on every run.

  • Pro and above

    Your own judge

    Point Relevant at any judge model endpoint you control and attach your own rubric. Judge model usage is billed by your provider.

Metrics every run reports

Retrieval and generation are measured separately, so a drop in one is never blamed on the other.

Outcome

Pass ratehigher is better
Share of samples the judge marked as passing all rubric checks.

Retrieval

Context precisionhigher is better
Rank-weighted precision of relevant chunks passed to the generator.
Context recallhigher is better
Share of gold-answer statements supported by the retrieved context.
Recall@5higher is better
Share of gold passages found in the top 5 retrieved chunks.
MRRhigher is better
Mean reciprocal rank of the first relevant chunk.
NDCG@10higher is better
Graded ranking quality of the top 10, normalized by the ideal order.

Generation

Faithfulnesshigher is better
Share of answer claims supported by the retrieved context.
Answer correctnesshigher is better
Agreement with the gold answer on facts, not style.
Instruction adherencehigher is better
Respect for tone, format, length, language and policy constraints.

Operational

Latency p50lower is better
Median end-to-end time to final token.
Latency p95lower is better
95th percentile end-to-end time to final token.
Cost per requestlower is better
Model, embedding and reranking cost per answered request (USD).

03 · Diagnose

Trace every failure to the chunk that caused it

Relevant does not stop at a score. Each failed answer is labeled with one category and linked to the retrieved chunks, so a regression points to a cause instead of an average.

Walk through a failure

Pick a failure to see the evidence the judges had and the staged checks that assigned its category.

Choose a failure

Showing F-2412: Superseded article ranks first. Assigned category: Retrieval mismatch.

F-2412

How much data can I use in Germany on Unlimited Plus before I get charged?

Retrieved chunks

  1. 1Roaming fair-use (2025 edition)Fair-use allowance · v2025.11 · superseded0.83Not relevant
  2. 2Roaming fair-useFair-use allowance · v2026.090.81Relevant
  3. 3Roaming zones and countriesZone 1: EU and EEA · v2026.090.74Not relevant

Model answer

Unlimited Plus includes Flagged by judge: 25 GB of roaming data per month in the EU kb-roaming-011. After that, data is charged at €0.0024 per MB.

Staged checks

Stages run in order. The first failing stage decides the category.

  1. 1 RetrievalFailed

    Was the gold passage retrieved and ranked first among the relevant chunks?

    The gold passage is at rank 2. A superseded 2025 edition of the article ranks first, and the generator used its 25 GB figure.

  2. 2 GroundingNot needed

    Is every claim in the answer supported by the retrieved context?

  3. 3 CorrectnessNot needed

    Does the answer agree with the gold answer on every key fact?

  4. 4 InstructionsNot needed

    Were the language, format and policy rules followed?

Assigned category

Retrieval mismatch

Root cause: The 2025 edition of the fair-use article is still indexed and outranks the current version by 0.02.

Suggested fix: Filter retrieval to effective-dated articles (version >= v2026.09) and remove kb-roaming-011-2025 from the index.

How categories are assigned

The signals below are computed for every failed answer. The first stage that fails decides the category; the rest are recorded but do not change it.

  • Unsupported claims525 open · 27%

    Detection signals

    • Claim-level faithfulness below 0.85
    • Numbers, fees or durations absent from every retrieved chunk
    • Citations that do not contain the cited sentence

    First fix: Require a citation after each factual sentence and refuse when none applies

  • Incorrect factual answers408 open · 21%

    Detection signals

    • Answer correctness below 0.7 while faithfulness stays high
    • Disagreement with gold on a key entity: plan, fee, date or zone
    • Reviewer edits the fact but keeps the structure

    First fix: Check for a stale article version before changing the prompt

  • Missing context369 open · 19%

    Detection signals

    • Context recall below 0.7
    • Gold passage outside the top 10
    • Generic or hedged answer where gold is specific

    First fix: Add customer-phrased synonyms and regional terms to query expansion

  • Retrieval mismatch350 open · 18%

    Detection signals

    • High similarity on the wrong article title
    • Stale-version flag on top-ranked chunks
    • Exception clause split from its rule across chunk boundaries

    First fix: Filter or boost by effective date and locale metadata

  • Instruction-following failures214 open · 11%

    Detection signals

    • Rule check fails: length, language, required closing line
    • Promises that policy reserves for agents, such as guaranteed refunds
    • Judge flags tone or format

    First fix: Move hard constraints into automated rule checks

  • Other78 open · 4%

    Detection signals

    • Refusal phrases on answerable items
    • Language detector mismatch
    • Output stops mid-sentence

    First fix: Relax refusal thresholds for intents with strong retrieval

Root-cause clusters

Failures that share a cause are grouped around a representative question, so one fix can close many cases.

  • Refund promises

    4 failures

    “Can I get a refund for the roaming pack I didn't use?”

    UnsupportedInstructions
  • Plan change proration

    3 failures

    “I upgraded from Essential to Unlimited on the 18th. Why is my bill €41.20 instead of €34.90?”

    UnsupportedIncorrectMismatch
  • Porting direction confusion

    3 failures

    “How long does it take to port my number from another operator?”

    IncorrectMismatch
  • SIM swap and fraud

    3 failures

    “Someone swapped my SIM without my permission. How fast can you block it?”

    IncorrectUnsupported

04 · Improve

Compare, fix the ground truth, and block regressions

Once you know why answers fail, change one thing, re-run the same dataset and read the difference. Wrong gold answers are corrected in review, and thresholds in CI keep fixed problems from coming back.

Compare before you ship

Every change is scored against the last release on the same dataset with the same judge. Gains and regressions are marked the same way.

support-assistant v3.3 to v3.4

support-golden-v4 · 320-token semantic chunks · support-prompt v7

Metric comparison between v3.3 and v3.4
Metricv3.3v3.4Change
Pass rate81.2%86.9%+5.7 pts(improvement)
Faithfulness0.870.91+0.04(improvement)
Answer correctness0.840.88+0.04(improvement)
Context recall0.790.84+0.05(improvement)
MRR0.690.74+0.05(improvement)
Latency p953.7s3.8s+0.1s(regression)
Cost per request$0.0046$0.0042−$0.0004(improvement)
Open model comparison

Fix the ground truth

Some failures are not the model's fault: the gold answer is wrong or out of date. Those go to the review queue, where approved and edited answers feed the next version of the dataset.

Awaiting review
412
Median review time
2.7 min
Open the review queue

Gate releases in CI

Thresholds live next to your code. The workflow runs your configuration on every pull request that touches prompts, retrieval or config, and fails the build when a threshold is missed. The CLI installs with pip install relevant and also runs as npx @relevant/cli in Node projects, with the same commands.

# .github/workflows/relevant-gate.yml
name: relevant-gate

on:
  pull_request:
    paths: ["prompts/**", "retrieval/**", "relevant.yaml"]

jobs:
  evaluate:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with:
          python-version: "3.12"
      - run: pip install relevant
      - name: Run regression gate
        env:
          RELEVANT_API_KEY: ${{ secrets.RELEVANT_API_KEY }}
        run: >
          relevant gate
          --config relevant.yaml
          --gates relevant.gates.yaml
          --baseline release/v3.3
          --fail-on-regression
          --comment

Regression gate · support-assistant v3.4

run_2026-10-06_v3-4-rc vs production baseline

7 of 7 checks passed
  • Passed: Pass rate ≥ 85.0%86.9%
  • Passed: Faithfulness ≥ 0.900.91
  • Passed: Answer correctness ≥ 0.850.88
  • Passed: Context recall ≥ 0.800.84
  • Passed: Instruction adherence ≥ 0.920.93
  • Passed: p95 latency ≤ 4.0s3.8s
  • Passed: No new critical failures in account-security0 new

FAQ

Questions about the workflow

Only to log traces. Call the SDK after your retrieval and generation steps, or import a JSONL export of existing traces. Datasets can be uploaded without touching any code.

Run the loop on your own data.

Upload a dataset or log a few traces, and see your first failure breakdown in the same session.

Free plan includes 2,000 evaluated samples a month. No credit card.

The loop, end to end

  • Log traces or upload CSV or JSONL
  • Evaluate with a rubric and judges
  • Diagnose each failure to its chunk
  • Gate your first release in CI