Skip to content

Regression gates for RAG in CI

Block merges that make answers worse, without blocking every merge. Thresholds, sample sizes and what to do when a gate fails.

Relevant teamEngineering

10 min read

On this page
  1. Pick absolute floors and relative deltas
  2. Size the gate dataset for signal, not coverage
  3. Make failures explainable in the pull request
  4. Handle flaky judges

Prompt and retrieval changes are code changes. They deserve the same protection as any other change that can break production. A regression gate runs a fixed evaluation on each pull request and fails the check when quality drops below an agreed threshold.

Pick absolute floors and relative deltas

Use two kinds of checks. Absolute floors catch catastrophic changes: pass rate must stay above 85%, faithfulness above 0.90. Relative deltas catch slow drift: no metric may drop more than 1.5 points below the last release.

yaml
# relevant.gates.yaml (baseline passed as --baseline release/v3.3)gates:  - metric: pass_rate    min: 0.85  - metric: faithfulness    min: 0.90  - metric: context_recall    min: 0.80  - metric: latency_p95_ms    max: 4000  - metric: all    max_drop_vs_baseline: 0.015

Size the gate dataset for signal, not coverage

A gate runs on every pull request, so it must be fast. A stratified slice of 300 to 400 items from your golden set detects a two-point drop in pass rate most of the time while finishing in under an hour.

Run the full golden set nightly and before releases. The gate is a smoke alarm, not an inspection.

Make failures explainable in the pull request

A red check with 'pass rate 83.1%' starts an argument. A red check with 'unsupported claims +14 on billing intents; top example F-2409' starts a fix. Link the check to the failing samples and their categories.

Handle flaky judges

LLM judges are not perfectly deterministic even at temperature zero. Pin the judge model and rubric version, cache judgments for unchanged samples, and require two consecutive failures before blocking on borderline deltas.

Share
XLinkedIn
All articles
    • Model comparison
    • Evaluation

    Choosing a generator with evidence

    Comparing models on the same retrieval, the same prompt and the same questions, then reading the per-category differences.

    7 min read

    • Evaluation
    • Retrieval

    Why aggregate scores hide retrieval failures

    A single accuracy number can rise while the answers your customers care about get worse. Here is how to break the number apart.

    9 min read

    • Multilingual
    • Evaluation

    Evaluating regional formatting in support answers

    Currency, dates, decimal separators and language coverage: the failure patterns that only show up when assistants answer customers across regions.

    8 min read

See why your retrieval fails

Run an evaluation on your own dataset and read every failure with its retrieved chunks, flagged claims and suggested fix.

Free plan includes 2,000 evaluated samples a month