Prompt and retrieval changes are code changes. They deserve the same protection as any other change that can break production. A regression gate runs a fixed evaluation on each pull request and fails the check when quality drops below an agreed threshold.
Pick absolute floors and relative deltas
Use two kinds of checks. Absolute floors catch catastrophic changes: pass rate must stay above 85%, faithfulness above 0.90. Relative deltas catch slow drift: no metric may drop more than 1.5 points below the last release.
# relevant.gates.yaml (baseline passed as --baseline release/v3.3)gates: - metric: pass_rate min: 0.85 - metric: faithfulness min: 0.90 - metric: context_recall min: 0.80 - metric: latency_p95_ms max: 4000 - metric: all max_drop_vs_baseline: 0.015Size the gate dataset for signal, not coverage
A gate runs on every pull request, so it must be fast. A stratified slice of 300 to 400 items from your golden set detects a two-point drop in pass rate most of the time while finishing in under an hour.
Run the full golden set nightly and before releases. The gate is a smoke alarm, not an inspection.
Make failures explainable in the pull request
A red check with 'pass rate 83.1%' starts an argument. A red check with 'unsupported claims +14 on billing intents; top example F-2409' starts a fix. Link the check to the failing samples and their categories.
Handle flaky judges
LLM judges are not perfectly deterministic even at temperature zero. Pin the judge model and rubric version, cache judgments for unchanged samples, and require two consecutive failures before blocking on borderline deltas.