Skip to content

Choosing a generator with evidence

Comparing models on the same retrieval, the same prompt and the same questions, then reading the per-category differences.

Relevant teamEngineering

7 min read

On this page
  1. Fix everything except the generator
  2. Read per-category deltas
  3. Include cost and latency in the decision
  4. Look at the samples

Model comparisons are easy to run badly. Change the retriever, the prompt and the model together and you learn nothing about the model. Hold everything else fixed, run the same questions, and compare per category, per dataset and per sample.

Fix everything except the generator

Use one retrieval configuration and one prompt version for every model in the comparison. Retrieval metrics such as Recall@5 and MRR should then be identical across runs, which is a useful sanity check.

Read per-category deltas

Two models with similar pass rates can fail very differently. One may produce fewer unsupported claims but break more format rules; another may handle regional formatting well and invent fees in billing answers. Per-category counts tell you which risks you are trading.

Include cost and latency in the decision

A smaller model that is two points worse at a third of the cost may be the right choice for high-volume intents, with a larger model reserved for billing and security questions. Plot quality against cost per request and decide per intent, not globally.

Look at the samples

Aggregate deltas are a starting point. Read at least twenty side-by-side samples where the models disagree before deciding. The patterns you see there become the next round of regression tests.

Share
XLinkedIn
All articles
    • Evaluation
    • Retrieval

    Why aggregate scores hide retrieval failures

    A single accuracy number can rise while the answers your customers care about get worse. Here is how to break the number apart.

    9 min read

    • CI
    • Evaluation

    Regression gates for RAG in CI

    Block merges that make answers worse, without blocking every merge. Thresholds, sample sizes and what to do when a gate fails.

    10 min read

    • Multilingual
    • Evaluation

    Evaluating regional formatting in support answers

    Currency, dates, decimal separators and language coverage: the failure patterns that only show up when assistants answer customers across regions.

    8 min read

See why your retrieval fails

Run an evaluation on your own dataset and read every failure with its retrieved chunks, flagged claims and suggested fix.

Free plan includes 2,000 evaluated samples a month