Model comparisons are easy to run badly. Change the retriever, the prompt and the model together and you learn nothing about the model. Hold everything else fixed, run the same questions, and compare per category, per dataset and per sample.
Fix everything except the generator
Use one retrieval configuration and one prompt version for every model in the comparison. Retrieval metrics such as Recall@5 and MRR should then be identical across runs, which is a useful sanity check.
Read per-category deltas
Two models with similar pass rates can fail very differently. One may produce fewer unsupported claims but break more format rules; another may handle regional formatting well and invent fees in billing answers. Per-category counts tell you which risks you are trading.
Include cost and latency in the decision
A smaller model that is two points worse at a third of the cost may be the right choice for high-volume intents, with a larger model reserved for billing and security questions. Plot quality against cost per request and decide per intent, not globally.
Look at the samples
Aggregate deltas are a starting point. Read at least twenty side-by-side samples where the models disagree before deciding. The patterns you see there become the next round of regression tests.