Skip to content

run_2026-10-06_gemini-2-5-pro

v3.4 · Gemini 2.5 Pro comparison

Gemini 2.5 Pro · bge-m3 · hybrid (BM25 + dense, RRF) · support-prompt v7 on support-golden-v4 v4.6

CompletedExperimentv3.4 · Gemini 2.5 Prov3 · 320 tokens, semantic
Compare

Pass rate

83.8%

−3.1 pts(regression)vs support-assistant v3.4 · release candidate1,039 of 1,240 samples passed

Faithfulness

0.88

−0.03(regression)vs support-assistant v3.4 · release candidate

Answer correctness

0.86

−0.02(regression)vs support-assistant v3.4 · release candidate

Context precision

0.79

0.00(no change)vs support-assistant v3.4 · release candidate

Context recall

0.84

0.00(no change)vs support-assistant v3.4 · release candidate

Latency p95

4.6s

+0.8s(regression)vs support-assistant v3.4 · release candidate

Failures by category

201 failed samples · largest share: unsupported

Share of failed samples by category, compared with support-assistant v3.4 · release candidate
CategorySharevs reference
Unsupported claims25.4%−1.8 pts
Incorrect factual answers22.9%+1.9 pts
Missing context18.9%−0.2 pts
Retrieval mismatch17.9%0.0 pts
Instruction-following failures10.9%−0.2 pts
Other4.0%+0.3 pts

Run details

Dataset
support-golden-v4 v4.6 · 1,240 items
Generator
Gemini 2.5 Pro
Retriever
bge-m3 · hybrid (BM25 + dense, RRF)
Prompt
support-prompt v7
Judge
LLM judge · Claude Opus 5.5 · rubric v3
Started
Oct 6, 14:06 UTC
Duration
2h 54m
Triggered by
Maya Collins
Tokens / request
3,420 in · 236 out
  • comparison

All metrics

Against support-assistant v3.4 · release candidate

All metrics for v3.4 · Gemini 2.5 Pro comparison
MetricThis runReferenceChange
Pass rate83.8%86.9%−3.1 pts
Faithfulness0.880.91−0.03
Answer correctness0.860.88−0.02
Instruction adherence0.910.93−0.02
Context precision0.790.790.00
Context recall0.840.840.00
Recall@50.880.880.00
MRR0.740.740.00
NDCG@100.810.810.00
Latency p502.3s1.9s+0.4s
Latency p954.6s3.8s+0.8s
Cost per request$0.0048$0.0042+0.0006

Nightly trend

Nightly runs on the support-golden-v4 slice, Jul 11 to Oct 8. Dashed line marks this run.

  • Pass rate · nightly
  • This run · 83.8%
Pass rate moved from 78.5% to 87.2% across 90 nightly runs. Configuration changes: v3.3 on Aug 4, v3.4 on Sep 8.

Across datasets

v3.4 · Gemini 2.5 Pro on every dataset in the project. Rows marked “matrix” come from the comparison matrix rather than a dedicated run.

Configuration results per dataset
DatasetItemsPass rateFaith.Ctx recallRecall@5p95Source
support-golden-v4this run1,240
83.8%
0.880.840.884.6sThis run
billing-edge-cases312
78.8%
0.840.820.864.8smatrix
roaming-and-travel188
78.2%
0.850.790.834.7smatrix
device-troubleshooting426
81.5%
0.860.830.875.3smatrix
account-security154
85.1%
0.890.850.894.5smatrix
locale-formatting96
78.1%
0.840.800.844.9smatrix
synthetic-hard-negatives540
73.3%
0.830.730.774.8smatrix

Traced samples

18 samples with full traces (5 failing, 13 passing), stratified by domain. Open a row to inspect retrieval and the judge's reasoning.

Traced samples for v3.4 · Gemini 2.5 Pro comparison

My eSIM stopped working after I reset my phone. Do I need to pay for a new one?

eSIM & activation

FailUnsupportedjudge 0.12

Why was I charged twice for my September bill?

Billing & payments

Passjudge 0.97

If I upgrade my plan mid-month, how is my next bill calculated?

Plans & upgrades

Passjudge 0.89

How much data can I use in the EU on the Unlimited Plus plan?

Roaming & travel

Passjudge 0.95

My eSIM QR code says it has already been used. What should I do?

eSIM & activation

Passjudge 0.94

What is the late fee if I pay my bill after the due date?

Billing & payments

Passjudge 0.90

What is the excess on a theft insurance claim?

Returns & warranty

Passjudge 0.94

Is SIM swap protection turned on by default?

Account & security

Passjudge 0.90

My phone shows 4G but mobile data does not work after the iOS update.

Device troubleshooting

Passjudge 0.96

Can I return a phone I bought online if I opened the box?

Returns & warranty

Passjudge 0.94