Skip to content

run_2026-10-06_gpt-5

v3.4 · GPT-5 comparison

GPT-5 · bge-m3 · hybrid (BM25 + dense, RRF) · support-prompt v7 on support-golden-v4 v4.6

CompletedExperimentv3.4 · GPT-5v3 · 320 tokens, semantic
Compare

Pass rate

85.3%

−1.6 pts(regression)vs support-assistant v3.4 · release candidate1,058 of 1,240 samples passed

Faithfulness

0.90

−0.01(regression)vs support-assistant v3.4 · release candidate

Answer correctness

0.87

−0.01(regression)vs support-assistant v3.4 · release candidate

Context precision

0.79

0.00(no change)vs support-assistant v3.4 · release candidate

Context recall

0.84

0.00(no change)vs support-assistant v3.4 · release candidate

Latency p95

5.2s

+1.4s(regression)vs support-assistant v3.4 · release candidate

Failures by category

182 failed samples · largest share: unsupported

Share of failed samples by category, compared with support-assistant v3.4 · release candidate
CategorySharevs reference
Unsupported claims22.0%−5.2 pts
Incorrect factual answers20.3%−0.7 pts
Missing context19.8%+0.6 pts
Retrieval mismatch19.2%+1.3 pts
Instruction-following failures14.8%+3.7 pts
Other3.8%+0.1 pts

Run details

Dataset
support-golden-v4 v4.6 · 1,240 items
Generator
GPT-5
Retriever
bge-m3 · hybrid (BM25 + dense, RRF)
Prompt
support-prompt v7
Judge
LLM judge · Claude Opus 5.5 · rubric v3
Started
Oct 6, 14:05 UTC
Duration
3h 8m
Triggered by
Maya Collins
Tokens / request
3,420 in · 251 out
  • comparison

All metrics

Against support-assistant v3.4 · release candidate

All metrics for v3.4 · GPT-5 comparison
MetricThis runReferenceChange
Pass rate85.3%86.9%−1.6 pts
Faithfulness0.900.91−0.01
Answer correctness0.870.88−0.01
Instruction adherence0.900.93−0.03
Context precision0.790.790.00
Context recall0.840.840.00
Recall@50.880.880.00
MRR0.740.740.00
NDCG@100.810.810.00
Latency p502.6s1.9s+0.7s
Latency p955.2s3.8s+1.4s
Cost per request$0.0061$0.0042+0.0019

Nightly trend

Nightly runs on the support-golden-v4 slice, Jul 11 to Oct 8. Dashed line marks this run.

  • Pass rate · nightly
  • This run · 85.3%
Pass rate moved from 78.5% to 87.2% across 90 nightly runs. Configuration changes: v3.3 on Aug 4, v3.4 on Sep 8.

Across datasets

v3.4 · GPT-5 on every dataset in the project. Rows marked “matrix” come from the comparison matrix rather than a dedicated run.

Configuration results per dataset
DatasetItemsPass rateFaith.Ctx recallRecall@5p95Source
support-golden-v4this run1,240
85.3%
0.900.840.885.2sThis run
billing-edge-cases312
80.1%
0.860.820.865.4smatrix
roaming-and-travel188
79.3%
0.870.790.835.3smatrix
device-troubleshooting426
82.9%
0.880.830.875.9smatrix
account-security154
86.4%
0.910.850.895.1smatrix
locale-formatting96
79.2%
0.850.800.845.5smatrix
synthetic-hard-negatives540
74.8%
0.850.730.775.4smatrix

Traced samples

17 samples with full traces (4 failing, 13 passing), stratified by domain. Open a row to inspect retrieval and the judge's reasoning.

Traced samples for v3.4 · GPT-5 comparison

How long does it take to port my number to Harbor?

Number porting

Passjudge 0.88

Please answer in two sentences: when is my bill due?

Billing & payments

FailInstructionsjudge 0.62

My mobile data isn't working after updating to iOS 26.1.

Device troubleshooting

FailUnsupportedjudge 0.52

My son is 15. Can he have his own line on my account?

Plans & upgrades

FailIncorrectjudge 0.10

Give me the steps as a numbered list, please: how do I reset my network settings on Android?

Device troubleshooting

FailInstructionsjudge 0.40

What is the excess on a theft insurance claim?

Returns & warranty

Passjudge 0.90

My bill says 41,20 €. Is that forty-one euros or four thousand?

Billing & payments

Passjudge 0.87

How do I get a copy of the personal data Harbor holds about me?

Data privacy requests

Passjudge 0.91

Can I return a phone I bought online if I opened the box?

Returns & warranty

Passjudge 0.88

My phone shows 4G but mobile data does not work after the iOS update.

Device troubleshooting

Passjudge 0.97