run_2026-10-06_gpt-5
v3.4 · GPT-5 comparison
GPT-5 · bge-m3 · hybrid (BM25 + dense, RRF) · support-prompt v7 on support-golden-v4 v4.6
Pass rate
85.3%
Faithfulness
0.90
Answer correctness
0.87
Context precision
0.79
Context recall
0.84
Latency p95
5.2s
Failures by category
182 failed samples · largest share: unsupported
| Category | Share | vs reference |
|---|---|---|
| Unsupported claims | 22.0% | −5.2 pts |
| Incorrect factual answers | 20.3% | −0.7 pts |
| Missing context | 19.8% | +0.6 pts |
| Retrieval mismatch | 19.2% | +1.3 pts |
| Instruction-following failures | 14.8% | +3.7 pts |
| Other | 3.8% | +0.1 pts |
Run details
- Dataset
- support-golden-v4 v4.6 · 1,240 items
- Generator
- GPT-5
- Retriever
- bge-m3 · hybrid (BM25 + dense, RRF)
- Prompt
- support-prompt v7
- Judge
- LLM judge · Claude Opus 5.5 · rubric v3
- Started
- Oct 6, 14:05 UTC
- Duration
- 3h 8m
- Triggered by
- Maya Collins
- Tokens / request
- 3,420 in · 251 out
- Compared with
- support-assistant v3.4 · release candidate
- comparison
All metrics
Against support-assistant v3.4 · release candidate
| Metric | Group | This run | Reference | Change |
|---|---|---|---|---|
| Pass rateShare of samples the judge marked as passing all rubric checks. | Outcome | 85.3% | 86.9% | −1.6 pts |
| FaithfulnessShare of answer claims supported by the retrieved context. | Generation | 0.90 | 0.91 | −0.01 |
| Answer correctnessAgreement with the gold answer on facts, not style. | Generation | 0.87 | 0.88 | −0.01 |
| Instruction adherenceRespect for tone, format, length, language and policy constraints. | Generation | 0.90 | 0.93 | −0.03 |
| Context precisionRank-weighted precision of relevant chunks passed to the generator. | Retrieval | 0.79 | 0.79 | 0.00 |
| Context recallShare of gold-answer statements supported by the retrieved context. | Retrieval | 0.84 | 0.84 | 0.00 |
| Recall@5Share of gold passages found in the top 5 retrieved chunks. | Retrieval | 0.88 | 0.88 | 0.00 |
| MRRMean reciprocal rank of the first relevant chunk. | Retrieval | 0.74 | 0.74 | 0.00 |
| NDCG@10Graded ranking quality of the top 10, normalized by the ideal order. | Retrieval | 0.81 | 0.81 | 0.00 |
| Latency p50Median end-to-end time to final token. | Operational | 2.6s | 1.9s | +0.7s |
| Latency p9595th percentile end-to-end time to final token. | Operational | 5.2s | 3.8s | +1.4s |
| Cost per requestModel, embedding and reranking cost per answered request (USD). | Operational | $0.0061 | $0.0042 | +0.0019 |
Nightly trend
Nightly runs on the support-golden-v4 slice, Jul 11 to Oct 8. Dashed line marks this run.
- Pass rate · nightly
- This run · 85.3%
Across datasets
v3.4 · GPT-5 on every dataset in the project. Rows marked “matrix” come from the comparison matrix rather than a dedicated run.
| Dataset | Items | Pass rate | Faith. | Ctx recall | Recall@5 | p95 | Source |
|---|---|---|---|---|---|---|---|
| support-golden-v4this run | 1,240 | 85.3% | 0.90 | 0.84 | 0.88 | 5.2s | This run |
| billing-edge-cases | 312 | 80.1% | 0.86 | 0.82 | 0.86 | 5.4s | matrix |
| roaming-and-travel | 188 | 79.3% | 0.87 | 0.79 | 0.83 | 5.3s | matrix |
| device-troubleshooting | 426 | 82.9% | 0.88 | 0.83 | 0.87 | 5.9s | matrix |
| account-security | 154 | 86.4% | 0.91 | 0.85 | 0.89 | 5.1s | matrix |
| locale-formatting | 96 | 79.2% | 0.85 | 0.80 | 0.84 | 5.5s | matrix |
| synthetic-hard-negatives | 540 | 74.8% | 0.85 | 0.73 | 0.77 | 5.4s | matrix |
Traced samples
17 samples with full traces (4 failing, 13 passing), stratified by domain. Open a row to inspect retrieval and the judge's reasoning.
How long does it take to port my number to Harbor? Number porting Passjudge 0.88 |
Please answer in two sentences: when is my bill due? Billing & payments FailInstructionsjudge 0.62 |
My mobile data isn't working after updating to iOS 26.1. Device troubleshooting FailUnsupportedjudge 0.52 |
My son is 15. Can he have his own line on my account? Plans & upgrades FailIncorrectjudge 0.10 |
Give me the steps as a numbered list, please: how do I reset my network settings on Android? Device troubleshooting FailInstructionsjudge 0.40 |
What is the excess on a theft insurance claim? Returns & warranty Passjudge 0.90 |
My bill says 41,20 €. Is that forty-one euros or four thousand? Billing & payments Passjudge 0.87 |
How do I get a copy of the personal data Harbor holds about me? Data privacy requests Passjudge 0.91 |
Can I return a phone I bought online if I opened the box? Returns & warranty Passjudge 0.88 |
My phone shows 4G but mobile data does not work after the iOS update. Device troubleshooting Passjudge 0.97 |