Skip to content

run_2026-10-07_haiku-5-5

v3.4 · Claude Haiku 5.5 comparison

Claude Haiku 5.5 · bge-m3 · hybrid (BM25 + dense, RRF) · support-prompt v7 on support-golden-v4 v4.6

FailedExperimentv3.4 · Claude Haiku 5.5v3 · 320 tokens, semantic
Compare

Pass rate

82.3%

−4.6 pts(regression)vs support-assistant v3.4 · release candidate344 of 418 samples passed

Faithfulness

0.89

−0.02(regression)vs support-assistant v3.4 · release candidate

Answer correctness

0.83

−0.05(regression)vs support-assistant v3.4 · release candidate

Context precision

0.79

0.00(no change)vs support-assistant v3.4 · release candidate

Context recall

0.84

0.00(no change)vs support-assistant v3.4 · release candidate

Latency p95

2.3s

−1.5s(improvement)vs support-assistant v3.4 · release candidate

Failures by category

74 failed samples · largest share: unsupported

Share of failed samples by category, compared with support-assistant v3.4 · release candidate
CategorySharevs reference
Unsupported claims28.4%+1.2 pts
Incorrect factual answers23.0%+2.0 pts
Missing context17.6%−1.6 pts
Retrieval mismatch17.6%−0.3 pts
Instruction-following failures9.5%−1.7 pts
Other4.1%+0.4 pts

Run details

Dataset
support-golden-v4 v4.6 · 1,240 items
Generator
Claude Haiku 5.5
Retriever
bge-m3 · hybrid (BM25 + dense, RRF)
Prompt
support-prompt v7
Judge
LLM judge · Claude Opus 5.5 · rubric v3
Started
Oct 7, 07:40 UTC
Duration
1h 41m
Triggered by
Sara Lindqvist
Tokens / request
3,420 in · 198 out
  • cost
  • comparison

All metrics

Against support-assistant v3.4 · release candidate

All metrics for v3.4 · Claude Haiku 5.5 comparison
MetricThis runReferenceChange
Pass rate82.3%86.9%−4.6 pts
Faithfulness0.890.91−0.02
Answer correctness0.830.88−0.05
Instruction adherence0.910.93−0.02
Context precision0.790.790.00
Context recall0.840.840.00
Recall@50.880.880.00
MRR0.740.740.00
NDCG@100.810.810.00
Latency p501.1s1.9s−0.8s
Latency p952.3s3.8s−1.5s
Cost per request$0.0013$0.0042−0.0029

Nightly trend

Nightly runs on the support-golden-v4 slice, Jul 11 to Oct 8. Dashed line marks this run.

  • Pass rate · nightly
  • This run · 82.3%
Pass rate moved from 78.5% to 87.2% across 90 nightly runs. Configuration changes: v3.3 on Aug 4, v3.4 on Sep 8.

Across datasets

v3.4 · Claude Haiku 5.5 on every dataset in the project. Rows marked “matrix” come from the comparison matrix rather than a dedicated run.

Configuration results per dataset
DatasetItemsPass rateFaith.Ctx recallRecall@5p95Source
support-golden-v4this run1,240
82.3%
0.890.840.882.3smatrix
billing-edge-cases312
77.2%
0.850.820.862.5smatrix
roaming-and-travel188
76.6%
0.860.790.832.4smatrix
device-troubleshooting426
79.8%
0.870.830.873.0smatrix
account-security154
83.8%
0.900.850.892.2smatrix
locale-formatting96
75.0%
0.830.800.842.6smatrix
synthetic-hard-negatives540
71.9%
0.840.730.772.5smatrix

Traced samples

17 samples with full traces (4 failing, 13 passing), stratified by domain. Open a row to inspect retrieval and the judge's reasoning.

Traced samples for v3.4 · Claude Haiku 5.5 comparison

What happens if I go over my fair-use limit on Unlimited at home?

Plans & upgrades

FailMissingjudge 0.07

Does the Fall 2026 upgrade offer work with a family plan discount?

Promotions

Passjudge 0.95

What is the late fee if I pay my bill after the due date?

Billing & payments

Passjudge 0.98

How long does it take to port my number to Harbor?

Number porting

Passjudge 0.90

How do I get a copy of the personal data Harbor holds about me?

Data privacy requests

Passjudge 0.89

My phone shows 4G but mobile data does not work after the iOS update.

Device troubleshooting

Passjudge 0.97

Is SIM swap protection turned on by default?

Account & security

Passjudge 0.97

What is the excess on a theft insurance claim?

Returns & warranty

Passjudge 0.87

My eSIM QR code says it has already been used. What should I do?

eSIM & activation

Passjudge 0.89

My bill says 41,20 €. Is that forty-one euros or four thousand?

Billing & payments

Passjudge 0.93