Skip to content

run_2026-09-27_nightly

Nightly · v3.4 · 2026-09-27

Claude Sonnet 5.5 · bge-m3 · hybrid (BM25 + dense, RRF) · support-prompt v7 on support-golden-v4 v4.6

CompletedNightlyProduction candidatesupport-assistant v3.4v3 · 320 tokens, semantic
Compare

Stratified slice of support-golden-v4 (420 items).

Pass rate

86.7%

−1.4 pts(regression)vs previous nightly364 of 420 samples passed

Faithfulness

0.92

+0.02(improvement)vs previous nightly

Answer correctness

0.89

0.00(no change)vs previous nightly

Context precision

0.78

0.00(no change)vs previous nightly

Context recall

0.85

0.00(no change)vs previous nightly

Latency p95

3.6s

0.0s(no change)vs previous nightly

Failures by category

56 failed samples · largest share: unsupported

Share of failed samples by category, compared with Nightly · v3.4 · 2026-09-26
CategorySharevs reference
Unsupported claims26.8%−1.2 pts
Incorrect factual answers21.4%−0.6 pts
Missing context19.6%+1.6 pts
Retrieval mismatch17.9%−0.1 pts
Instruction-following failures10.7%+0.7 pts
Other3.6%−0.4 pts

Run details

Dataset
support-golden-v4 v4.6 · 420 items
Generator
Claude Sonnet 5.5
Retriever
bge-m3 · hybrid (BM25 + dense, RRF)
Prompt
support-prompt v7
Judge
LLM judge · Claude Opus 5.5 · rubric v3
Started
Sep 27, 02:00 UTC
Duration
43m 27s
Triggered by
Scheduler
Tokens / request
3,420 in · 214 out
  • nightly
  • golden-slice

All metrics

Against Nightly · v3.4 · 2026-09-26

All metrics for Nightly · v3.4 · 2026-09-27
MetricThis runReferenceChange
Pass rate86.7%88.1%−1.4 pts
Faithfulness0.920.90+0.02
Answer correctness0.890.890.00
Instruction adherence0.930.930.00
Context precision0.780.780.00
Context recall0.850.850.00
Recall@50.890.87+0.01
MRR0.750.74+0.01
NDCG@100.800.82−0.01
Latency p502.0s1.9s+0.1s
Latency p953.6s3.6s0.0s
Cost per request$0.0041$0.00410.0000

Nightly trend

Nightly runs on the support-golden-v4 slice, Jul 11 to Oct 8. This run is one point in the series.

Pass rate moved from 78.5% to 87.2% across 90 nightly runs. Configuration changes: v3.3 on Aug 4, v3.4 on Sep 8.

Across datasets

support-assistant v3.4 on every dataset in the project. Rows marked “matrix” come from the comparison matrix rather than a dedicated run.

Configuration results per dataset
DatasetItemsPass rateFaith.Ctx recallRecall@5p95Source
support-golden-v4this run1,240
86.7%
0.920.850.893.6sThis run
billing-edge-cases312
81.7%
0.870.820.864.0sRun · Oct 2
roaming-and-travel188
80.9%
0.880.790.833.9sRun · Sep 28
device-troubleshooting426
84.5%
0.890.830.874.5sRun · Sep 30
account-security154
88.3%
0.920.850.893.7sRun · Sep 25
locale-formatting96
80.2%
0.860.800.844.1sRun · Oct 4
synthetic-hard-negatives540
76.5%
0.860.730.774.0smatrix

Traced samples

14 samples with full traces (1 failing, 13 passing), stratified by domain. Open a row to inspect retrieval and the judge's reasoning.

Traced samples for Nightly · v3.4 · 2026-09-27

My eSIM QR code says it has already been used. What should I do?

eSIM & activation

Passjudge 0.95

Is SIM swap protection turned on by default?

Account & security

Passjudge 0.88

Can I return a phone I bought online if I opened the box?

Returns & warranty

Passjudge 0.97

My bill says 41,20 €. Is that forty-one euros or four thousand?

Billing & payments

Passjudge 0.93

Why was I charged twice for my September bill?

Billing & payments

Passjudge 0.94

If I upgrade my plan mid-month, how is my next bill calculated?

Plans & upgrades

Passjudge 0.97

How much data can I use in the EU on the Unlimited Plus plan?

Roaming & travel

Passjudge 0.89

What is the late fee if I pay my bill after the due date?

Billing & payments

Passjudge 0.97

My phone shows 4G but mobile data does not work after the iOS update.

Device troubleshooting

Passjudge 0.91

How do I get a copy of the personal data Harbor holds about me?

Data privacy requests

Passjudge 0.86