Skip to content

run_2026-09-15_nightly

Nightly · v3.4 · 2026-09-15

Claude Sonnet 5.5 · bge-m3 · hybrid (BM25 + dense, RRF) · support-prompt v7 on support-golden-v4 v4.6

CompletedNightlyProduction candidatesupport-assistant v3.4v3 · 320 tokens, semantic
Compare

Stratified slice of support-golden-v4 (419 items).

Pass rate

86.6%

−1.0 pts(regression)vs previous nightly363 of 419 samples passed

Faithfulness

0.91

0.00(no change)vs previous nightly

Answer correctness

0.87

−0.01(regression)vs previous nightly

Context precision

0.79

−0.01(regression)vs previous nightly

Context recall

0.83

−0.01(regression)vs previous nightly

Latency p95

3.8s

0.0s(no change)vs previous nightly

Failures by category

56 failed samples · largest share: unsupported

Share of failed samples by category, compared with Nightly · v3.4 · 2026-09-14
CategorySharevs reference
Unsupported claims26.8%−0.1 pts
Incorrect factual answers19.6%+0.4 pts
Missing context19.6%+0.4 pts
Retrieval mismatch17.9%+0.5 pts
Instruction-following failures10.7%−0.8 pts
Other5.4%−0.4 pts

Run details

Dataset
support-golden-v4 v4.6 · 419 items
Generator
Claude Sonnet 5.5
Retriever
bge-m3 · hybrid (BM25 + dense, RRF)
Prompt
support-prompt v7
Judge
LLM judge · Claude Opus 5.5 · rubric v3
Started
Sep 15, 02:00 UTC
Duration
42m 48s
Triggered by
Scheduler
Tokens / request
3,420 in · 214 out
  • nightly
  • golden-slice

All metrics

Against Nightly · v3.4 · 2026-09-14

All metrics for Nightly · v3.4 · 2026-09-15
MetricThis runReferenceChange
Pass rate86.6%87.6%−1.0 pts
Faithfulness0.910.910.00
Answer correctness0.870.88−0.01
Instruction adherence0.930.930.00
Context precision0.790.80−0.01
Context recall0.830.84−0.01
Recall@50.880.89−0.01
MRR0.740.740.00
NDCG@100.800.81−0.01
Latency p502.0s2.0s0.0s
Latency p953.8s3.8s0.0s
Cost per request$0.0041$0.00410.0000

Nightly trend

Nightly runs on the support-golden-v4 slice, Jul 11 to Oct 8. This run is one point in the series.

Pass rate moved from 78.5% to 87.2% across 90 nightly runs. Configuration changes: v3.3 on Aug 4, v3.4 on Sep 8.

Across datasets

support-assistant v3.4 on every dataset in the project. Rows marked “matrix” come from the comparison matrix rather than a dedicated run.

Configuration results per dataset
DatasetItemsPass rateFaith.Ctx recallRecall@5p95Source
support-golden-v4this run1,240
86.6%
0.910.830.883.8sThis run
billing-edge-cases312
81.7%
0.870.820.864.0sRun · Oct 2
roaming-and-travel188
80.9%
0.880.790.833.9sRun · Sep 28
device-troubleshooting426
84.5%
0.890.830.874.5sRun · Sep 30
account-security154
88.3%
0.920.850.893.7sRun · Sep 25
locale-formatting96
80.2%
0.860.800.844.1sRun · Oct 4
synthetic-hard-negatives540
76.5%
0.860.730.774.0smatrix

Traced samples

17 samples with full traces (4 failing, 13 passing), stratified by domain. Open a row to inspect retrieval and the judge's reasoning.

Traced samples for Nightly · v3.4 · 2026-09-15

How much data can I use in Germany on Unlimited Plus before I get charged?

Roaming & travel

FailMismatchjudge 0.18

Can I get a refund for the roaming pack I didn't use?

Roaming & travel

FailUnsupportedjudge 0.04

My bill says 41,20 €. Is that forty-one euros or four thousand?

Billing & payments

Passjudge 0.87

How do I get a copy of the personal data Harbor holds about me?

Data privacy requests

Passjudge 0.97

Can I return a phone I bought online if I opened the box?

Returns & warranty

Passjudge 0.91

My phone shows 4G but mobile data does not work after the iOS update.

Device troubleshooting

Passjudge 0.91

Is SIM swap protection turned on by default?

Account & security

Passjudge 0.91

What is the late fee if I pay my bill after the due date?

Billing & payments

Passjudge 0.86

My eSIM QR code says it has already been used. What should I do?

eSIM & activation

Passjudge 0.96

How much data can I use in the EU on the Unlimited Plus plan?

Roaming & travel

Passjudge 0.88