Skip to content

run_2026-07-24_nightly

Nightly · v3.2 · 2026-07-24

Claude Sonnet 5.5 · bge-m3 · dense · support-prompt v5 on support-golden-v4 v4.6

CompletedNightlysupport-assistant v3.2v2 · 512 tokens / 64 overlap
Compare

Stratified slice of support-golden-v4 (411 items).

Pass rate

78.1%

0.0 pts(no change)vs previous nightly321 of 411 samples passed

Faithfulness

0.85

−0.01(regression)vs previous nightly

Answer correctness

0.81

+0.01(improvement)vs previous nightly

Context precision

0.71

+0.02(improvement)vs previous nightly

Context recall

0.74

−0.01(regression)vs previous nightly

Latency p95

3.5s

0.0s(no change)vs previous nightly

Failures by category

90 failed samples · largest share: unsupported

Share of failed samples by category, compared with Nightly · v3.2 · 2026-07-23
CategorySharevs reference
Unsupported claims24.4%+0.5 pts
Incorrect factual answers18.9%−0.7 pts
Missing context23.3%+0.5 pts
Retrieval mismatch23.3%−0.6 pts
Instruction-following failures6.7%+0.1 pts
Other3.3%+0.1 pts

Run details

Dataset
support-golden-v4 v4.6 · 411 items
Generator
Claude Sonnet 5.5
Retriever
bge-m3 · dense
Prompt
support-prompt v5
Judge
LLM judge · Claude Opus 5.5 · rubric v3
Started
Jul 24, 02:00 UTC
Duration
40m 46s
Triggered by
Scheduler
Tokens / request
3,870 in · 230 out
  • nightly
  • golden-slice

All metrics

Against Nightly · v3.2 · 2026-07-23

All metrics for Nightly · v3.2 · 2026-07-24
MetricThis runReferenceChange
Pass rate78.1%78.1%0.0 pts
Faithfulness0.850.86−0.01
Answer correctness0.810.80+0.01
Instruction adherence0.870.870.00
Context precision0.710.69+0.02
Context recall0.740.75−0.01
Recall@50.770.79−0.02
MRR0.630.65−0.01
NDCG@100.720.71+0.02
Latency p501.7s1.8s0.0s
Latency p953.5s3.5s0.0s
Cost per request$0.0044$0.0045−0.0001

Nightly trend

Nightly runs on the support-golden-v4 slice, Jul 11 to Oct 8. This run is one point in the series.

Pass rate moved from 78.5% to 87.2% across 90 nightly runs. Configuration changes: v3.3 on Aug 4, v3.4 on Sep 8.

Across datasets

support-assistant v3.2 on every dataset in the project. Rows marked “matrix” come from the comparison matrix rather than a dedicated run.

Configuration results per dataset
DatasetItemsPass rateFaith.Ctx recallRecall@5p95Source
support-golden-v4this run1,240
78.1%
0.850.740.773.5sThis run
billing-edge-cases312Not evaluated with this configuration
roaming-and-travel188Not evaluated with this configuration
device-troubleshooting426Not evaluated with this configuration
account-security154Not evaluated with this configuration
locale-formatting96Not evaluated with this configuration
synthetic-hard-negatives540Not evaluated with this configuration

Traced samples

17 samples with full traces (4 failing, 13 passing), stratified by domain. Open a row to inspect retrieval and the judge's reasoning.

Traced samples for Nightly · v3.2 · 2026-07-24

What happens if I go over my fair-use limit on Unlimited at home?

Plans & upgrades

FailMissingjudge 0.07

How much data can I use in Germany on Unlimited Plus before I get charged?

Roaming & travel

FailMismatchjudge 0.18

Can I keep my old number if I move from prepaid to a contract plan?

Plans & upgrades

FailIncorrectjudge 0.48

My eSIM stopped working after I reset my phone. Do I need to pay for a new one?

eSIM & activation

FailUnsupportedjudge 0.12

What is the late fee if I pay my bill after the due date?

Billing & payments

Passjudge 0.90

How do I get a copy of the personal data Harbor holds about me?

Data privacy requests

Passjudge 0.91

My phone shows 4G but mobile data does not work after the iOS update.

Device troubleshooting

Passjudge 0.91

Is SIM swap protection turned on by default?

Account & security

Passjudge 0.89

What is the excess on a theft insurance claim?

Returns & warranty

Passjudge 0.93

My eSIM QR code says it has already been used. What should I do?

eSIM & activation

Passjudge 0.87