Skip to content

run_2026-08-02_nightly

Nightly · v3.2 · 2026-08-02

Claude Sonnet 5.5 · bge-m3 · dense · support-prompt v5 on support-golden-v4 v4.6

CompletedNightlysupport-assistant v3.2v2 · 512 tokens / 64 overlap
Compare

Stratified slice of support-golden-v4 (423 items).

Pass rate

76.4%

−2.2 pts(regression)vs previous nightly323 of 423 samples passed

Faithfulness

0.85

+0.01(improvement)vs previous nightly

Answer correctness

0.80

−0.01(regression)vs previous nightly

Context precision

0.71

0.00(no change)vs previous nightly

Context recall

0.74

−0.02(regression)vs previous nightly

Latency p95

3.6s

+0.1s(regression)vs previous nightly

Failures by category

100 failed samples · largest share: unsupported

Share of failed samples by category, compared with Nightly · v3.2 · 2026-08-01
CategorySharevs reference
Unsupported claims24.0%−0.4 pts
Incorrect factual answers19.0%+0.1 pts
Missing context23.0%−0.3 pts
Retrieval mismatch24.0%+0.7 pts
Instruction-following failures7.0%+0.3 pts
Other3.0%−0.3 pts

Run details

Dataset
support-golden-v4 v4.6 · 423 items
Generator
Claude Sonnet 5.5
Retriever
bge-m3 · dense
Prompt
support-prompt v5
Judge
LLM judge · Claude Opus 5.5 · rubric v3
Started
Aug 2, 02:00 UTC
Duration
44m 18s
Triggered by
Scheduler
Tokens / request
3,870 in · 230 out
  • nightly
  • golden-slice

All metrics

Against Nightly · v3.2 · 2026-08-01

All metrics for Nightly · v3.2 · 2026-08-02
MetricThis runReferenceChange
Pass rate76.4%78.6%−2.2 pts
Faithfulness0.850.85+0.01
Answer correctness0.800.81−0.01
Instruction adherence0.870.87−0.01
Context precision0.710.700.00
Context recall0.740.76−0.02
Recall@50.780.780.00
MRR0.640.640.00
NDCG@100.710.72−0.01
Latency p501.6s1.6s0.0s
Latency p953.6s3.4s+0.1s
Cost per request$0.0043$0.0044−0.0001

Nightly trend

Nightly runs on the support-golden-v4 slice, Jul 11 to Oct 8. This run is one point in the series.

Pass rate moved from 78.5% to 87.2% across 90 nightly runs. Configuration changes: v3.3 on Aug 4, v3.4 on Sep 8.

Across datasets

support-assistant v3.2 on every dataset in the project. Rows marked “matrix” come from the comparison matrix rather than a dedicated run.

Configuration results per dataset
DatasetItemsPass rateFaith.Ctx recallRecall@5p95Source
support-golden-v4this run1,240
76.4%
0.850.740.783.6sThis run
billing-edge-cases312Not evaluated with this configuration
roaming-and-travel188Not evaluated with this configuration
device-troubleshooting426Not evaluated with this configuration
account-security154Not evaluated with this configuration
locale-formatting96Not evaluated with this configuration
synthetic-hard-negatives540Not evaluated with this configuration

Traced samples

17 samples with full traces (4 failing, 13 passing), stratified by domain. Open a row to inspect retrieval and the judge's reasoning.

Traced samples for Nightly · v3.2 · 2026-08-02

Can I keep my old number if I move from prepaid to a contract plan?

Plans & upgrades

FailIncorrectjudge 0.48

How much data can I use in Germany on Unlimited Plus before I get charged?

Roaming & travel

FailMismatchjudge 0.18

My eSIM stopped working after I reset my phone. Do I need to pay for a new one?

eSIM & activation

FailUnsupportedjudge 0.12

My eSIM QR code says it has already been used. What should I do?

eSIM & activation

Passjudge 0.92

Is SIM swap protection turned on by default?

Account & security

Passjudge 0.95

Can I return a phone I bought online if I opened the box?

Returns & warranty

Passjudge 0.95

My bill says 41,20 €. Is that forty-one euros or four thousand?

Billing & payments

Passjudge 0.98

Why was I charged twice for my September bill?

Billing & payments

Passjudge 0.90

If I upgrade my plan mid-month, how is my next bill calculated?

Plans & upgrades

Passjudge 0.92

How much data can I use in the EU on the Unlimited Plus plan?

Roaming & travel

Passjudge 0.95