Skip to content

run_2026-08-04_nightly

Nightly · v3.3 · 2026-08-04

Claude Sonnet 5.5 · bge-m3 · hybrid (BM25 + dense, RRF) · support-prompt v6 on support-golden-v4 v4.6

CompletedNightlyIn productionsupport-assistant v3.3v2 · 512 tokens / 64 overlap
Compare

Stratified slice of support-golden-v4 (397 items).

Pass rate

80.4%

+3.8 pts(improvement)vs previous nightly319 of 397 samples passed

Faithfulness

0.87

+0.02(improvement)vs previous nightly

Answer correctness

0.85

+0.05(improvement)vs previous nightly

Context precision

0.74

+0.04(improvement)vs previous nightly

Context recall

0.78

+0.02(improvement)vs previous nightly

Latency p95

3.8s

+0.2s(regression)vs previous nightly

Failures by category

78 failed samples · largest share: unsupported

Share of failed samples by category, compared with Nightly · v3.2 · 2026-08-03
CategorySharevs reference
Unsupported claims25.6%+1.7 pts
Incorrect factual answers20.5%+1.8 pts
Missing context21.8%−1.1 pts
Retrieval mismatch20.5%−3.4 pts
Instruction-following failures9.0%+1.7 pts
Other2.6%−0.6 pts

Run details

Dataset
support-golden-v4 v4.6 · 397 items
Generator
Claude Sonnet 5.5
Retriever
bge-m3 · hybrid (BM25 + dense, RRF)
Prompt
support-prompt v6
Judge
LLM judge · Claude Opus 5.5 · rubric v3
Started
Aug 4, 02:00 UTC
Duration
44m 18s
Triggered by
Scheduler
Tokens / request
3,980 in · 221 out
  • nightly
  • golden-slice

All metrics

Against Nightly · v3.2 · 2026-08-03

All metrics for Nightly · v3.3 · 2026-08-04
MetricThis runReferenceChange
Pass rate80.4%76.6%+3.8 pts
Faithfulness0.870.84+0.02
Answer correctness0.850.80+0.05
Instruction adherence0.890.87+0.02
Context precision0.740.70+0.04
Context recall0.780.76+0.02
Recall@50.830.78+0.05
MRR0.700.64+0.06
NDCG@100.770.73+0.04
Latency p501.8s1.6s+0.2s
Latency p953.8s3.6s+0.2s
Cost per request$0.0047$0.0044+0.0003

Nightly trend

Nightly runs on the support-golden-v4 slice, Jul 11 to Oct 8. This run is one point in the series.

Pass rate moved from 78.5% to 87.2% across 90 nightly runs. Configuration changes: v3.3 on Aug 4, v3.4 on Sep 8.

Across datasets

support-assistant v3.3 on every dataset in the project. Rows marked “matrix” come from the comparison matrix rather than a dedicated run.

Configuration results per dataset
DatasetItemsPass rateFaith.Ctx recallRecall@5p95Source
support-golden-v4this run1,240
80.4%
0.870.780.833.8sThis run
billing-edge-cases312
76.0%
0.830.770.813.9sRun · Aug 20
roaming-and-travel188
75.5%
0.840.740.783.8smatrix
device-troubleshooting426
78.9%
0.850.780.824.4smatrix
account-security154
82.5%
0.880.800.843.6smatrix
locale-formatting96
75.0%
0.820.750.794.0smatrix
synthetic-hard-negatives540
70.7%
0.820.680.723.9smatrix

Traced samples

17 samples with full traces (4 failing, 13 passing), stratified by domain. Open a row to inspect retrieval and the judge's reasoning.

Traced samples for Nightly · v3.3 · 2026-08-04

Can I return my phone after 20 days if I don't like it?

Returns & warranty

FailIncorrectjudge 0.11

What happens if I go over my fair-use limit on Unlimited at home?

Plans & upgrades

FailMissingjudge 0.07

Does the Fall 2026 upgrade offer apply if I'm on a family plan?

Promotions

FailMismatchjudge 0.09

My mobile data isn't working after updating to iOS 26.1.

Device troubleshooting

FailUnsupportedjudge 0.52

Can I return a phone I bought online if I opened the box?

Returns & warranty

Passjudge 0.92

My bill says 41,20 €. Is that forty-one euros or four thousand?

Billing & payments

Passjudge 0.87

My eSIM QR code says it has already been used. What should I do?

eSIM & activation

Passjudge 0.92

Is SIM swap protection turned on by default?

Account & security

Passjudge 0.94

If I upgrade my plan mid-month, how is my next bill calculated?

Plans & upgrades

Passjudge 0.97

How much data can I use in the EU on the Unlimited Plus plan?

Roaming & travel

Passjudge 0.98