Skip to content

run_2026-10-08_synthetic-hn

v3.4 · synthetic-hard-negatives v1.3

Claude Sonnet 5.5 · bge-m3 · hybrid (BM25 + dense, RRF) · support-prompt v7 on synthetic-hard-negatives v1.3

RunningExperimentProduction candidatesupport-assistant v3.4v3 · 320 tokens, semantic
Compare

Judging samples · 340 of 540

63% complete · about 50m remaining

Metrics update as results arrive. Partial figures are provisional until the run completes.

Pass rate

76.5%

Share of samples the judge marked as passing all rubric checks.260 of 340 samples passed

Faithfulness

0.86

Share of answer claims supported by the retrieved context.

Answer correctness

0.83

Agreement with the gold answer on facts, not style.

Context precision

0.68

Rank-weighted precision of relevant chunks passed to the generator.

Context recall

0.73

Share of gold-answer statements supported by the retrieved context.

Latency p95

4.0s

95th percentile end-to-end time to final token.

Failures by category

80 failed samples · largest share: unsupported

Share of failed samples by category
CategoryShare
Unsupported claims27.5%
Incorrect factual answers21.3%
Missing context18.8%
Retrieval mismatch17.5%
Instruction-following failures11.3%
Other3.8%

Run details

Dataset
synthetic-hard-negatives v1.3 · 540 items
Generator
Claude Sonnet 5.5
Retriever
bge-m3 · hybrid (BM25 + dense, RRF)
Prompt
support-prompt v7
Judge
LLM judge · Claude Opus 5.5 · rubric v3
Started
Oct 8, 08:05 UTC
Duration
In progress
Triggered by
Sara Lindqvist
Tokens / request
3,420 in · 214 out
  • adversarial
  • ranking

All metrics

Retrieval, generation and operational metrics

All metrics for v3.4 · synthetic-hard-negatives v1.3
MetricThis run
Pass rate76.5%
Faithfulness0.86
Answer correctness0.83
Instruction adherence0.91
Context precision0.68
Context recall0.73
Recall@50.77
MRR0.63
NDCG@100.70
Latency p502.0s
Latency p954.0s
Cost per request$0.0043

Nightly trend

Nightly runs on the support-golden-v4 slice, Jul 11 to Oct 8. Dashed line marks this run.

  • Pass rate · nightly
  • This run · 76.5%
Pass rate moved from 78.5% to 87.2% across 90 nightly runs. Configuration changes: v3.3 on Aug 4, v3.4 on Sep 8.

Across datasets

support-assistant v3.4 on every dataset in the project. Rows marked “matrix” come from the comparison matrix rather than a dedicated run.

Configuration results per dataset
DatasetItemsPass rateFaith.Ctx recallRecall@5p95Source
support-golden-v41,240
87.2%
0.920.850.884.3sRun · Oct 8
billing-edge-cases312
81.7%
0.870.820.864.0sRun · Oct 2
roaming-and-travel188
80.9%
0.880.790.833.9sRun · Sep 28
device-troubleshooting426
84.5%
0.890.830.874.5sRun · Sep 30
account-security154
88.3%
0.920.850.893.7sRun · Sep 25
locale-formatting96
80.2%
0.860.800.844.1sRun · Oct 4
synthetic-hard-negativesthis run540
76.5%
0.860.730.774.0smatrix

Traced samples

4 of 5 traced samples judged so far. Open a row to see retrieval and the judge's reasoning.

Traced samples for v3.4 · synthetic-hard-negatives v1.3

How much roaming data is included on Unlimited (not Unlimited Plus) in the EU?

Roaming & travel

Passjudge 0.97

Can I use my Unlimited hotspot allowance abroad?

Roaming & travel

FailMismatchjudge 0.16

Can I combine the refer-a-friend credit with the Fall 2026 upgrade offer?

Promotions

Passjudge 0.98

Which plan includes Disney+?

Plans & upgrades

FailMissingjudge 0.05

4 rows