Skip to content

run_2026-09-28_roaming-and-travel

v3.4 · roaming-and-travel v1.4

Claude Sonnet 5.5 · bge-m3 · hybrid (BM25 + dense, RRF) · support-prompt v7 on roaming-and-travel v1.4

CompletedExperimentProduction candidatesupport-assistant v3.4v3 · 320 tokens, semantic
Compare

Re-run after the roaming fair-use article update (kb-roaming-011).

Pass rate

80.9%

Share of samples the judge marked as passing all rubric checks.152 of 188 samples passed

Faithfulness

0.88

Share of answer claims supported by the retrieved context.

Answer correctness

0.85

Agreement with the gold answer on facts, not style.

Context precision

0.74

Rank-weighted precision of relevant chunks passed to the generator.

Context recall

0.79

Share of gold-answer statements supported by the retrieved context.

Latency p95

3.9s

95th percentile end-to-end time to final token.

Failures by category

36 failed samples · largest share: unsupported

Share of failed samples by category
CategoryShare
Unsupported claims27.8%
Incorrect factual answers22.2%
Missing context19.4%
Retrieval mismatch16.7%
Instruction-following failures11.1%
Other2.8%

Run details

Dataset
roaming-and-travel v1.4 · 188 items
Generator
Claude Sonnet 5.5
Retriever
bge-m3 · hybrid (BM25 + dense, RRF)
Prompt
support-prompt v7
Judge
LLM judge · Claude Opus 5.5 · rubric v3
Started
Sep 28, 11:00 UTC
Duration
33m
Triggered by
Daniel Okafor
Tokens / request
3,420 in · 214 out
  • roaming
  • policy-update

All metrics

Retrieval, generation and operational metrics

All metrics for v3.4 · roaming-and-travel v1.4
MetricThis run
Pass rate80.9%
Faithfulness0.88
Answer correctness0.85
Instruction adherence0.92
Context precision0.74
Context recall0.79
Recall@50.83
MRR0.69
NDCG@100.76
Latency p502.0s
Latency p953.9s
Cost per request$0.0043

Nightly trend

Nightly runs on the support-golden-v4 slice, Jul 11 to Oct 8. Dashed line marks this run.

  • Pass rate · nightly
  • This run · 80.9%
Pass rate moved from 78.5% to 87.2% across 90 nightly runs. Configuration changes: v3.3 on Aug 4, v3.4 on Sep 8.

Across datasets

support-assistant v3.4 on every dataset in the project. Rows marked “matrix” come from the comparison matrix rather than a dedicated run.

Configuration results per dataset
DatasetItemsPass rateFaith.Ctx recallRecall@5p95Source
support-golden-v41,240
87.2%
0.920.850.884.3sRun · Oct 8
billing-edge-cases312
81.7%
0.870.820.864.0sRun · Oct 2
roaming-and-travelthis run188
80.9%
0.880.790.833.9sThis run
device-troubleshooting426
84.5%
0.890.830.874.5sRun · Sep 30
account-security154
88.3%
0.920.850.893.7sRun · Sep 25
locale-formatting96
80.2%
0.860.800.844.1sRun · Oct 4
synthetic-hard-negatives540
76.5%
0.860.730.774.0smatrix

Traced samples

7 samples with full traces (3 failing, 4 passing), stratified by domain. Open a row to inspect retrieval and the judge's reasoning.

Traced samples for v3.4 · roaming-and-travel v1.4

Do I need a roaming pack for a weekend in Zurich?

Roaming & travel

FailIncorrectjudge 0.05

I'm going on a cruise in the Mediterranean. Will my phone work?

Roaming & travel

FailMissingjudge 0.06

Is Switzerland included in Zone 1 roaming?

Roaming & travel

Passjudge 0.94

How do I stop satellite roaming charges on a flight?

Roaming & travel

FailMissingjudge 0.50

Can I use my phone on a cruise ship?

Roaming & travel

Passjudge 0.91

If I don't use my weekly roaming pack, do the GBs carry over?

Roaming & travel

Passjudge 0.98

Do I need a roaming pack for a trip to London?

Roaming & travel

Passjudge 0.97

7 rows