run_2026-09-28_roaming-and-travel
v3.4 · roaming-and-travel v1.4
Claude Sonnet 5.5 · bge-m3 · hybrid (BM25 + dense, RRF) · support-prompt v7 on roaming-and-travel v1.4
Re-run after the roaming fair-use article update (kb-roaming-011).
Pass rate
80.9%
Faithfulness
0.88
Answer correctness
0.85
Context precision
0.74
Context recall
0.79
Latency p95
3.9s
Failures by category
36 failed samples · largest share: unsupported
| Category | Share |
|---|---|
| Unsupported claims | 27.8% |
| Incorrect factual answers | 22.2% |
| Missing context | 19.4% |
| Retrieval mismatch | 16.7% |
| Instruction-following failures | 11.1% |
| Other | 2.8% |
Run details
- Dataset
- roaming-and-travel v1.4 · 188 items
- Generator
- Claude Sonnet 5.5
- Retriever
- bge-m3 · hybrid (BM25 + dense, RRF)
- Prompt
- support-prompt v7
- Judge
- LLM judge · Claude Opus 5.5 · rubric v3
- Started
- Sep 28, 11:00 UTC
- Duration
- 33m
- Triggered by
- Daniel Okafor
- Tokens / request
- 3,420 in · 214 out
- roaming
- policy-update
All metrics
Retrieval, generation and operational metrics
| Metric | Group | This run |
|---|---|---|
| Pass rateShare of samples the judge marked as passing all rubric checks. | Outcome | 80.9% |
| FaithfulnessShare of answer claims supported by the retrieved context. | Generation | 0.88 |
| Answer correctnessAgreement with the gold answer on facts, not style. | Generation | 0.85 |
| Instruction adherenceRespect for tone, format, length, language and policy constraints. | Generation | 0.92 |
| Context precisionRank-weighted precision of relevant chunks passed to the generator. | Retrieval | 0.74 |
| Context recallShare of gold-answer statements supported by the retrieved context. | Retrieval | 0.79 |
| Recall@5Share of gold passages found in the top 5 retrieved chunks. | Retrieval | 0.83 |
| MRRMean reciprocal rank of the first relevant chunk. | Retrieval | 0.69 |
| NDCG@10Graded ranking quality of the top 10, normalized by the ideal order. | Retrieval | 0.76 |
| Latency p50Median end-to-end time to final token. | Operational | 2.0s |
| Latency p9595th percentile end-to-end time to final token. | Operational | 3.9s |
| Cost per requestModel, embedding and reranking cost per answered request (USD). | Operational | $0.0043 |
Nightly trend
Nightly runs on the support-golden-v4 slice, Jul 11 to Oct 8. Dashed line marks this run.
- Pass rate · nightly
- This run · 80.9%
Across datasets
support-assistant v3.4 on every dataset in the project. Rows marked “matrix” come from the comparison matrix rather than a dedicated run.
| Dataset | Items | Pass rate | Faith. | Ctx recall | Recall@5 | p95 | Source |
|---|---|---|---|---|---|---|---|
| support-golden-v4 | 1,240 | 87.2% | 0.92 | 0.85 | 0.88 | 4.3s | Run · Oct 8 |
| billing-edge-cases | 312 | 81.7% | 0.87 | 0.82 | 0.86 | 4.0s | Run · Oct 2 |
| roaming-and-travelthis run | 188 | 80.9% | 0.88 | 0.79 | 0.83 | 3.9s | This run |
| device-troubleshooting | 426 | 84.5% | 0.89 | 0.83 | 0.87 | 4.5s | Run · Sep 30 |
| account-security | 154 | 88.3% | 0.92 | 0.85 | 0.89 | 3.7s | Run · Sep 25 |
| locale-formatting | 96 | 80.2% | 0.86 | 0.80 | 0.84 | 4.1s | Run · Oct 4 |
| synthetic-hard-negatives | 540 | 76.5% | 0.86 | 0.73 | 0.77 | 4.0s | matrix |
Traced samples
7 samples with full traces (3 failing, 4 passing), stratified by domain. Open a row to inspect retrieval and the judge's reasoning.
Do I need a roaming pack for a weekend in Zurich? Roaming & travel FailIncorrectjudge 0.05 |
I'm going on a cruise in the Mediterranean. Will my phone work? Roaming & travel FailMissingjudge 0.06 |
Is Switzerland included in Zone 1 roaming? Roaming & travel Passjudge 0.94 |
How do I stop satellite roaming charges on a flight? Roaming & travel FailMissingjudge 0.50 |
Can I use my phone on a cruise ship? Roaming & travel Passjudge 0.91 |
If I don't use my weekly roaming pack, do the GBs carry over? Roaming & travel Passjudge 0.98 |
Do I need a roaming pack for a trip to London? Roaming & travel Passjudge 0.97 |
7 rows