run_2026-08-20_v3-3-billing
v3.3 · billing-edge-cases v2.0
Claude Sonnet 5.5 · bge-m3 · hybrid (BM25 + dense, RRF) · support-prompt v6 on billing-edge-cases v2.1
Pass rate
76.0%
Faithfulness
0.83
Answer correctness
0.80
Context precision
0.72
Context recall
0.77
Latency p95
3.9s
Failures by category
75 failed samples · largest share: unsupported
| Category | Share |
|---|---|
| Unsupported claims | 25.3% |
| Incorrect factual answers | 20.0% |
| Missing context | 21.3% |
| Retrieval mismatch | 21.3% |
| Instruction-following failures | 9.3% |
| Other | 2.7% |
Run details
- Dataset
- billing-edge-cases v2.1 · 312 items
- Generator
- Claude Sonnet 5.5
- Retriever
- bge-m3 · hybrid (BM25 + dense, RRF)
- Prompt
- support-prompt v6
- Judge
- LLM judge · Claude Opus 5.5 · rubric v3
- Started
- Aug 20, 10:10 UTC
- Duration
- 48m
- Triggered by
- Sara Lindqvist
- Tokens / request
- 3,980 in · 221 out
- billing
All metrics
Retrieval, generation and operational metrics
| Metric | Group | This run |
|---|---|---|
| Pass rateShare of samples the judge marked as passing all rubric checks. | Outcome | 76.0% |
| FaithfulnessShare of answer claims supported by the retrieved context. | Generation | 0.83 |
| Answer correctnessAgreement with the gold answer on facts, not style. | Generation | 0.80 |
| Instruction adherenceRespect for tone, format, length, language and policy constraints. | Generation | 0.87 |
| Context precisionRank-weighted precision of relevant chunks passed to the generator. | Retrieval | 0.72 |
| Context recallShare of gold-answer statements supported by the retrieved context. | Retrieval | 0.77 |
| Recall@5Share of gold passages found in the top 5 retrieved chunks. | Retrieval | 0.81 |
| MRRMean reciprocal rank of the first relevant chunk. | Retrieval | 0.67 |
| NDCG@10Graded ranking quality of the top 10, normalized by the ideal order. | Retrieval | 0.74 |
| Latency p50Median end-to-end time to final token. | Operational | 1.9s |
| Latency p9595th percentile end-to-end time to final token. | Operational | 3.9s |
| Cost per requestModel, embedding and reranking cost per answered request (USD). | Operational | $0.0047 |
Nightly trend
Nightly runs on the support-golden-v4 slice, Jul 11 to Oct 8. Dashed line marks this run.
- Pass rate · nightly
- This run · 76.0%
Across datasets
support-assistant v3.3 on every dataset in the project. Rows marked “matrix” come from the comparison matrix rather than a dedicated run.
| Dataset | Items | Pass rate | Faith. | Ctx recall | Recall@5 | p95 | Source |
|---|---|---|---|---|---|---|---|
| support-golden-v4 | 1,240 | 81.1% | 0.87 | 0.80 | 0.82 | 3.6s | Run · Sep 7 |
| billing-edge-casesthis run | 312 | 76.0% | 0.83 | 0.77 | 0.81 | 3.9s | This run |
| roaming-and-travel | 188 | 75.5% | 0.84 | 0.74 | 0.78 | 3.8s | matrix |
| device-troubleshooting | 426 | 78.9% | 0.85 | 0.78 | 0.82 | 4.4s | matrix |
| account-security | 154 | 82.5% | 0.88 | 0.80 | 0.84 | 3.6s | matrix |
| locale-formatting | 96 | 75.0% | 0.82 | 0.75 | 0.79 | 4.0s | matrix |
| synthetic-hard-negatives | 540 | 70.7% | 0.82 | 0.68 | 0.72 | 3.9s | matrix |
Traced samples
8 samples with full traces (4 failing, 4 passing), stratified by domain. Open a row to inspect retrieval and the judge's reasoning.
I downgraded on the 12th. Why is my bill higher this month? Plans & upgrades Passjudge 0.93 |
When will I get my refund after canceling my contract? Billing & payments Passjudge 0.96 |
Does changing my billing date affect my first bill? Billing & payments Passjudge 0.93 |
My direct debit failed. Will I be charged a fee? Billing & payments Passjudge 0.94 |
Why is my bill higher this month after I downgraded on the 12th? Billing & payments FailMismatchjudge 0.05 |
I upgraded from Essential to Unlimited on the 18th. Why is my bill €41.20 instead of €34.90? Billing & payments FailUnsupportedjudge 0.06 |
What's the early termination fee if I cancel my 24-month contract after 10 months? Plans & upgrades FailMissingjudge 0.14 |
Is there a fee to change my billing date? Billing & payments FailIncorrectjudge 0.46 |
8 rows