Skip to content

run_2026-10-02_billing-edge-cases

v3.4 · billing-edge-cases v2.1

Claude Sonnet 5.5 · bge-m3 · hybrid (BM25 + dense, RRF) · support-prompt v7 on billing-edge-cases v2.1

CompletedExperimentProduction candidatesupport-assistant v3.4v3 · 320 tokens, semantic
Compare

Pass rate

81.7%

Share of samples the judge marked as passing all rubric checks.255 of 312 samples passed

Faithfulness

0.87

Share of answer claims supported by the retrieved context.

Answer correctness

0.84

Agreement with the gold answer on facts, not style.

Context precision

0.77

Rank-weighted precision of relevant chunks passed to the generator.

Context recall

0.82

Share of gold-answer statements supported by the retrieved context.

Latency p95

4.0s

95th percentile end-to-end time to final token.

Failures by category

57 failed samples · largest share: unsupported

Share of failed samples by category
CategoryShare
Unsupported claims28.1%
Incorrect factual answers21.1%
Missing context19.3%
Retrieval mismatch17.5%
Instruction-following failures10.5%
Other3.5%

Run details

Dataset
billing-edge-cases v2.1 · 312 items
Generator
Claude Sonnet 5.5
Retriever
bge-m3 · hybrid (BM25 + dense, RRF)
Prompt
support-prompt v7
Judge
LLM judge · Claude Opus 5.5 · rubric v3
Started
Oct 2, 09:30 UTC
Duration
49m
Triggered by
Sara Lindqvist
Tokens / request
3,420 in · 214 out
  • billing

All metrics

Retrieval, generation and operational metrics

All metrics for v3.4 · billing-edge-cases v2.1
MetricThis run
Pass rate81.7%
Faithfulness0.87
Answer correctness0.84
Instruction adherence0.91
Context precision0.77
Context recall0.82
Recall@50.86
MRR0.72
NDCG@100.79
Latency p502.0s
Latency p954.0s
Cost per request$0.0043

Nightly trend

Nightly runs on the support-golden-v4 slice, Jul 11 to Oct 8. Dashed line marks this run.

  • Pass rate · nightly
  • This run · 81.7%
Pass rate moved from 78.5% to 87.2% across 90 nightly runs. Configuration changes: v3.3 on Aug 4, v3.4 on Sep 8.

Across datasets

support-assistant v3.4 on every dataset in the project. Rows marked “matrix” come from the comparison matrix rather than a dedicated run.

Configuration results per dataset
DatasetItemsPass rateFaith.Ctx recallRecall@5p95Source
support-golden-v41,240
87.2%
0.920.850.884.3sRun · Oct 8
billing-edge-casesthis run312
81.7%
0.870.820.864.0sThis run
roaming-and-travel188
80.9%
0.880.790.833.9sRun · Sep 28
device-troubleshooting426
84.5%
0.890.830.874.5sRun · Sep 30
account-security154
88.3%
0.920.850.893.7sRun · Sep 25
locale-formatting96
80.2%
0.860.800.844.1sRun · Oct 4
synthetic-hard-negatives540
76.5%
0.860.730.774.0smatrix

Traced samples

8 samples with full traces (4 failing, 4 passing), stratified by domain. Open a row to inspect retrieval and the judge's reasoning.

Traced samples for v3.4 · billing-edge-cases v2.1

Is there a fee to change my billing date?

Billing & payments

FailIncorrectjudge 0.46

Why is my bill higher this month after I downgraded on the 12th?

Billing & payments

FailMismatchjudge 0.05

What's the early termination fee if I cancel my 24-month contract after 10 months?

Plans & upgrades

FailMissingjudge 0.14

I downgraded on the 12th. Why is my bill higher this month?

Plans & upgrades

Passjudge 0.96

Does changing my billing date affect my first bill?

Billing & payments

Passjudge 0.92

When will I get my refund after canceling my contract?

Billing & payments

Passjudge 0.92

My direct debit failed. Will I be charged a fee?

Billing & payments

Passjudge 0.94

I upgraded from Essential to Unlimited on the 18th. Why is my bill €41.20 instead of €34.90?

Billing & payments

FailUnsupportedjudge 0.06

8 rows