Skip to content

run_2026-08-20_v3-3-billing

v3.3 · billing-edge-cases v2.0

Claude Sonnet 5.5 · bge-m3 · hybrid (BM25 + dense, RRF) · support-prompt v6 on billing-edge-cases v2.1

CompletedExperimentIn productionsupport-assistant v3.3v2 · 512 tokens / 64 overlap
Compare

Pass rate

76.0%

Share of samples the judge marked as passing all rubric checks.237 of 312 samples passed

Faithfulness

0.83

Share of answer claims supported by the retrieved context.

Answer correctness

0.80

Agreement with the gold answer on facts, not style.

Context precision

0.72

Rank-weighted precision of relevant chunks passed to the generator.

Context recall

0.77

Share of gold-answer statements supported by the retrieved context.

Latency p95

3.9s

95th percentile end-to-end time to final token.

Failures by category

75 failed samples · largest share: unsupported

Share of failed samples by category
CategoryShare
Unsupported claims25.3%
Incorrect factual answers20.0%
Missing context21.3%
Retrieval mismatch21.3%
Instruction-following failures9.3%
Other2.7%

Run details

Dataset
billing-edge-cases v2.1 · 312 items
Generator
Claude Sonnet 5.5
Retriever
bge-m3 · hybrid (BM25 + dense, RRF)
Prompt
support-prompt v6
Judge
LLM judge · Claude Opus 5.5 · rubric v3
Started
Aug 20, 10:10 UTC
Duration
48m
Triggered by
Sara Lindqvist
Tokens / request
3,980 in · 221 out
  • billing

All metrics

Retrieval, generation and operational metrics

All metrics for v3.3 · billing-edge-cases v2.0
MetricThis run
Pass rate76.0%
Faithfulness0.83
Answer correctness0.80
Instruction adherence0.87
Context precision0.72
Context recall0.77
Recall@50.81
MRR0.67
NDCG@100.74
Latency p501.9s
Latency p953.9s
Cost per request$0.0047

Nightly trend

Nightly runs on the support-golden-v4 slice, Jul 11 to Oct 8. Dashed line marks this run.

  • Pass rate · nightly
  • This run · 76.0%
Pass rate moved from 78.5% to 87.2% across 90 nightly runs. Configuration changes: v3.3 on Aug 4, v3.4 on Sep 8.

Across datasets

support-assistant v3.3 on every dataset in the project. Rows marked “matrix” come from the comparison matrix rather than a dedicated run.

Configuration results per dataset
DatasetItemsPass rateFaith.Ctx recallRecall@5p95Source
support-golden-v41,240
81.1%
0.870.800.823.6sRun · Sep 7
billing-edge-casesthis run312
76.0%
0.830.770.813.9sThis run
roaming-and-travel188
75.5%
0.840.740.783.8smatrix
device-troubleshooting426
78.9%
0.850.780.824.4smatrix
account-security154
82.5%
0.880.800.843.6smatrix
locale-formatting96
75.0%
0.820.750.794.0smatrix
synthetic-hard-negatives540
70.7%
0.820.680.723.9smatrix

Traced samples

8 samples with full traces (4 failing, 4 passing), stratified by domain. Open a row to inspect retrieval and the judge's reasoning.

Traced samples for v3.3 · billing-edge-cases v2.0

I downgraded on the 12th. Why is my bill higher this month?

Plans & upgrades

Passjudge 0.93

When will I get my refund after canceling my contract?

Billing & payments

Passjudge 0.96

Does changing my billing date affect my first bill?

Billing & payments

Passjudge 0.93

My direct debit failed. Will I be charged a fee?

Billing & payments

Passjudge 0.94

Why is my bill higher this month after I downgraded on the 12th?

Billing & payments

FailMismatchjudge 0.05

I upgraded from Essential to Unlimited on the 18th. Why is my bill €41.20 instead of €34.90?

Billing & payments

FailUnsupportedjudge 0.06

What's the early termination fee if I cancel my 24-month contract after 10 months?

Plans & upgrades

FailMissingjudge 0.14

Is there a fee to change my billing date?

Billing & payments

FailIncorrectjudge 0.46

8 rows