Skip to content

run_2026-07-13_v3-2-release

support-assistant v3.2 · release

Claude Sonnet 5.5 · bge-m3 · dense · support-prompt v5 on support-golden-v4 v4.6

CompletedReleasesupport-assistant v3.2v2 · 512 tokens / 64 overlap
Compare

Pass rate

77.5%

−9.4 pts(regression)vs support-assistant v3.4 · release candidate961 of 1,240 samples passed

Faithfulness

0.85

−0.06(regression)vs support-assistant v3.4 · release candidate

Answer correctness

0.81

−0.07(regression)vs support-assistant v3.4 · release candidate

Context precision

0.70

−0.09(regression)vs support-assistant v3.4 · release candidate

Context recall

0.75

−0.09(regression)vs support-assistant v3.4 · release candidate

Latency p95

3.5s

−0.3s(improvement)vs support-assistant v3.4 · release candidate

Failures by category

279 failed samples · largest share: unsupported

Share of failed samples by category, compared with support-assistant v3.4 · release candidate
CategorySharevs reference
Unsupported claims24.0%−3.1 pts
Incorrect factual answers19.0%−2.0 pts
Missing context22.9%+3.8 pts
Retrieval mismatch24.0%+6.1 pts
Instruction-following failures7.2%−3.9 pts
Other2.9%−0.8 pts

Run details

Dataset
support-golden-v4 v4.6 · 1,240 items
Generator
Claude Sonnet 5.5
Retriever
bge-m3 · dense
Prompt
support-prompt v5
Judge
LLM judge · Claude Opus 5.5 · rubric v3
Started
Jul 13, 12:30 UTC
Duration
2h 37m
Triggered by
Maya Collins
Tokens / request
3,870 in · 230 out
  • release

All metrics

Against support-assistant v3.4 · release candidate

All metrics for support-assistant v3.2 · release
MetricThis runReferenceChange
Pass rate77.5%86.9%−9.4 pts
Faithfulness0.850.91−0.06
Answer correctness0.810.88−0.07
Instruction adherence0.870.93−0.06
Context precision0.700.79−0.09
Context recall0.750.84−0.09
Recall@50.780.88−0.10
MRR0.640.74−0.10
NDCG@100.720.81−0.09
Latency p501.7s1.9s−0.2s
Latency p953.5s3.8s−0.3s
Cost per request$0.0044$0.0042+0.0002

Nightly trend

Nightly runs on the support-golden-v4 slice, Jul 11 to Oct 8. Dashed line marks this run.

  • Pass rate · nightly
  • This run · 77.5%
Pass rate moved from 78.5% to 87.2% across 90 nightly runs. Configuration changes: v3.3 on Aug 4, v3.4 on Sep 8.

Across datasets

support-assistant v3.2 on every dataset in the project. Rows marked “matrix” come from the comparison matrix rather than a dedicated run.

Configuration results per dataset
DatasetItemsPass rateFaith.Ctx recallRecall@5p95Source
support-golden-v4this run1,240
77.5%
0.850.750.783.5sThis run
billing-edge-cases312Not evaluated with this configuration
roaming-and-travel188Not evaluated with this configuration
device-troubleshooting426Not evaluated with this configuration
account-security154Not evaluated with this configuration
locale-formatting96Not evaluated with this configuration
synthetic-hard-negatives540Not evaluated with this configuration

Traced samples

17 samples with full traces (4 failing, 13 passing), stratified by domain. Open a row to inspect retrieval and the judge's reasoning.

Traced samples for support-assistant v3.2 · release

Can I return my phone after 20 days if I don't like it?

Returns & warranty

FailIncorrectjudge 0.11

What happens if I go over my fair-use limit on Unlimited at home?

Plans & upgrades

FailMissingjudge 0.07

Does the Fall 2026 upgrade offer apply if I'm on a family plan?

Promotions

FailMismatchjudge 0.09

My mobile data isn't working after updating to iOS 26.1.

Device troubleshooting

FailUnsupportedjudge 0.52

Can I return a phone I bought online if I opened the box?

Returns & warranty

Passjudge 0.97

My bill says 41,20 €. Is that forty-one euros or four thousand?

Billing & payments

Passjudge 0.98

My eSIM QR code says it has already been used. What should I do?

eSIM & activation

Passjudge 0.97

Is SIM swap protection turned on by default?

Account & security

Passjudge 0.89

If I upgrade my plan mid-month, how is my next bill calculated?

Plans & upgrades

Passjudge 0.86

How much data can I use in the EU on the Unlimited Plus plan?

Roaming & travel

Passjudge 0.92