run_2026-07-13_v3-2-release
support-assistant v3.2 · release
Claude Sonnet 5.5 · bge-m3 · dense · support-prompt v5 on support-golden-v4 v4.6
Pass rate
77.5%
Faithfulness
0.85
Answer correctness
0.81
Context precision
0.70
Context recall
0.75
Latency p95
3.5s
Failures by category
279 failed samples · largest share: unsupported
| Category | Share | vs reference |
|---|---|---|
| Unsupported claims | 24.0% | −3.1 pts |
| Incorrect factual answers | 19.0% | −2.0 pts |
| Missing context | 22.9% | +3.8 pts |
| Retrieval mismatch | 24.0% | +6.1 pts |
| Instruction-following failures | 7.2% | −3.9 pts |
| Other | 2.9% | −0.8 pts |
Run details
- Dataset
- support-golden-v4 v4.6 · 1,240 items
- Generator
- Claude Sonnet 5.5
- Retriever
- bge-m3 · dense
- Prompt
- support-prompt v5
- Judge
- LLM judge · Claude Opus 5.5 · rubric v3
- Started
- Jul 13, 12:30 UTC
- Duration
- 2h 37m
- Triggered by
- Maya Collins
- Tokens / request
- 3,870 in · 230 out
- Compared with
- support-assistant v3.4 · release candidate
- release
All metrics
Against support-assistant v3.4 · release candidate
| Metric | Group | This run | Reference | Change |
|---|---|---|---|---|
| Pass rateShare of samples the judge marked as passing all rubric checks. | Outcome | 77.5% | 86.9% | −9.4 pts |
| FaithfulnessShare of answer claims supported by the retrieved context. | Generation | 0.85 | 0.91 | −0.06 |
| Answer correctnessAgreement with the gold answer on facts, not style. | Generation | 0.81 | 0.88 | −0.07 |
| Instruction adherenceRespect for tone, format, length, language and policy constraints. | Generation | 0.87 | 0.93 | −0.06 |
| Context precisionRank-weighted precision of relevant chunks passed to the generator. | Retrieval | 0.70 | 0.79 | −0.09 |
| Context recallShare of gold-answer statements supported by the retrieved context. | Retrieval | 0.75 | 0.84 | −0.09 |
| Recall@5Share of gold passages found in the top 5 retrieved chunks. | Retrieval | 0.78 | 0.88 | −0.10 |
| MRRMean reciprocal rank of the first relevant chunk. | Retrieval | 0.64 | 0.74 | −0.10 |
| NDCG@10Graded ranking quality of the top 10, normalized by the ideal order. | Retrieval | 0.72 | 0.81 | −0.09 |
| Latency p50Median end-to-end time to final token. | Operational | 1.7s | 1.9s | −0.2s |
| Latency p9595th percentile end-to-end time to final token. | Operational | 3.5s | 3.8s | −0.3s |
| Cost per requestModel, embedding and reranking cost per answered request (USD). | Operational | $0.0044 | $0.0042 | +0.0002 |
Nightly trend
Nightly runs on the support-golden-v4 slice, Jul 11 to Oct 8. Dashed line marks this run.
- Pass rate · nightly
- This run · 77.5%
Across datasets
support-assistant v3.2 on every dataset in the project. Rows marked “matrix” come from the comparison matrix rather than a dedicated run.
| Dataset | Items | Pass rate | Faith. | Ctx recall | Recall@5 | p95 | Source |
|---|---|---|---|---|---|---|---|
| support-golden-v4this run | 1,240 | 77.5% | 0.85 | 0.75 | 0.78 | 3.5s | This run |
| billing-edge-cases | 312 | Not evaluated with this configuration | |||||
| roaming-and-travel | 188 | Not evaluated with this configuration | |||||
| device-troubleshooting | 426 | Not evaluated with this configuration | |||||
| account-security | 154 | Not evaluated with this configuration | |||||
| locale-formatting | 96 | Not evaluated with this configuration | |||||
| synthetic-hard-negatives | 540 | Not evaluated with this configuration | |||||
Traced samples
17 samples with full traces (4 failing, 13 passing), stratified by domain. Open a row to inspect retrieval and the judge's reasoning.
Can I return my phone after 20 days if I don't like it? Returns & warranty FailIncorrectjudge 0.11 |
What happens if I go over my fair-use limit on Unlimited at home? Plans & upgrades FailMissingjudge 0.07 |
Does the Fall 2026 upgrade offer apply if I'm on a family plan? Promotions FailMismatchjudge 0.09 |
My mobile data isn't working after updating to iOS 26.1. Device troubleshooting FailUnsupportedjudge 0.52 |
Can I return a phone I bought online if I opened the box? Returns & warranty Passjudge 0.97 |
My bill says 41,20 €. Is that forty-one euros or four thousand? Billing & payments Passjudge 0.98 |
My eSIM QR code says it has already been used. What should I do? eSIM & activation Passjudge 0.97 |
Is SIM swap protection turned on by default? Account & security Passjudge 0.89 |
If I upgrade my plan mid-month, how is my next bill calculated? Plans & upgrades Passjudge 0.86 |
How much data can I use in the EU on the Unlimited Plus plan? Roaming & travel Passjudge 0.92 |