Skip to content

run_2026-08-03_v3-3-release

support-assistant v3.3 · release

Claude Sonnet 5.5 · bge-m3 · hybrid (BM25 + dense, RRF) · support-prompt v6 on support-golden-v4 v4.6

CompletedReleaseIn productionsupport-assistant v3.3v2 · 512 tokens / 64 overlap
Compare

Prompt v6 and hybrid retrieval. +3.7 pts pass rate vs v3.2.

Pass rate

81.2%

+3.7 pts(improvement)vs support-assistant v3.2 · release1,007 of 1,240 samples passed

Faithfulness

0.87

+0.02(improvement)vs support-assistant v3.2 · release

Answer correctness

0.84

+0.03(improvement)vs support-assistant v3.2 · release

Context precision

0.74

+0.04(improvement)vs support-assistant v3.2 · release

Context recall

0.79

+0.04(improvement)vs support-assistant v3.2 · release

Latency p95

3.7s

+0.2s(regression)vs support-assistant v3.2 · release

Failures by category

233 failed samples · largest share: unsupported

Share of failed samples by category, compared with support-assistant v3.2 · release
CategorySharevs reference
Unsupported claims25.8%+1.7 pts
Incorrect factual answers20.2%+1.2 pts
Missing context21.0%−1.9 pts
Retrieval mismatch21.0%−3.0 pts
Instruction-following failures9.0%+1.8 pts
Other3.0%+0.1 pts

Run details

Dataset
support-golden-v4 v4.6 · 1,240 items
Generator
Claude Sonnet 5.5
Retriever
bge-m3 · hybrid (BM25 + dense, RRF)
Prompt
support-prompt v6
Judge
LLM judge · Claude Opus 5.5 · rubric v3
Started
Aug 3, 13:00 UTC
Duration
2h 46m
Triggered by
Maya Collins
Tokens / request
3,980 in · 221 out
  • release
  • regression-gate

All metrics

Against support-assistant v3.2 · release

All metrics for support-assistant v3.3 · release
MetricThis runReferenceChange
Pass rate81.2%77.5%+3.7 pts
Faithfulness0.870.85+0.02
Answer correctness0.840.81+0.03
Instruction adherence0.890.87+0.02
Context precision0.740.70+0.04
Context recall0.790.75+0.04
Recall@50.830.78+0.05
MRR0.690.64+0.05
NDCG@100.760.72+0.04
Latency p501.8s1.7s+0.1s
Latency p953.7s3.5s+0.2s
Cost per request$0.0046$0.0044+0.0002

Nightly trend

Nightly runs on the support-golden-v4 slice, Jul 11 to Oct 8. Dashed line marks this run.

  • Pass rate · nightly
  • This run · 81.2%
Pass rate moved from 78.5% to 87.2% across 90 nightly runs. Configuration changes: v3.3 on Aug 4, v3.4 on Sep 8.

Across datasets

support-assistant v3.3 on every dataset in the project. Rows marked “matrix” come from the comparison matrix rather than a dedicated run.

Configuration results per dataset
DatasetItemsPass rateFaith.Ctx recallRecall@5p95Source
support-golden-v4this run1,240
81.2%
0.870.790.833.7sThis run
billing-edge-cases312
76.0%
0.830.770.813.9sRun · Aug 20
roaming-and-travel188
75.5%
0.840.740.783.8smatrix
device-troubleshooting426
78.9%
0.850.780.824.4smatrix
account-security154
82.5%
0.880.800.843.6smatrix
locale-formatting96
75.0%
0.820.750.794.0smatrix
synthetic-hard-negatives540
70.7%
0.820.680.723.9smatrix

Traced samples

17 samples with full traces (4 failing, 13 passing), stratified by domain. Open a row to inspect retrieval and the judge's reasoning.

Traced samples for support-assistant v3.3 · release

Is there a restocking fee if I return an opened phone?

Returns & warranty

FailMissingjudge 0.13

Can I pay my bill with a credit card from another person?

Billing & payments

FailIncorrectjudge 0.58

Does the Fall 2026 upgrade offer work with a family plan discount?

Promotions

Passjudge 0.93

What is the late fee if I pay my bill after the due date?

Billing & payments

Passjudge 0.95

How long does it take to port my number to Harbor?

Number porting

Passjudge 0.90

How do I get a copy of the personal data Harbor holds about me?

Data privacy requests

Passjudge 0.89

My phone shows 4G but mobile data does not work after the iOS update.

Device troubleshooting

Passjudge 0.86

Why was I charged twice for my September bill?

Billing & payments

Passjudge 0.93

How much data can I use in the EU on the Unlimited Plus plan?

Roaming & travel

Passjudge 0.91

If I upgrade my plan mid-month, how is my next bill calculated?

Plans & upgrades

Passjudge 0.91