Skip to content

run_2026-09-26_cohere-rerank

v3.4 · Cohere embed-v4 + rerank-3.5

Claude Sonnet 5.5 · Cohere embed-v4 · dense + rerank-3.5 · support-prompt v7 on support-golden-v4 v4.6

CompletedExperimentv3.4 · Cohere rerankv3 · 320 tokens, semantic
Compare

Pass rate

86.1%

−0.8 pts(regression)vs support-assistant v3.4 · release candidate1,068 of 1,240 samples passed

Faithfulness

0.91

0.00(no change)vs support-assistant v3.4 · release candidate

Answer correctness

0.87

−0.01(regression)vs support-assistant v3.4 · release candidate

Context precision

0.83

+0.04(improvement)vs support-assistant v3.4 · release candidate

Context recall

0.85

+0.01(improvement)vs support-assistant v3.4 · release candidate

Latency p95

4.3s

+0.5s(regression)vs support-assistant v3.4 · release candidate

Failures by category

172 failed samples · largest share: unsupported

Share of failed samples by category, compared with support-assistant v3.4 · release candidate
CategorySharevs reference
Unsupported claims27.9%+0.7 pts
Incorrect factual answers22.1%+1.1 pts
Missing context16.9%−2.3 pts
Retrieval mismatch15.1%−2.8 pts
Instruction-following failures12.8%+1.7 pts
Other5.2%+1.5 pts

Run details

Dataset
support-golden-v4 v4.6 · 1,240 items
Generator
Claude Sonnet 5.5
Retriever
Cohere embed-v4 · dense + rerank-3.5
Prompt
support-prompt v7
Judge
LLM judge · Claude Opus 5.5 · rubric v3
Started
Sep 26, 09:40 UTC
Duration
3h 2m
Triggered by
Sara Lindqvist
Tokens / request
3,410 in · 212 out
  • retriever
  • rerank

All metrics

Against support-assistant v3.4 · release candidate

All metrics for v3.4 · Cohere embed-v4 + rerank-3.5
MetricThis runReferenceChange
Pass rate86.1%86.9%−0.8 pts
Faithfulness0.910.910.00
Answer correctness0.870.88−0.01
Instruction adherence0.930.930.00
Context precision0.830.79+0.04
Context recall0.850.84+0.01
Recall@50.900.88+0.02
MRR0.790.74+0.05
NDCG@100.840.81+0.03
Latency p502.2s1.9s+0.3s
Latency p954.3s3.8s+0.5s
Cost per request$0.0049$0.0042+0.0007

Nightly trend

Nightly runs on the support-golden-v4 slice, Jul 11 to Oct 8. Dashed line marks this run.

  • Pass rate · nightly
  • This run · 86.1%
Pass rate moved from 78.5% to 87.2% across 90 nightly runs. Configuration changes: v3.3 on Aug 4, v3.4 on Sep 8.

Across datasets

v3.4 · Cohere rerank on every dataset in the project. Rows marked “matrix” come from the comparison matrix rather than a dedicated run.

Configuration results per dataset
DatasetItemsPass rateFaith.Ctx recallRecall@5p95Source
support-golden-v4this run1,240
86.1%
0.910.850.904.3sThis run
billing-edge-cases312
81.1%
0.870.830.884.5smatrix
roaming-and-travel188
80.3%
0.880.800.854.4smatrix
device-troubleshooting426
83.8%
0.890.840.895.0smatrix
account-security154
87.0%
0.920.860.914.2smatrix
locale-formatting96
79.2%
0.860.810.864.6smatrix
synthetic-hard-negatives540
75.7%
0.860.740.794.5smatrix

Traced samples

17 samples with full traces (4 failing, 13 passing), stratified by domain. Open a row to inspect retrieval and the judge's reasoning.

Traced samples for v3.4 · Cohere embed-v4 + rerank-3.5

How long does it take to port my number to Harbor?

Number porting

Passjudge 0.93

Does the Fall 2026 upgrade offer work with a family plan discount?

Promotions

Passjudge 0.92

Can I return a phone I bought online if I opened the box?

Returns & warranty

Passjudge 0.95

My bill says 41,20 €. Is that forty-one euros or four thousand?

Billing & payments

Passjudge 0.90

My eSIM QR code says it has already been used. What should I do?

eSIM & activation

Passjudge 0.93

Is SIM swap protection turned on by default?

Account & security

Passjudge 0.94

If I upgrade my plan mid-month, how is my next bill calculated?

Plans & upgrades

Passjudge 0.89

How much data can I use in the EU on the Unlimited Plus plan?

Roaming & travel

Passjudge 0.98

Why was I charged twice for my September bill?

Billing & payments

Passjudge 0.87

What is the excess on a theft insurance claim?

Returns & warranty

Passjudge 0.88