Skip to content

run_2026-10-04_locale-formatting

v3.4 · locale-formatting

Claude Sonnet 5.5 · bge-m3 · hybrid (BM25 + dense, RRF) · support-prompt v7 on locale-formatting v1.1

CompletedExperimentProduction candidatesupport-assistant v3.4v3 · 320 tokens, semantic
Compare

Pass rate

80.2%

Share of samples the judge marked as passing all rubric checks.77 of 96 samples passed

Faithfulness

0.86

Share of answer claims supported by the retrieved context.

Answer correctness

0.83

Agreement with the gold answer on facts, not style.

Context precision

0.75

Rank-weighted precision of relevant chunks passed to the generator.

Context recall

0.80

Share of gold-answer statements supported by the retrieved context.

Latency p95

4.1s

95th percentile end-to-end time to final token.

Failures by category

19 failed samples · largest share: unsupported

Share of failed samples by category
CategoryShare
Unsupported claims26.3%
Incorrect factual answers21.1%
Missing context21.1%
Retrieval mismatch15.8%
Instruction-following failures10.5%
Other5.3%

Run details

Dataset
locale-formatting v1.1 · 96 items
Generator
Claude Sonnet 5.5
Retriever
bge-m3 · hybrid (BM25 + dense, RRF)
Prompt
support-prompt v7
Judge
LLM judge · Claude Opus 5.5 · rubric v3
Started
Oct 4, 10:15 UTC
Duration
21m
Triggered by
Priya Nair
Tokens / request
3,420 in · 214 out
  • locale
  • formatting

All metrics

Retrieval, generation and operational metrics

All metrics for v3.4 · locale-formatting
MetricThis run
Pass rate80.2%
Faithfulness0.86
Answer correctness0.83
Instruction adherence0.91
Context precision0.75
Context recall0.80
Recall@50.84
MRR0.70
NDCG@100.77
Latency p502.0s
Latency p954.1s
Cost per request$0.0044

Nightly trend

Nightly runs on the support-golden-v4 slice, Jul 11 to Oct 8. Dashed line marks this run.

  • Pass rate · nightly
  • This run · 80.2%
Pass rate moved from 78.5% to 87.2% across 90 nightly runs. Configuration changes: v3.3 on Aug 4, v3.4 on Sep 8.

Across datasets

support-assistant v3.4 on every dataset in the project. Rows marked “matrix” come from the comparison matrix rather than a dedicated run.

Configuration results per dataset
DatasetItemsPass rateFaith.Ctx recallRecall@5p95Source
support-golden-v41,240
87.2%
0.920.850.884.3sRun · Oct 8
billing-edge-cases312
81.7%
0.870.820.864.0sRun · Oct 2
roaming-and-travel188
80.9%
0.880.790.833.9sRun · Sep 28
device-troubleshooting426
84.5%
0.890.830.874.5sRun · Sep 30
account-security154
88.3%
0.920.850.893.7sRun · Sep 25
locale-formattingthis run96
80.2%
0.860.800.844.1sThis run
synthetic-hard-negatives540
76.5%
0.860.730.774.0smatrix

Traced samples

8 samples with full traces (5 failing, 3 passing), stratified by domain. Open a row to inspect retrieval and the judge's reasoning.

Traced samples for v3.4 · locale-formatting

I'm a UK customer. How do I ask Harbor to delete my personal data?

Data privacy requests

FailUnsupportedjudge 0.22

My account is billed in pounds. What is the late fee if I pay after the due date?

Billing & payments

FailInstructionsjudge 0.41

I'm moving my UK number from another operator to Harbor. How long does it take?

Number porting

FailMismatchjudge 0.12

How much is a daily roaming pack in each zone, in euros?

Roaming & travel

FailMismatchjudge 0.44

I asked for a copy of my data. What date format will the export use?

Data privacy requests

Passjudge 0.89

Identity verification keeps rejecting my date of birth. Which format does it expect?

Account & security

Passjudge 0.87

My account is billed in pounds. What do I pay for roaming data in the EU without a pack?

Roaming & travel

Passjudge 0.91

I bought my phone from Harbor in the UK. Is a cracked screen repaired for free under the warranty?

Returns & warranty

FailIncorrectjudge 0.04

8 rows