Skip to content

run_2026-09-25_account-security

v3.4 · account-security v1.2

Claude Sonnet 5.5 · bge-m3 · hybrid (BM25 + dense, RRF) · support-prompt v7 on account-security v1.2

CompletedExperimentProduction candidatesupport-assistant v3.4v3 · 320 tokens, semantic
Compare

Pass rate

88.3%

Share of samples the judge marked as passing all rubric checks.136 of 154 samples passed

Faithfulness

0.92

Share of answer claims supported by the retrieved context.

Answer correctness

0.89

Agreement with the gold answer on facts, not style.

Context precision

0.80

Rank-weighted precision of relevant chunks passed to the generator.

Context recall

0.85

Share of gold-answer statements supported by the retrieved context.

Latency p95

3.7s

95th percentile end-to-end time to final token.

Failures by category

18 failed samples · largest share: unsupported

Share of failed samples by category
CategoryShare
Unsupported claims27.8%
Incorrect factual answers22.2%
Missing context16.7%
Retrieval mismatch16.7%
Instruction-following failures11.1%
Other5.6%

Run details

Dataset
account-security v1.2 · 154 items
Generator
Claude Sonnet 5.5
Retriever
bge-m3 · hybrid (BM25 + dense, RRF)
Prompt
support-prompt v7
Judge
LLM judge · Claude Opus 5.5 · rubric v3
Started
Sep 25, 15:20 UTC
Duration
25m
Triggered by
Maya Collins
Tokens / request
3,420 in · 214 out
  • security

All metrics

Retrieval, generation and operational metrics

All metrics for v3.4 · account-security v1.2
MetricThis run
Pass rate88.3%
Faithfulness0.92
Answer correctness0.89
Instruction adherence0.94
Context precision0.80
Context recall0.85
Recall@50.89
MRR0.75
NDCG@100.82
Latency p501.9s
Latency p953.7s
Cost per request$0.0042

Nightly trend

Nightly runs on the support-golden-v4 slice, Jul 11 to Oct 8. Dashed line marks this run.

  • Pass rate · nightly
  • This run · 88.3%
Pass rate moved from 78.5% to 87.2% across 90 nightly runs. Configuration changes: v3.3 on Aug 4, v3.4 on Sep 8.

Across datasets

support-assistant v3.4 on every dataset in the project. Rows marked “matrix” come from the comparison matrix rather than a dedicated run.

Configuration results per dataset
DatasetItemsPass rateFaith.Ctx recallRecall@5p95Source
support-golden-v41,240
87.2%
0.920.850.884.3sRun · Oct 8
billing-edge-cases312
81.7%
0.870.820.864.0sRun · Oct 2
roaming-and-travel188
80.9%
0.880.790.833.9sRun · Sep 28
device-troubleshooting426
84.5%
0.890.830.874.5sRun · Sep 30
account-securitythis run154
88.3%
0.920.850.893.7sThis run
locale-formatting96
80.2%
0.860.800.844.1sRun · Oct 4
synthetic-hard-negatives540
76.5%
0.860.730.774.0smatrix

Traced samples

6 samples with full traces (2 failing, 4 passing), stratified by domain. Open a row to inspect retrieval and the judge's reasoning.

Traced samples for v3.4 · account-security v1.2

How quickly is a line blocked after reporting an unauthorized SIM swap?

Account & security

Passjudge 0.90

Someone swapped my SIM without my permission. How fast can you block it?

Account & security

FailIncorrectjudge 0.08

Can the 2FA requirement be removed over chat if I lose my phone?

Account & security

Passjudge 0.91

I lost my phone with my 2FA app. How do I sign in to My Harbor?

Account & security

Passjudge 0.93

Someone called pretending to be Harbor and asked for my code. What do I do?

Account & security

Passjudge 0.91

Do I need to keep my old SIM active while porting my number in?

Number porting

FailMismatchjudge 0.03

6 rows