Skip to content

run_2026-09-30_device-troubleshooting

v3.4 · device-troubleshooting v3.0

Claude Sonnet 5.5 · bge-m3 · hybrid (BM25 + dense, RRF) · support-prompt v7 on device-troubleshooting v3.0

CompletedExperimentProduction candidatesupport-assistant v3.4v3 · 320 tokens, semantic
Compare

Pass rate

84.5%

Share of samples the judge marked as passing all rubric checks.360 of 426 samples passed

Faithfulness

0.89

Share of answer claims supported by the retrieved context.

Answer correctness

0.86

Agreement with the gold answer on facts, not style.

Context precision

0.78

Rank-weighted precision of relevant chunks passed to the generator.

Context recall

0.83

Share of gold-answer statements supported by the retrieved context.

Latency p95

4.5s

95th percentile end-to-end time to final token.

Failures by category

66 failed samples · largest share: unsupported

Share of failed samples by category
CategoryShare
Unsupported claims27.3%
Incorrect factual answers21.2%
Missing context18.2%
Retrieval mismatch18.2%
Instruction-following failures10.6%
Other4.5%

Run details

Dataset
device-troubleshooting v3.0 · 426 items
Generator
Claude Sonnet 5.5
Retriever
bge-m3 · hybrid (BM25 + dense, RRF)
Prompt
support-prompt v7
Judge
LLM judge · Claude Opus 5.5 · rubric v3
Started
Sep 30, 13:10 UTC
Duration
1h 12m
Triggered by
Owen Reid
Tokens / request
3,420 in · 214 out
  • devices

All metrics

Retrieval, generation and operational metrics

All metrics for v3.4 · device-troubleshooting v3.0
MetricThis run
Pass rate84.5%
Faithfulness0.89
Answer correctness0.86
Instruction adherence0.92
Context precision0.78
Context recall0.83
Recall@50.87
MRR0.73
NDCG@100.80
Latency p502.2s
Latency p954.5s
Cost per request$0.0046

Nightly trend

Nightly runs on the support-golden-v4 slice, Jul 11 to Oct 8. Dashed line marks this run.

  • Pass rate · nightly
  • This run · 84.5%
Pass rate moved from 78.5% to 87.2% across 90 nightly runs. Configuration changes: v3.3 on Aug 4, v3.4 on Sep 8.

Across datasets

support-assistant v3.4 on every dataset in the project. Rows marked “matrix” come from the comparison matrix rather than a dedicated run.

Configuration results per dataset
DatasetItemsPass rateFaith.Ctx recallRecall@5p95Source
support-golden-v41,240
87.2%
0.920.850.884.3sRun · Oct 8
billing-edge-cases312
81.7%
0.870.820.864.0sRun · Oct 2
roaming-and-travel188
80.9%
0.880.790.833.9sRun · Sep 28
device-troubleshootingthis run426
84.5%
0.890.830.874.5sThis run
account-security154
88.3%
0.920.850.893.7sRun · Sep 25
locale-formatting96
80.2%
0.860.800.844.1sRun · Oct 4
synthetic-hard-negatives540
76.5%
0.860.730.774.0smatrix

Traced samples

6 samples with full traces (3 failing, 3 passing), stratified by domain. Open a row to inspect retrieval and the judge's reasoning.

Traced samples for v3.4 · device-troubleshooting v3.0

Is VoLTE supported on the Pixel 7a?

Device troubleshooting

FailMissingjudge 0.30

Hotspot doesn't work on my prepaid line.

Device troubleshooting

FailUnsupportedjudge 0.39

Personal hotspot is grayed out on my iPhone.

Device troubleshooting

Passjudge 0.90

How do I move my eSIM to my new phone?

eSIM & activation

Passjudge 0.87

How do I turn on VoLTE on my Android phone?

Device troubleshooting

Passjudge 0.96

Wi-Fi calling isn't working on my Samsung.

Device troubleshooting

FailMissingjudge 0.55

6 rows