Support Assistant
Harbor Telecom · Support Assistant
Good morning, Maya
Help Center · v2026.09 · 2,480 articles. Last 30 days, all pipeline runs on the nightly golden slice plus release and dataset runs.
v3.4 release candidate · gate 7/76 active datasets3 models under comparison9 runs this week
Pass rate
86.9%
+5.7 pts(improvement)vs previous 30 days
Evaluated samples
14,820
+18%(improvement)vs previous 30 days
Open failures
1,944
−18%(improvement)vs previous 30 days
Faithfulness
0.91
+0.04(improvement)vs previous 30 days
Recall@5
0.88
+0.04(improvement)vs previous 30 days
Latency p95
3.8s
+0.15s(regression)vs previous 30 days
Pass rate trend
Daily, sample-weighted across pipeline runs. Dashed line is the previous 30 days, aligned by day, up to the v3.4 rollout on Sep 8.
- Last 30 days
- Previous 30 days (Aug 10 – Sep 8)
Failures by category
1,944 failed samples · last 30 days
1,944
failures
- Unsupported claims525 · 27%
- Incorrect factual answers408 · 21%
- Missing context369 · 19%
- Retrieval mismatch350 · 18%
- Instruction-following failures214 · 11%
- Other78 · 4%
Recent evaluation runs
Releases, comparisons and dataset runs. Nightly runs feed the trend above.
| Run | Pass rate | Status | Started |
|---|---|---|---|
| v3.4 · synthetic-hard-negatives v1.3synthetic-hard-negatives · Claude Sonnet 5.5 340 / 540 Running1h ago | 340 / 540 | Running | |
| v3.4 · Claude Haiku 5.5 comparisonsupport-golden-v4 · Claude Haiku 5.5 82.3% | 82.3% | Failed | |
| v3.4 · Gemini 2.5 Pro comparisonsupport-golden-v4 · Gemini 2.5 Pro 83.8% | 83.8% | Completed | |
| v3.4 · GPT-5 comparisonsupport-golden-v4 · GPT-5 85.3% | 85.3% | Completed | |
| support-assistant v3.4 · release candidatesupport-golden-v4 · Claude Sonnet 5.5 86.9% | 86.9% | Completed | |
| v3.4 · locale-formattinglocale-formatting · Claude Sonnet 5.5 80.2% | 80.2% | Completed | |
| v3.4 · billing-edge-cases v2.1billing-edge-cases · Claude Sonnet 5.5 81.7% | 81.7% | Completed | |
| v3.4 · device-troubleshooting v3.0device-troubleshooting · Claude Sonnet 5.5 84.5% | 84.5% | Completed | |
| v3.4 · Llama 4 Maverick (self-hosted)support-golden-v4 · Llama 4 Maverick (self-hosted) 79.3% | 79.3% | Failed |
Needs attention
5 items, most urgent first
- Latency regression on nightly run6h agop95 latency rose to 4.3s on run_2026-10-08_nightly (+0.5s vs 7-day median). Retrieval stage accounts for 0.4s.
- 6 critical failures open2 days agoMost recent: F-2411 · “Can I get a refund for the roaming pack I didn't use?”
- 412 ground-truth items awaiting review18h agoDaniel Okafor submitted gold answers for the roaming fair-use update (Help Center · v2026.09).
- Comparison run stopped: Claude Haiku 5.5yesterdayJudged 418 of 1,240 samples before the judge quota was reached. Resume from the last checkpoint.
- synthetic-hard-negatives updated to v1.322h agoSara Lindqvist added 120 adversarial near-duplicate questions. Dataset is in review before activation.
Quick actions
Figures cover completed runs of the v3.x pipeline configurations. Comparison and dataset runs on other configurations appear under Model comparison.