Skip to content

run_2026-09-30_llama-4-maverick

v3.4 · Llama 4 Maverick (self-hosted)

Llama 4 Maverick (self-hosted) · bge-m3 · hybrid (BM25 + dense, RRF) · support-prompt v7 on support-golden-v4 v4.6

FailedExperimentv3.4 · Llama 4 Maverickv3 · 320 tokens, semantic
Compare

Pass rate

79.3%

−7.6 pts(regression)vs support-assistant v3.4 · release candidate165 of 208 samples passed

Faithfulness

0.86

−0.05(regression)vs support-assistant v3.4 · release candidate

Answer correctness

0.80

−0.08(regression)vs support-assistant v3.4 · release candidate

Context precision

0.79

0.00(no change)vs support-assistant v3.4 · release candidate

Context recall

0.84

0.00(no change)vs support-assistant v3.4 · release candidate

Latency p95

3.9s

+0.1s(regression)vs support-assistant v3.4 · release candidate

Failures by category

43 failed samples · largest share: unsupported

Share of failed samples by category, compared with support-assistant v3.4 · release candidate
CategorySharevs reference
Unsupported claims27.9%+0.7 pts
Incorrect factual answers20.9%−0.1 pts
Missing context16.3%−2.9 pts
Retrieval mismatch14.0%−3.9 pts
Instruction-following failures16.3%+5.2 pts
Other4.7%+0.9 pts

Run details

Dataset
support-golden-v4 v4.6 · 1,240 items
Generator
Llama 4 Maverick (self-hosted)
Retriever
bge-m3 · hybrid (BM25 + dense, RRF)
Prompt
support-prompt v7
Judge
LLM judge · Claude Opus 5.5 · rubric v3
Started
Sep 30, 08:20 UTC
Duration
45m
Triggered by
Owen Reid
Tokens / request
3,420 in · 242 out
  • self-hosted
  • comparison

All metrics

Against support-assistant v3.4 · release candidate

All metrics for v3.4 · Llama 4 Maverick (self-hosted)
MetricThis runReferenceChange
Pass rate79.3%86.9%−7.6 pts
Faithfulness0.860.91−0.05
Answer correctness0.800.88−0.08
Instruction adherence0.850.93−0.08
Context precision0.790.790.00
Context recall0.840.840.00
Recall@50.880.880.00
MRR0.740.740.00
NDCG@100.810.810.00
Latency p501.6s1.9s−0.3s
Latency p953.9s3.8s+0.1s
Cost per request$0.0019$0.0042−0.0023

Nightly trend

Nightly runs on the support-golden-v4 slice, Jul 11 to Oct 8. Dashed line marks this run.

  • Pass rate · nightly
  • This run · 79.3%
Pass rate moved from 78.5% to 87.2% across 90 nightly runs. Configuration changes: v3.3 on Aug 4, v3.4 on Sep 8.

Across datasets

v3.4 · Llama 4 Maverick on every dataset in the project. Rows marked “matrix” come from the comparison matrix rather than a dedicated run.

Configuration results per dataset
DatasetItemsPass rateFaith.Ctx recallRecall@5p95Source
support-golden-v4this run1,240Not evaluated with this configuration
billing-edge-cases312Not evaluated with this configuration
roaming-and-travel188Not evaluated with this configuration
device-troubleshooting426Not evaluated with this configuration
account-security154Not evaluated with this configuration
locale-formatting96Not evaluated with this configuration
synthetic-hard-negatives540Not evaluated with this configuration

Traced samples

17 samples with full traces (4 failing, 13 passing), stratified by domain. Open a row to inspect retrieval and the judge's reasoning.

Traced samples for v3.4 · Llama 4 Maverick (self-hosted)

What is the excess on a theft insurance claim?

Returns & warranty

Passjudge 0.92

What's included in my data access request?

Data privacy requests

FailUnsupportedjudge 0.29

What happens if I go over my fair-use limit on Unlimited at home?

Plans & upgrades

FailMissingjudge 0.07

How many lines can I add to a family plan?

Plans & upgrades

FailIncorrectjudge 0.07

My phone shows 4G but mobile data does not work after the iOS update.

Device troubleshooting

Passjudge 0.94

How do I get a copy of the personal data Harbor holds about me?

Data privacy requests

Passjudge 0.96

What is the late fee if I pay my bill after the due date?

Billing & payments

Passjudge 0.94

How long does it take to port my number to Harbor?

Number porting

Passjudge 0.94

Does the Fall 2026 upgrade offer work with a family plan discount?

Promotions

Passjudge 0.87

Can I return a phone I bought online if I opened the box?

Returns & warranty

Passjudge 0.91