Featured article
- Evaluation
- Retrieval
Why aggregate scores hide retrieval failures
A single accuracy number can rise while the answers your customers care about get worse. Here is how to break the number apart.
9 min read
Read articleBlog
Methods, metrics and process from the team building Relevant: how to measure RAG systems, explain their failures and ship changes with evidence.
7 articles
Featured article
A single accuracy number can rise while the answers your customers care about get worse. Here is how to break the number apart.
9 min read
Read articleComparing models on the same retrieval, the same prompt and the same questions, then reading the per-category differences.
7 min read
Six categories, clear detection signals and one owner per category: a failure taxonomy your support team will actually use.
8 min read
Block merges that make answers worse, without blocking every merge. Thresholds, sample sizes and what to do when a gate fails.
10 min read
On a telecom help center we evaluated, moving from 512-token fixed windows to 320-token semantic chunks cut chunk-boundary loss from 7.2% to 3.1%. Here is what changed.
7 min read
Currency, dates, decimal separators and language coverage: the failure patterns that only show up when assistants answer customers across regions.
8 min read
Queues, keyboard shortcuts and reason codes: how to keep ground truth fresh without building up a review backlog every week.
6 min read
Upload a dataset, run an evaluation and see every failure traced to a cause. The free plan includes 2,000 evaluated samples a month.
No credit card required