A failure taxonomy is only useful if two reviewers put the same failure in the same bucket and if each bucket maps to someone who can fix it. Ours has six categories and is the one used in the Support Assistant project shown in the dashboard.
Start from the fix, not the symptom
Categories that describe symptoms ('bad answer', 'confusing') are easy to assign and impossible to act on. Categories that describe causes route work: retrieval problems go to whoever owns the index, generation problems to whoever owns the prompt and model.
- Unsupported claims: tighten grounding, add numeric checks.
- Incorrect factual answers: check versions, add contrastive examples.
- Missing context: query expansion, larger k, fill knowledge-base gaps.
- Retrieval mismatch: metadata filters, heading-aware chunking.
- Instruction-following failures: rule checks, locale enforcement.
- Other: over-refusal, truncation, wrong language without a rule.
Write detection signals for each category
Each category needs signals a judge or a rule can check. For unsupported claims, the signal is a sentence with no supporting chunk. For retrieval mismatch, it is a stale-version flag or a high score on the wrong article title.
Signals make the taxonomy auditable. When a reviewer disagrees with a label, the conversation is about whether the signal fired, not about taste.
Keep 'Other' small and visible
An 'Other' bucket is necessary and dangerous. Cap it: if it grows past 5% of failures, split out the largest recurring pattern into its own category. Over-refusal is the usual first candidate.
Review the taxonomy each quarter
Products change. When a new feature launches (eSIM transfers, a new roaming zone), add its failure patterns as clusters within existing categories first. Promote a cluster to a category only when it needs a different owner or fix.