Skip to content

A failure taxonomy for support assistants

Six categories, clear detection signals and one owner per category: a failure taxonomy your support team will actually use.

Relevant teamProduct

8 min read

On this page
  1. Start from the fix, not the symptom
  2. Write detection signals for each category
  3. Keep 'Other' small and visible
  4. Review the taxonomy each quarter

A failure taxonomy is only useful if two reviewers put the same failure in the same bucket and if each bucket maps to someone who can fix it. Ours has six categories and is the one used in the Support Assistant project shown in the dashboard.

Start from the fix, not the symptom

Categories that describe symptoms ('bad answer', 'confusing') are easy to assign and impossible to act on. Categories that describe causes route work: retrieval problems go to whoever owns the index, generation problems to whoever owns the prompt and model.

  • Unsupported claims: tighten grounding, add numeric checks.
  • Incorrect factual answers: check versions, add contrastive examples.
  • Missing context: query expansion, larger k, fill knowledge-base gaps.
  • Retrieval mismatch: metadata filters, heading-aware chunking.
  • Instruction-following failures: rule checks, locale enforcement.
  • Other: over-refusal, truncation, wrong language without a rule.

Write detection signals for each category

Each category needs signals a judge or a rule can check. For unsupported claims, the signal is a sentence with no supporting chunk. For retrieval mismatch, it is a stale-version flag or a high score on the wrong article title.

Signals make the taxonomy auditable. When a reviewer disagrees with a label, the conversation is about whether the signal fired, not about taste.

Keep 'Other' small and visible

An 'Other' bucket is necessary and dangerous. Cap it: if it grows past 5% of failures, split out the largest recurring pattern into its own category. Over-refusal is the usual first candidate.

Review the taxonomy each quarter

Products change. When a new feature launches (eSIM transfers, a new roaming zone), add its failure patterns as clusters within existing categories first. Promote a cluster to a category only when it needs a different owner or fix.

Share
XLinkedIn
All articles
    • Ground truth
    • Process

    Human review that scales

    Queues, keyboard shortcuts and reason codes: how to keep ground truth fresh without building up a review backlog every week.

    6 min read

    • Model comparison
    • Evaluation

    Choosing a generator with evidence

    Comparing models on the same retrieval, the same prompt and the same questions, then reading the per-category differences.

    7 min read

    • Evaluation
    • Retrieval

    Why aggregate scores hide retrieval failures

    A single accuracy number can rise while the answers your customers care about get worse. Here is how to break the number apart.

    9 min read

See why your retrieval fails

Run an evaluation on your own dataset and read every failure with its retrieved chunks, flagged claims and suggested fix.

Free plan includes 2,000 evaluated samples a month