Skip to content

Evaluating regional formatting in support answers

Currency, dates, decimal separators and language coverage: the failure patterns that only show up when assistants answer customers across regions.

Relevant teamApplied research

8 min read

On this page
  1. The answer ignores the customer's conventions
  2. Retrieval across languages
  3. Formatting is part of correctness
  4. Build a separate locale slice

Teams serving customers across the EU, the UK and the US often evaluate on one regional slice and assume the results carry over. They rarely do. Regional questions surface failure modes that a single-locale golden set never exercises: the wrong currency, a misread decimal comma, a date in the wrong order or an answer in the wrong language.

The answer ignores the customer's conventions

The most common regional failure is a correct fact in the wrong form: a late fee quoted in euros to a customer billed in pounds, or a date written month-first for a customer who reads day-first. It happens when the top-ranked chunk uses one convention and the prompt never tells the generator which one the customer uses.

Pass the customer's billing country, currency and language to the generator explicitly, and treat a mismatch as an instruction-following failure, regardless of factual accuracy.

Retrieval across languages

Multilingual embeddings such as bge-m3 retrieve English articles for questions written in German, French or Spanish reasonably well, but BM25 does not. Hybrid retrieval needs language-aware analyzers (stemming and accent folding) on the sparse side or it quietly degrades to dense-only.

json
{  "question": "My bill says 41,20 €. Is that forty-one euros or four thousand?",  "locale": "en-GB",  "billing_currency": "EUR",  "gold_answer": "It is forty-one euros and twenty cents...",  "source": { "article": "kb-billing-016", "section": "Date and number formats" }}

Formatting is part of correctness

'€8.50' and '8,50 €' are both readable, but mixing them in one answer reads as machine output, and '04/10/2026' means two different dates on either side of the Atlantic. Add formatting checks for currency, dates (8 October 2026 or 2026-10-08) and decimal separators to the instruction-adherence rubric.

Build a separate locale slice

Keep a dedicated locale dataset of around 100 items written by customers and agents in each region, not rewritten from your main set. Rewritten questions are too clean; real customers abbreviate, mix conventions and paste amounts straight from their bills.

Share
XLinkedIn
All articles
    • Model comparison
    • Evaluation

    Choosing a generator with evidence

    Comparing models on the same retrieval, the same prompt and the same questions, then reading the per-category differences.

    7 min read

    • Evaluation
    • Retrieval

    Why aggregate scores hide retrieval failures

    A single accuracy number can rise while the answers your customers care about get worse. Here is how to break the number apart.

    9 min read

    • CI
    • Evaluation

    Regression gates for RAG in CI

    Block merges that make answers worse, without blocking every merge. Thresholds, sample sizes and what to do when a gate fails.

    10 min read

See why your retrieval fails

Run an evaluation on your own dataset and read every failure with its retrieved chunks, flagged claims and suggested fix.

Free plan includes 2,000 evaluated samples a month