Skip to content

About Relevant

We help teams understand why their AI answers fail

Relevant is a small team of ML engineers, applied scientists and designers working remotely across European and US time zones. We build evaluation and failure-analysis software for LLM and retrieval systems, so a pass rate always arrives with an explanation.

Our mission

Every team shipping an AI assistant should be able to answer one question: why did this answer fail?

Teams that ship retrieval-augmented assistants usually measure accuracy early. What they cannot do is explain a drop: was the right passage never retrieved, was it retrieved and ignored, or did the answer break a rule the prompt already stated?

Relevant exists to answer that. It ingests traces and datasets, runs evaluations, attributes every failure to a cause and keeps the evidence next to the verdict, so improvements can be shipped with confidence.

The problem

A score is not an explanation

Aggregate metrics tell you that quality moved. They do not tell you which part of the system moved it, who should fix it, or whether the fix worked.

  • Retrieval and generation errors are entangled. A wrong answer can start in the index, the retriever, the prompt or the model.

  • Manual review does not scale. Reading transcripts finds problems, but not fast enough to keep up with releases.

  • Regressions ship silently. Without a gate, a prompt change that fixes one intent can quietly break another.

What one pass rate hides

86.9%

pass rate across 14,820 evaluated samples, with 1,944 open failures in six causes.

  • Unsupported claims525 · 27%
  • Incorrect factual answers408 · 21%
  • Missing context369 · 19%
  • Retrieval mismatch350 · 18%
  • Instruction-following failures214 · 11%
  • Other78 · 4%

From the dashboard: Support Assistant, last 30 days. Explore the failures

Principles

Four ideas that shape the product

They decide what we build, what we leave out and how the product behaves when the evidence is uncomfortable.

  1. 01

    Explain, do not only score

    A number that cannot be traced to its causes is a liability. Every metric links to the samples that moved it, and every failure carries a category, a cause and a suggested fix.

  2. 02

    Show the evidence

    Verdicts come with rationales, retrieved chunks and source articles. Engineers and reviewers can disagree with a judgment and see exactly why it was made.

  3. 03

    Keep people in the loop

    Ground truth is reviewed by people, judges are calibrated against human labels, and every change to a gold answer records a reason. Automation proposes; people decide.

  4. 04

    Respect the data

    Customer questions and answers are sensitive. Data is processed in the region you choose, and retention and deletion stay under your control.

How we work

Small team, clear writing, steady releases

  • Small and senior

    A compact team of engineers, scientists and designers who ship end to end and talk to the people who use the product.

  • Written first

    Decisions, metric definitions and designs are written down. Documentation and release notes are part of the product, not an afterthought.

    Read the changelog
  • Evidence over opinion

    We use the method we advocate: measure, categorize the failures, fix the cause and measure again. Our articles show the working.

    Read the blog
  • Remote by default

    One team across European and US time zones. We keep a few hours of overlap each day and work asynchronously the rest of the time, so focus time is protected.

The team

Four disciplines, one product

We work as one team across research, engineering, design and developer experience.

  • Applied research

    Metrics, judges and failure taxonomies: how to measure retrieval and generation, and how to know when a judge can be trusted.

    • Judge calibration
    • Failure taxonomy
    • Multilingual evaluation
  • Platform engineering

    The ingestion pipeline, SDKs and CI integrations that turn production traces and datasets into reliable evaluation runs.

    • SDKs
    • Regression gates
    • Regional processing
  • Product and design

    Workflows and interfaces that put a verdict, its evidence and its fix on one screen, for engineers and reviewers alike.

    • Review queues
    • Data-dense UI
    • Accessibility
  • Developer experience

    Documentation, templates, guides and support that help a platform team go from first upload to a CI gate.

    • Documentation
    • Templates
    • Support

Careers

Help us build the measurement layer for AI

We hire engineers and scientists who care about measurement and clear writing. Roles are remote-first, across European and US time zones.

careers@relevant.com.tr

Do not see your role? Write to careers@relevant.com.tr and tell us what you would build.

Talk to the people building it

Questions about your evaluation setup, enterprise requirements or how Relevant would fit your stack? Write to us and we will reply within one business day.