About Relevant
We help teams understand why their AI answers fail
Relevant is a small team of ML engineers, applied scientists and designers working remotely across European and US time zones. We build evaluation and failure-analysis software for LLM and retrieval systems, so a pass rate always arrives with an explanation.
Our mission
Every team shipping an AI assistant should be able to answer one question: why did this answer fail?
Teams that ship retrieval-augmented assistants usually measure accuracy early. What they cannot do is explain a drop: was the right passage never retrieved, was it retrieved and ignored, or did the answer break a rule the prompt already stated?
Relevant exists to answer that. It ingests traces and datasets, runs evaluations, attributes every failure to a cause and keeps the evidence next to the verdict, so improvements can be shipped with confidence.
The problem
A score is not an explanation
Aggregate metrics tell you that quality moved. They do not tell you which part of the system moved it, who should fix it, or whether the fix worked.
Retrieval and generation errors are entangled. A wrong answer can start in the index, the retriever, the prompt or the model.
Manual review does not scale. Reading transcripts finds problems, but not fast enough to keep up with releases.
Regressions ship silently. Without a gate, a prompt change that fixes one intent can quietly break another.
86.9%
pass rate across 14,820 evaluated samples, with 1,944 open failures in six causes.
- Unsupported claims525 · 27%
- Incorrect factual answers408 · 21%
- Missing context369 · 19%
- Retrieval mismatch350 · 18%
- Instruction-following failures214 · 11%
- Other78 · 4%
From the dashboard: Support Assistant, last 30 days. Explore the failures
Principles
Four ideas that shape the product
They decide what we build, what we leave out and how the product behaves when the evidence is uncomfortable.
01
Explain, do not only score
A number that cannot be traced to its causes is a liability. Every metric links to the samples that moved it, and every failure carries a category, a cause and a suggested fix.
02
Show the evidence
Verdicts come with rationales, retrieved chunks and source articles. Engineers and reviewers can disagree with a judgment and see exactly why it was made.
03
Keep people in the loop
Ground truth is reviewed by people, judges are calibrated against human labels, and every change to a gold answer records a reason. Automation proposes; people decide.
04
Respect the data
Customer questions and answers are sensitive. Data is processed in the region you choose, and retention and deletion stay under your control.
How we work
Small team, clear writing, steady releases
Small and senior
A compact team of engineers, scientists and designers who ship end to end and talk to the people who use the product.
Written first
Decisions, metric definitions and designs are written down. Documentation and release notes are part of the product, not an afterthought.
Read the changelogEvidence over opinion
We use the method we advocate: measure, categorize the failures, fix the cause and measure again. Our articles show the working.
Read the blogRemote by default
One team across European and US time zones. We keep a few hours of overlap each day and work asynchronously the rest of the time, so focus time is protected.
The team
Four disciplines, one product
We work as one team across research, engineering, design and developer experience.
Applied research
Metrics, judges and failure taxonomies: how to measure retrieval and generation, and how to know when a judge can be trusted.
- Judge calibration
- Failure taxonomy
- Multilingual evaluation
Platform engineering
The ingestion pipeline, SDKs and CI integrations that turn production traces and datasets into reliable evaluation runs.
- SDKs
- Regression gates
- Regional processing
Product and design
Workflows and interfaces that put a verdict, its evidence and its fix on one screen, for engineers and reviewers alike.
- Review queues
- Data-dense UI
- Accessibility
Developer experience
Documentation, templates, guides and support that help a platform team go from first upload to a CI gate.
- Documentation
- Templates
- Support
Careers
Help us build the measurement layer for AI
We hire engineers and scientists who care about measurement and clear writing. Roles are remote-first, across European and US time zones.
Do not see your role? Write to careers@relevant.com.tr and tell us what you would build.
Talk to the people building it
Questions about your evaluation setup, enterprise requirements or how Relevant would fit your stack? Write to us and we will reply within one business day.