Skip to content

Pricing

Plans that scale with your evaluation volume

Start free with 2,000 evaluated samples a month. Move up when you need more volume, more seats or review workflows for your team.

Plans

Save ~17% with annual billing

Showing monthly prices.

  • Free

    For individual engineers evaluating a first RAG pipeline.

    $0

    Free, with no credit card

    Start free

    Includes

    • 2,000 evaluated samples per month
    • 1 seat and 1 project
    • 3 datasets
    • LLM judge with the default rubric
    • Failure taxonomy and retrieval metrics
    • 7-day data retention
  • Pro

    For small teams running evaluations on every change.

    $49/ month

    Billed monthly

    Start with Pro

    Includes

    • 25,000 evaluated samples per month
    • Up to 3 seats
    • Unlimited datasets
    • Custom rubrics and judge models
    • Model and retriever comparison
    • CI regression gates
    • 30-day data retention
  • Most popular

    Team

    For ML and support-quality teams sharing one evaluation loop.

    $149/ month

    Billed monthly

    Start with Team

    Includes

    • 150,000 evaluated samples per month
    • Up to 10 seats
    • Ground-truth review queues
    • Human and LLM judges side by side
    • Scheduled reports and CSV exports
    • Role-based access control
    • 90-day data retention
  • Enterprise

    For organizations with residency, security and volume requirements.

    Custom

    Annual contract, volume pricing

    Contact sales

    Includes

    • Custom sample volume
    • Unlimited seats
    • SSO (SAML, OIDC) and SCIM
    • Audit logs
    • On-premises and VPC connectors
    • Dedicated regional deployments
    • Support SLA and named contact

Prices are in USD and exclude taxes. VAT or sales tax is calculated and added at checkout where it applies.

On every plan

  • Dataset versioning
  • Failure taxonomy and clustering
  • Chunk-level retrieval attribution
  • Processing in the region you choose
  • Your data is never used to train models

Compare plans

Every feature, plan by plan

Volume, seats and retention scale with the plan. The diagnostic features that explain failures are on every plan.

Team

$149/mo

Evaluation

Evaluated samples per month
150,000
Projects
Unlimited
LLM judgesDefault rubric on Free; custom rubrics and judge models on paid plans
Custom
Human review judges
Included
CI regression gates
Included

Datasets

Datasets
Unlimited
Dataset versioning
Included
Ground-truth review queues
Included

Analysis

Failure taxonomy and clustering
Included
Chunk-level retrieval attribution
Included
Model and retriever comparison
Included
Scheduled reports
Included

Collaboration

Seats
10
Role-based access control
Included

Security and support

Data retention
90 days
SSO (SAML, OIDC) and SCIM
Not included
Audit logs
Not included
On-premises and VPC connectors
Not included
Support
Email, 1 business day
Start with Team

Plan finder

Estimate your monthly volume

An evaluated sample is one question scored by one judge in one run. Estimate yours from the size of your dataset and how often you run it.

1,200
10010,000
3
140
Judges per run

Each judge that scores a question counts as one evaluated sample.

1,200 questions × 3 runs × 1 judge × 4.33 weeks

Estimated volume

15,600evaluated samples a month

  • Free13,600 over 2,000
  • ProRecommended62% of 25,000
  • Team10% of 150,000

Pro covers this volume

$49 a month. After the allowance, extra samples cost $12 per 10,000 samples.

Start with Pro

Add-ons

Add capacity without changing plans

Go past an allowance, or keep history for longer, without moving to the next tier.

  • Additional samples

    $12 per 10,000 samples

    Available on Pro and Team when you exceed the monthly allowance.

  • Additional seats

    $19 per seat / month

    Add seats on Team beyond the 10 included.

  • Extended retention

    $99 / month

    Keep runs, traces and reviews for 12 months on Team.

FAQ

Billing and plan questions

Cannot find your question? Contact us and we will reply within one business day.

One question scored by one judge in one run. Re-running the same dataset counts again; viewing results, filtering and exporting do not.

Start with the free plan.

Upload a dataset, run your first evaluation and see which failures matter. Upgrade when your volume does.

No credit card. 2,000 evaluated samples a month on Free.

Free includes

  • 2,000 evaluated samples a month
  • Failure taxonomy and chunk-level attribution
  • Dataset versioning and the default LLM judge rubric
  • An upgrade path when your volume grows