Pricing
Plans that scale with your evaluation volume
Start free with 2,000 evaluated samples a month. Move up when you need more volume, more seats or review workflows for your team.
Plans
Showing monthly prices.
Free
For individual engineers evaluating a first RAG pipeline.
Start free$0
Free, with no credit card
Includes
- 2,000 evaluated samples per month
- 1 seat and 1 project
- 3 datasets
- LLM judge with the default rubric
- Failure taxonomy and retrieval metrics
- 7-day data retention
Pro
For small teams running evaluations on every change.
Start with Pro$49/ month
Billed monthly
Includes
- 25,000 evaluated samples per month
- Up to 3 seats
- Unlimited datasets
- Custom rubrics and judge models
- Model and retriever comparison
- CI regression gates
- 30-day data retention
- Most popular
Team
For ML and support-quality teams sharing one evaluation loop.
Start with Team$149/ month
Billed monthly
Includes
- 150,000 evaluated samples per month
- Up to 10 seats
- Ground-truth review queues
- Human and LLM judges side by side
- Scheduled reports and CSV exports
- Role-based access control
- 90-day data retention
Enterprise
For organizations with residency, security and volume requirements.
Contact salesCustom
Annual contract, volume pricing
Includes
- Custom sample volume
- Unlimited seats
- SSO (SAML, OIDC) and SCIM
- Audit logs
- On-premises and VPC connectors
- Dedicated regional deployments
- Support SLA and named contact
Prices are in USD and exclude taxes. VAT or sales tax is calculated and added at checkout where it applies.
On every plan
- Dataset versioning
- Failure taxonomy and clustering
- Chunk-level retrieval attribution
- Processing in the region you choose
- Your data is never used to train models
Compare plans
Every feature, plan by plan
Volume, seats and retention scale with the plan. The diagnostic features that explain failures are on every plan.
| Feature | Free$0 | Pro$49/mo | Team$149/mo | EnterpriseCustom |
|---|---|---|---|---|
| Evaluation | ||||
| Evaluated samples per month | 2,000 | 25,000 | 150,000 | Custom |
| Projects | 1 | 5 | Unlimited | Unlimited |
| LLM judgesDefault rubric on Free; custom rubrics and judge models on paid plans | Default rubric | Custom | Custom | Custom + private models |
| Human review judges | Not included | Not included | Included | Included |
| CI regression gates | Not included | Included | Included | Included |
| Datasets | ||||
| Datasets | 3 | Unlimited | Unlimited | Unlimited |
| Dataset versioning | Included | Included | Included | Included |
| Ground-truth review queues | Not included | Not included | Included | Included |
| Analysis | ||||
| Failure taxonomy and clustering | Included | Included | Included | Included |
| Chunk-level retrieval attribution | Included | Included | Included | Included |
| Model and retriever comparison | Not included | Included | Included | Included |
| Scheduled reports | Not included | Not included | Included | Included |
| Collaboration | ||||
| Seats | 1 | 3 | 10 | Unlimited |
| Role-based access control | Not included | Not included | Included | Included |
| Security and support | ||||
| Data retention | 7 days | 30 days | 90 days | Custom |
| SSO (SAML, OIDC) and SCIM | Not included | Not included | Not included | Included |
| Audit logs | Not included | Not included | Not included | Included |
| On-premises and VPC connectors | Not included | Not included | Not included | Included |
| Support | Documentation and email | Email, 2 business days | Email, 1 business day | SLA and named contact |
| Start free | Start with Pro | Start with Team | Contact sales | |
Team
$149/moEvaluation
- Evaluated samples per month
- 150,000
- Projects
- Unlimited
- LLM judgesDefault rubric on Free; custom rubrics and judge models on paid plans
- Custom
- Human review judges
- Included
- CI regression gates
- Included
Datasets
- Datasets
- Unlimited
- Dataset versioning
- Included
- Ground-truth review queues
- Included
Analysis
- Failure taxonomy and clustering
- Included
- Chunk-level retrieval attribution
- Included
- Model and retriever comparison
- Included
- Scheduled reports
- Included
Collaboration
- Seats
- 10
- Role-based access control
- Included
Security and support
- Data retention
- 90 days
- SSO (SAML, OIDC) and SCIM
- Not included
- Audit logs
- Not included
- On-premises and VPC connectors
- Not included
- Support
- Email, 1 business day
Plan finder
Estimate your monthly volume
An evaluated sample is one question scored by one judge in one run. Estimate yours from the size of your dataset and how often you run it.
1,200 questions × 3 runs × 1 judge × 4.33 weeks
Estimated volume
15,600evaluated samples a month
- Free13,600 over 2,000
- ProRecommended62% of 25,000
- Team10% of 150,000
Pro covers this volume
$49 a month. After the allowance, extra samples cost $12 per 10,000 samples.
Start with ProAdd-ons
Add capacity without changing plans
Go past an allowance, or keep history for longer, without moving to the next tier.
Additional samples
$12 per 10,000 samples
Available on Pro and Team when you exceed the monthly allowance.
Additional seats
$19 per seat / month
Add seats on Team beyond the 10 included.
Extended retention
$99 / month
Keep runs, traces and reviews for 12 months on Team.
FAQ
Billing and plan questions
Cannot find your question? Contact us and we will reply within one business day.
Start with the free plan.
Upload a dataset, run your first evaluation and see which failures matter. Upgrade when your volume does.
No credit card. 2,000 evaluated samples a month on Free.
Free includes
- 2,000 evaluated samples a month
- Failure taxonomy and chunk-level attribution
- Dataset versioning and the default LLM judge rubric
- An upgrade path when your volume grows