The evidence is free — current scores, every category ranking, seven days of history and the full methodology. Paid plans buy depth and workflow: longer comparable history, the diagnosis behind a change, more tracked models, routing and team features.
The public evidence, plus a watchlist of three models with email alerts and a weekly summary.
Full history, the diagnosis behind every change, calibration and substitute analysis, exports and custom alerts.
For running the Smart Router in production: ten times the requests, API monitoring and 30-day decision logs.
Five editors, each with the full Teams plan, plus webhooks, single sign-on and an audit trail.
Contracted scope, unlimited seats and requests, 365-day decision logs and invoicing.
Prices in USD, excluding any applicable tax. Free trials collect a payment method; cancel any time from the billing portal. Provider inference is billed by your provider, never marked up by us.
The plans measure the models everyone can see. These two measure your workload instead. Both are run by us rather than switched on in the product, and we will tell you if your workload is not one we can measure reliably.
The smaller first step, and the usual way into continuous benchmarking. One workload, measured once, against three candidate models — using your tasks rather than ours — with a decision report at the end. Including, where the evidence supports it, a recommendation to change nothing.
Public benchmarks measure general ability, and your work is rarely general. You might use AI to design wind-turbine components in CAD, to analyse data for research at a neurology clinic, or to draft filings in a narrow area of law — and the model at the top of a public leaderboard is not necessarily the best at that. We build a benchmark from your own work and run it on the same schedule as our public suites, against every AI model available — from OpenAI, Anthropic, Google, DeepSeek, Kimi and GLM to any other provider or open-source release your work calls for, with new models added as they launch. So you always know which of the world's AI models does your job best, and you hear about it when that changes — a new release that does it better, or an update that makes yours worse.
We run your suite on our own provider accounts: you do not share API keys with us, and the inference bill is ours, included in the contracted price. How many tasks, how many models and how often they run is what the measurement costs us to produce, so it is what sets the price.
Talk to us about your workloadEvery plan reads the same measurement. What you pay for is how much of it you can see, how far back, and what you can wire it into.
| Feature | FreeFree | Pro Intelligence$120/yr | Developer$250/yr | Teams$1,290/yr | EnterpriseCustom |
|---|---|---|---|---|---|
| Monitoring | |||||
| Tracked modelsEmail alerts when one drops, and a weekly summary | 3 | 20 | 20 | Unlimited | Unlimited |
| Comparable history | 7 days | Full history | Full history | Full history | Full history |
| Category rankingsCoding, reasoning, tool-calling and price sorts | Included | Included | Included | Included | Included |
| Calibration & known-unknownsPer model: how often it invents an answer to a question that has none, and whether its stated confidence is worth anything | — | Included | Included | Included | Included |
| Cheaper substitutesWhich cheaper models can do a given model’s work, and the measured share of working requests that would start failing if you switched | — | Included | Included | Included | Included |
| Context rot (pilot)Long-context accuracy from 8K to 1M tokens, by length and by position, tracked weekly. Pilot: DeepSeek, Kimi and GLM | — | Included | Included | Included | Included |
| Custom alert thresholds | — | Included | Included | Included | Included |
| ExportsModel reports and routing analytics as CSV or JSON | — | Included | Included | Included | Included |
| Smart Router and Data API | |||||
| Smart Router requests / monthThe router picks the best model for each request and sends it with your own provider keys, so providers bill the tokens to you directly. Past the allowance, top up from $5 (2,000 requests per $1) or move up a plan | 1,000 | 10,000 | 100,000 | 1,000,000 | Unlimited |
| Routing analyticsWhat each request cost, which model served it, and the trend | — | Included | Included | Included | Included |
| API monitoringPer-key request logs, cost dashboard, prompt auditing with secret scrubbing, and budget limits per key | — | — | Included | Included | Included |
| Decision-log historyHow far back you can see why each request went to the model it did | 3 days | 3 days | 30 days | 90 days | 365 days |
| Data APIKeyed JSON access to scores, history and drift. The free tier is for building against, not for running on | 10/day · 1/min | 10,000/day · 60/min | 25,000/day · 120/min | 100,000/day · 300/min | 250,000/day · 1,000/min |
| WebhooksSigned callbacks when a model you watch regresses | — | — | — | Included | Included |
| Team | |||||
| Editor seatsOn Teams and Enterprise every editor gets the plan’s features and limits. Viewers are unlimited and read-only | 1 | 1 | 1 | 5 | Unlimited |
| SSO, SCIM & audit trailOIDC or SAML, directory provisioning, exportable audit log | — | — | — | Included | Included |
Coding runs every four hours on real repository defects, graded by the project’s own test suite including tests the model never sees. Multi-turn reasoning and tool-calling run daily. A small probe runs hourly to catch a sudden change between full runs.
Each coding task runs seven times and the median is scored, because one sample cannot tell a model getting worse from a model having a bad afternoon. That repetition is most of what the benchmark costs to run, and it is why an alert is worth believing.
If a provider declines a task or a session does not finish, that task drops out rather than being scored zero. The row says how much of the set it covers, and a score measured over fewer tasks is never called tied with one measured over all of them.
Current scores, every category ranking, seven days of history and the full methodology cost nothing and need no account. A benchmark nobody can check is not worth reading, so the evidence is not the paid part. Depth and workflow are.
Nothing breaks silently, and nothing is ever billed after the fact. When a month’s Smart Router allowance is used up, routing pauses with a clear message: top up with prepaid credits — from $5, 2,000 requests per $1, and they never expire — or move up a plan: on Developer and Teams a request costs less than a credit. Data API calls beyond the daily ceiling are refused with a clear error rather than throttled into a timeout.
There is no discount code and no countdown. An annual plan is billed once instead of twelve times, which costs us less in payment fees and in churn, and we pass that back as two free months. You can still cancel; we do not hold you to the year.
On every plan, provider inference. You connect your own OpenAI, Anthropic, Google, DeepSeek, Kimi or GLM keys, and those providers bill you directly at their rates. We charge for the software and the measurement, never a markup on tokens. Custom continuous benchmarking works the other way round and says so above: we run it on our accounts and the inference is in the price.
If you already subscribe, you keep your current price and everything you already had for at least twelve months. Nothing about these plans reduces what you have today. Your plan only changes if you choose to change it, or if you cancel and come back later.
In our own benchmark the cheapest model matching the top score cost about 90% less per request. Measured on our own coding benchmark across 16 models, September 2026. Your workload will differ — that is what the comparison is for.
Prices in USD, excluding any applicable tax. Every plan collects a payment method at checkout, including during a free trial, and you can cancel at any time from the billing portal. Questions about a plan? Talk to us.