… visits today
ONLINE
Pricing

Know when your model decisions stop being right

The evidence is free — current scores, every category ranking, seven days of history and the full methodology. Paid plans buy depth and workflow: longer comparable history, the diagnosis behind a change, more tracked models, routing and team features.

Free

The public evidence, plus a watchlist of three models with email alerts and a weekly summary.

$0forever
No card required
  • 3 tracked models, with email alerts
  • 7 days of history
  • Every category ranking
  • Smart Router: 1,000 requests a month
  • Data API: 10 requests a day

Pro Intelligence

Full history, the diagnosis behind every change, calibration and substitute analysis, exports and custom alerts.

$120/ year
$10 a month · save $24
  • Everything in Free, plus
  • 20 tracked models and custom alert thresholds
  • Full history and the drift diagnosis behind every change
  • Calibration and cheaper-substitute analysis
  • Context-rot pilot: long-context accuracy up to 1M tokens
  • Routing analytics and CSV/JSON exports
  • Smart Router: 10,000 requests a month
  • Data API: 10,000 requests a day

Developer

For running the Smart Router in production: ten times the requests, API monitoring and 30-day decision logs.

$250/ year
$20.83 a month · save $50
  • Everything in Pro, plus
  • Smart Router: 100,000 requests a month
  • API monitoring: per-key logs, costs, prompt auditing and budget limits
  • 30-day decision logs
  • Data API: 25,000 requests a day

Enterprise

Contracted scope, unlimited seats and requests, 365-day decision logs and invoicing.

Talk to us

Prices in USD, excluding any applicable tax. Free trials collect a payment method; cancel any time from the billing portal. Provider inference is billed by your provider, never marked up by us.

Beyond the plans

The plans measure the models everyone can see. These two measure your workload instead. Both are run by us rather than switched on in the product, and we will tell you if your workload is not one we can measure reliably.

Workload assessment

$1,290one-off, fixed scope

The smaller first step, and the usual way into continuous benchmarking. One workload, measured once, against three candidate models — using your tasks rather than ours — with a decision report at the end. Including, where the evidence supports it, a recommendation to change nothing.

  • Up to 20 tasks you supply, three candidate models
  • A decision report within seven business days of us having what we need
  • Scope confirmed within two business days — or a full refund if we cannot measure your workload
  • Credited against an annual plan bought within 30 days, up to that plan’s price
  • Refunded if we cannot deliver the agreed report
Book an assessment

Custom continuous benchmarking

Contracted · priced on scope

Public benchmarks measure general ability, and your work is rarely general. You might use AI to design wind-turbine components in CAD, to analyse data for research at a neurology clinic, or to draft filings in a narrow area of law — and the model at the top of a public leaderboard is not necessarily the best at that. We build a benchmark from your own work and run it on the same schedule as our public suites, against every AI model available — from OpenAI, Anthropic, Google, DeepSeek, Kimi and GLM to any other provider or open-source release your work calls for, with new models added as they launch. So you always know which of the world's AI models does your job best, and you hear about it when that changes — a new release that does it better, or an update that makes yours worse.

  • A task set built from your real work, with the answer keys held out of the prompt
  • Run repeatedly against every model worth considering for your work, scored on the median of several trials
  • The same change detection the public board runs, with a baseline for your suite alone
  • Results through the alerts, webhooks, exports and Data API your plan already has
  • A scheduled review of what moved, and whether it is worth changing model
  • Your tasks stay yours: never published, never folded into the public corpus

We run your suite on our own provider accounts: you do not share API keys with us, and the inference bill is ours, included in the contracted price. How many tasks, how many models and how often they run is what the measurement costs us to produce, so it is what sets the price.

Talk to us about your workload

Compare plans

Every plan reads the same measurement. What you pay for is how much of it you can see, how far back, and what you can wire it into.

Swipe to compare →
FeatureFreeFreePro Intelligence$120/yrDeveloper$250/yrTeams$1,290/yrEnterpriseCustom
Monitoring
Tracked modelsEmail alerts when one drops, and a weekly summary32020Unlimited
Comparable history7 daysFull historyFull historyFull history
Category rankingsCoding, reasoning, tool-calling and price sortsIncludedIncludedIncludedIncluded
Calibration & known-unknownsPer model: how often it invents an answer to a question that has none, and whether its stated confidence is worth anything—IncludedIncludedIncluded
Cheaper substitutesWhich cheaper models can do a given model’s work, and the measured share of working requests that would start failing if you switched—IncludedIncludedIncluded
Context rot (pilot)Long-context accuracy from 8K to 1M tokens, by length and by position, tracked weekly. Pilot: DeepSeek, Kimi and GLM—IncludedIncludedIncluded
Custom alert thresholds—IncludedIncludedIncluded
ExportsModel reports and routing analytics as CSV or JSON—IncludedIncludedIncluded
Smart Router and Data API
Smart Router requests / monthThe router picks the best model for each request and sends it with your own provider keys, so providers bill the tokens to you directly. Past the allowance, top up from $5 (2,000 requests per $1) or move up a plan1,00010,000100,000Unlimited
Routing analyticsWhat each request cost, which model served it, and the trend—IncludedIncludedIncluded
API monitoringPer-key request logs, cost dashboard, prompt auditing with secret scrubbing, and budget limits per key——IncludedIncluded
Decision-log historyHow far back you can see why each request went to the model it did3 days3 days30 days365 days
Data APIKeyed JSON access to scores, history and drift. The free tier is for building against, not for running on10/day · 1/min10,000/day · 60/min25,000/day · 120/min250,000/day · 1,000/min
WebhooksSigned callbacks when a model you watch regresses———Included
Team
Editor seatsOn Teams and Enterprise every editor gets the plan’s features and limits. Viewers are unlimited and read-only111Unlimited
SSO, SCIM & audit trailOIDC or SAML, directory provisioning, exportable audit log———Included

What you are actually buying

Four suites, on a fixed schedule

Coding runs every four hours on real repository defects, graded by the project’s own test suite including tests the model never sees. Multi-turn reasoning and tool-calling run daily. A small probe runs hourly to catch a sudden change between full runs.

Repeated, then taken as a median

Each coding task runs seven times and the median is scored, because one sample cannot tell a model getting worse from a model having a bad afternoon. That repetition is most of what the benchmark costs to run, and it is why an alert is worth believing.

Gaps are shown, not filled

If a provider declines a task or a session does not finish, that task drops out rather than being scored zero. The row says how much of the set it covers, and a score measured over fewer tasks is never called tied with one measured over all of them.

Free is a real tier

Current scores, every category ranking, seven days of history and the full methodology cost nothing and need no account. A benchmark nobody can check is not worth reading, so the evidence is not the paid part. Depth and workflow are.

What happens at a limit

Nothing breaks silently, and nothing is ever billed after the fact. When a month’s Smart Router allowance is used up, routing pauses with a clear message: top up with prepaid credits — from $5, 2,000 requests per $1, and they never expire — or move up a plan: on Developer and Teams a request costs less than a credit. Data API calls beyond the daily ceiling are refused with a clear error rather than throttled into a timeout.

Why annual is cheaper

There is no discount code and no countdown. An annual plan is billed once instead of twelve times, which costs us less in payment fees and in churn, and we pass that back as two free months. You can still cancel; we do not hold you to the year.

What we do not charge for

On every plan, provider inference. You connect your own OpenAI, Anthropic, Google, DeepSeek, Kimi or GLM keys, and those providers bill you directly at their rates. We charge for the software and the measurement, never a markup on tokens. Custom continuous benchmarking works the other way round and says so above: we run it on our accounts and the inference is in the price.

Existing subscribers

If you already subscribe, you keep your current price and everything you already had for at least twelve months. Nothing about these plans reduces what you have today. Your plan only changes if you choose to change it, or if you cancel and come back later.

Does routing save money?

In our own benchmark the cheapest model matching the top score cost about 90% less per request. Measured on our own coding benchmark across 16 models, September 2026. Your workload will differ — that is what the comparison is for.

Prices in USD, excluding any applicable tax. Every plan collects a payment method at checkout, including during a free trial, and you can cancel at any time from the billing portal. Questions about a plan? Talk to us.