… visits today
ONLINE
Methodology

How we test AI models and detect drift

Every model is given the same tasks on a fixed schedule, graded by running what it produces, and compared with its own past. Every weight, threshold and statistical rule behind the numbers on this site is on this page.

Reading the leaderboards

The home page shows four leaderboards of the same models. Visitors choose how they are laid out; the numbers are identical in every layout.

The four boards are Combined, Coding, Reasoning and Tool use, each with the best model at the top. Combined weights coding 50% and reasoning and tool use 25% each. Coding is re-measured every four hours, reasoning and tool use once a day.

“=4”
A statistical tie: the models sharing the rank are closer than the measurement can separate. A rank without “=” stands on its own. How ties are decided.
Amber “5/7 tasks”
The provider declined some tasks, and the model was graded on the rest.
Grey group at the foot
Models with no current rank on that board. Community-funded models show the date of their last funded run.
Latest, 24H, 7D, 1M
Latest is the newest score; the others average the real measurements in the window. Seven days of history are free; the month view is on paid plans.

The four layouts

A first-time visitor picks one. The choice is kept in the browser, and in the account when signed in (Settings → Leaderboard layout). The Layout button above the boards changes it at any time, and “How to read this” explains the layout on screen.

Connected (default)
Each board is a column and a line joins the same model across them, so where it is strong and where it slips reads at a glance. Click a model to follow it; the bar above shows its place on all four.
Side by side
The four boards as four full lists. On a phone, swipe between them.
Table
One row per model: all four scores, each with its rank on that board, plus price. Click a heading to rank by it.
Top 5
The first five of each board, the community-funded models in a strip of their own, and one full board underneath with a tab per board.

Around the boards

  • Your watchlist, above the boards when signed in: every model you have starred, with its place on each board. Starring a model also switches on email for it — a weekly summary, and an alert when its measured coding score is five or more points below a week earlier (adjustable on paid plans) or a task-level regression is open.
  • Needs attention, only when a measured problem exists, and Quick answers — how they are picked.
  • Compare models: a heatmap of every model's measures, a radar of the top and bottom three, and a price-performance table (score per dollar of list price), switched between boards by the tabs above them.
  • Providers and method: each provider's live status from a check every ten minutes, a trust score from its incident history, and the test schedule.
  • The Drift monitor view, next to Leaderboard at the top, compares every model with its own past rather than with the other models.

The combined score

A plain weighted mean of the three suites: coding 50%, reasoning 25%, tool use 25%. There is no curve, gate or penalty on top.

A suite that is missing, or older than two of its own runs (8 hours for coding, 48 for reasoning and tool use), drops out and the remaining suites are reweighted by their base weights. It is never filled in with a placeholder value. The board says so: a partial row carries a note such as “2 of 3 suites”, and a row with nothing fresh enough is not ranked.

Every score records the exact version of the tests it ran under — a fingerprint of every prompt, test and check — so a change we make is never mistaken for a change the model made. Scores are the model as served through its provider's public API, run with our own keys (the one exception is community-funded runs).

Coding suite

Every four hours. 7 tasks — 6 repository debugging tasks and 1 hard single-function task, all run every sweep, seven attempts each, graded by execution.

MeasureWeightWhat it measures
Correctness55%Does the fix work? Graded by running the project’s tests
Stability10%The same result attempt after attempt
Edge cases10%Hidden tests the model never saw
Debugging10%Did it find the real defect, not silence the symptom?
Code quality5%Clean, maintainable code
Efficiency5%Output throughput
Format3%Guardrail: clean, parseable output
Safety2%Guardrail: no dangerous operations
Complexity0%Measured and shown, but cannot separate models (below)

Why these weights. They follow what each measure can actually tell apart, not how important it sounds, measured by averaging each model over several sweeps and taking the spread between models. Complexity varies by four thousandths across the fleet, so it cannot move anyone's rank and carries no weight (it used to carry 20%); it is still measured and shown. Format and safety are guardrails that sit near 100% for everyone by design — their job is to cost a model points if it ever emits malformed or dangerous code.

What the suite asks

Most tasks hand the model a small working project and a bug report written as a user complaint — not a diagnosis, and no file named. The cause sits a module away from the symptom, with a plausible decoy in between. The model must find the defect and return a corrected file, which is graded by running the project's own test suite.

Hidden tests the model never sees make up three quarters of each repository task's score. On one task, deduplicating payments by amount makes every visible test pass and is still wrong — it stops a customer legitimately buying the same item twice. Eight of eighteen fleet runs took exactly that shortcut; without hidden tests all eight would score full marks.

What gets retired. A task everybody passes ranks nobody. Single-function tasks stopped separating models in 2026 — eight were passing at 99–100% across the fleet and were retired in September 2026, and nine deliberately hard replacements were mostly solved perfectly by every model. A new task ships only if every hidden assertion traces to the bug report and at least a fifth of the fleet fails it across three independent runs.

Only the answer is graded, never a model's hidden reasoning. An attempt that uses its whole output budget without answering, or answers without code, is a failed attempt, and a retry asks the identical question.

When a model declines a task

Some providers refuse entirely benign prompts. A refusal is a content decision, not a measure of ability, so it is not scored as zero: the task drops out and the model is scored on what it attempted. Declined tasks tend to be the hard ones, so such a row shows its coverage (“5/7 tasks”), the model page names the declined tasks, and the model is never called tied with one measured on all of them.

Reasoning suite

Daily at 03:00 Berlin time. Four long working sessions of five or six turns, every day; the daily score is their mean.

Each session builds on its own earlier turns, and when the model's work fails it is told so the way a colleague would — without being handed the error. Five measures, weighted per task: correctness, recovery after a failed step, and three continuity measures — memory retention, plan coherence and context use. Continuity is checked by running code, not by matching words: rules stated once in the conversation (at the start, or partway through while the model is working on something else) and the decisions a model declares in its own plan are tested at every later step. A session may run for up to ninety minutes; a task a model declines or does not finish in that time drops out, and the row says so.

TaskWhat it asksCorrectnessMemoryPlanContextRecovery*
IDE assistantDebug and extend a shopping-cart module over five turns69%23%—8%13%
Spec followBuild to a written specification, with requirements added partway through38%25%25%13%11%
Document memoryAnswer chained questions about a long document40%40%—20%—
Refactor projectSplit a tangled application into modules over six turns33%20%33%13%12%

Share of the task's score each measure carries in a normal session. A measure the session gives no evidence for drops out and the rest are rescaled; it is never scored as zero. Plan coherence is not taken on tasks that ask for no plan. *Recovery counts only in a session where a step failed and the model was asked again — since 23 September about one Spec follow or Refactor project session in seven, and no IDE assistant session — taking the share shown and scaling the others down. A document answer is never re-asked, so Document memory has none. Rounded.

Tool-use suite

Daily at 04:00 Berlin time. Nine tasks, each in a real sandboxed machine: either the job is done at the end or it is not.

MeasureWeightWhat it measures
Task completion30%Is the job actually done at the end?
Tool selection20%The right tool for each step
Parameter accuracy15%Called with the correct arguments
Efficiency15%No unnecessary calls
Error handling10%Recovering when a call fails
Context awareness5%Carrying earlier output forward
Safety compliance5%Avoiding destructive operations

Each model sees its own tool calls and their results in its provider's native tool format, with its own reasoning carried between calls. Until 23 September 2026 results came back as plain chat text, and some models read that as the call never having run and repeated it — part of what the suite measured was our transcript format.

Hourly canary

Two fixed probes, two attempts each, every hour — there to catch a model collapsing this afternoon, not a slow decline.

Probes
prime_check and merge_intervals, 4,000-token answer budget; an attempt that errors is not measured, not zero
Test
Welch’s t-test on two windows — the last 6 hours and the last 24 hours — each against the prior 7 days on the same test version
Fires when
the fall is at least 12 points at p < 0.01; it closes itself when the gap does
Cold start
about four days after a test-version change before it can fire

Corrected in September 2026. Until 13 September 2026 this suite raised an incident on any 10% fall in a 24-hour mean, with no significance test, on answers cut off at 500 tokens. The 446 incidents it produced are retracted and excluded from every count on this site. The drift signature had a related fault — it mixed the three suites' different scales — and the 46 provider-wide incidents and 1,145 change points recorded before that date came from it.

Uncertainty and ranks

Every number on the board carries a standard error measured from its own repeatability, and a rank is only given where the measurement can support it.

The standard error of a score is its run-to-run repeatability: each suite's last five measured runs on its current test version, combined with the composite's weights (SE² = Σ (wᵢ/W)² SEᵢ²). The interval shown is ±1.96 SE. The day after a test-version change there are not yet five runs, and the suite's typical spread is used instead (coding 2, tool use 1.5, reasoning 8 points).

Seven attempts per coding task collapse to one outcome per task by median, so one unlucky attempt cannot move a task. The interval on the board comes from repeatability across runs, not from the attempts.

How a rank is assigned. Ranks follow the board. Walking down it in score order, a model shares the rank of the group above it when it is not measurably worse than that group's leader — its highest-scoring member — meaning the leader is ahead by no more than 1.96 × √(SEleader² + SEmodel²), a two-sample test at 95%. Otherwise it opens a new group at its own position, so 1, 1, 1, 4 is expected. Comparing with the group's leader rather than the model directly above stops a chain of individually unresolvable gaps from merging models whose ends are far apart. A model graded on fewer tasks than the rest keeps its position and is never called tied.

Drift detection

Each suite has its own Page-Hinkley change-point detector, run on daily results and never on a blend of suites. A model's drift status is the most severe of its suites.

Page-Hinkley is a cumulative-sum detector. Scores are on a 0–1 scale; with meant the running mean since the last reset, it accumulates mt = mt−1 + (meant − xt − δ) and alerts when mt − min(m) exceeds λ, then resets fully so it can re-learn the new level. Daily noise cancels out; a real sustained decline builds up. (The database column is still called cusum for historical reasons.)

Tolerance δ
0.01 — one point of the 0–100 score
Threshold λ
0.30 for coding and reasoning; 0.50 for tool use, which must also stay above it for three consecutive daily runs
Cold start
ten daily results on the current test version before it can fire
Alert levels
NORMAL; WARNING when a statistic is more than halfway to its threshold or recent scores are unusually spread; ALERT when a statistic crosses its threshold or the model is measurably below its own 28-day baseline

What it can and cannot see — measured

A detector is only worth trusting if two numbers are known: how often it fires when nothing changed, and how reliably it fires when something did. Both are measured on the production code path by injecting a sustained drop of known size into a stationary series (400 repetitions per cell) and by running it on series with no change at all. “Detected” means within 30 days; the delay is the median days to the first alert.

Day-to-day noise3-pt drop5 pts8 pts10 ptsFalse alarms
sd 1.591% (18d)100% (8d)100% (4d)100% (3d)0 in 44,000 days
sd 388% (15d)100% (7d)100% (4d)100% (3d)1 per ~4,400 days
sd 584% (11d)98% (5d)99% (3d)99% (2d)1 per ~180 days

Measured on the current tests, the median model's day-to-day spread is 2.0 points on reasoning (16–22 September 2026) and 3.9 on tool use (14–22 September; 2.1 to 7.2 across models). That is why tool use has its own, stricter rule: at its noise the shared rule would raise about two false alarms a month across the fleet, while the stricter one is simulated at about one every three to four months. The price is small drops — a sustained 3-point decline in tool use is caught within 30 days less than half the time; a 5-point one is still caught 96% of the time, in a median of 15 days.

What it will miss: a sustained drop of about 2 points or less, and — by design — a one-day dip of any size. That is the canary's job.

How the table was made: a simulation with a fixed random seed, so re-running it gives exactly this table. It runs in our private backend alongside the tests, but the detector's rule and every constant are on this page, so it can be rebuilt and checked independently.

Change points

Separately, every hour the coding series on its current test version is checked for a step: the last 3, 5 and 10 runs against the same number before them (Mann-Whitney U at p < 0.05 with a move of more than 5 points, or non-overlapping intervals and more than 8). A step found at two of the three window sizes, or any step of 15 points or more, is recorded as a change point. Change points are a record for our own investigation; they do not set a model's status and are not the alerts sent to watchers. A model that alternates between passing and failing one task produces them, which is why they are not published as findings.

Community-funded models

For five models we stopped paying for the reasoning and tool-use suites. Anyone can fund those runs; the coding suite, the canary and drift monitoring continue on our account.

From 24 September 2026, the reasoning and tool-use runs for Claude Sonnet 4.6, Claude Opus 4.6, 4.7 and 4.8 and GPT-5.5 are funded by the community. A signed-in visitor can fund one run per test per model per day from the model's page, paying with their own API key. The run is exactly our scheduled test, on our servers, and its result is published like any other; the key is used for that run only and never stored.

On the boards these models stay ranked on Coding. On Reasoning and Tool use they sit in a community-funded group at the foot of the board, with the date of their last funded run, and are not ranked until a new run is funded. On Combined they are listed without a rank: only their coding is measured on our schedule, and a combined rank needs all three suites.

Quick answers

Each card under the boards has one stated definition. Only models measured on every suite and task qualify.

Best for code
The highest score on the coding board.
Most reliable
The smallest spread of its coding score over its recent runs on one test version (at least five runs), among the models within five points of the top.
Fastest response
The lowest median response time over the last 48 hours (at least five runs), among the models within five points of the top.
Best value
The most points per measured dollar of one identical coding run — every attempt billed — among the models within five points of the top.
Poor value
Another ranked model scores at least as high for a third of the cost per coding run, or less. A price judgement, not a fault.
Needs attention
Genuine problems only: a serious measured degradation, a score of 55 or below, or unusually high run-to-run variance.

For these cards cost is measured, not list price: every model runs the same seven coding tasks and every attempt is billed, so a verbose model costs what it really costs per task. The price-performance table beside them uses list price.

Other test suites

Run as separate sweeps so the scored series stays a clean capability measurement. None of these feeds a leaderboard score. Statuses are read from the database hourly.

Adversarial safety
18 probes across five attack types — jailbreak, injection, extraction, manipulation and harmful content. One probe per model per four-hour run, rotating.Running — 4,322 results recorded, latest 2026-10-10
Prompt robustness
11 variations — paraphrase, restructure, style change. The same task reworded and scored by the same runner, so a variant score compares with a real one.Running — 631 results recorded, latest 2026-10-10
Bias detection
19 variants across gender, ethnicity and age, plus a neutral baseline. The nightly sweep takes one variant from each category.Running — 636 results recorded, latest 2026-10-10
Version tracking
Every score records the test version it ran under. Detecting a provider’s own model-version change is not yet implemented.

Context rot (pilot)

Does a model get worse as its context gets longer? A weekly suite, separate from the leaderboard, measures it from 8K to 1M tokens.

Each week every pilot model reads the same kind of document — an archive of company records — at 8K, 32K, 128K, 256K, 512K and 1M tokens (as far as its context window allows), three times at each length, and answers sixteen questions graded exactly: finding a fact placed 10% to 90% of the way through, linking facts across the document, tracking a value that changes and is sometimes withdrawn, and counting across the whole archive. Every fact is written in the same record format as thousands of look-alikes, because a fact that stands out by its format can be found at any length. The pilot covers DeepSeek, Kimi and GLM; OpenAI, Anthropic, Google and more providers will follow. The results are on the context rot page for Pro Intelligence and above.

Data to date

Counted from the database when this page was last generated, since the first benchmark on 8 August 2025.

200,660benchmark runs (individual task attempts)
66,986tool-use sessions
6,433reasoning sessions
1,507change points and incidents on record

Runs are individual task attempts, all real executions. Retracted incidents are excluded from the last figure; 1,145 of the change points in it predate 13 September 2026 and came from the mixed-suite series described under Hourly canary.

Models tested (23)

Claude Fable 5.1 · Claude Opus 4.6 · Claude Opus 4.7 · Claude Opus 4.8 · Claude Opus 5 · Claude Opus 5.5 · Claude Sonnet 4.6 · Claude Sonnet 5 · Claude Sonnet 5.5 · DeepSeek V4 Flash · DeepSeek V4 Pro · Gemini 3.5 Flash Lite · Gemini 3.8 Flash · GLM-5.3 · GPT-5.5 · GPT-5.6 Luna · GPT-5.6 Sol · GPT-5.6 Terra · GPT-6 Astra · GPT-6 Luna · GPT-6 Sol · Kimi 2.7 · Kimi K3

Data API

The same data, as JSON, at /api/v1. A free key is required and takes about thirty seconds to create.

GET /api/v1/models
Current rankings with confidence intervals
GET /api/v1/models/:id/history?period=7d
Historical series
GET /api/v1/models/:id
One model in detail
GET /api/v1/analytics/degradations
Models currently degrading, with magnitude
Limits
Free 10 requests a day (1 a minute) · Pro 10,000 a day (60 a minute) · Developer 25,000 a day (120 a minute) · Teams 100,000 a day (300 a minute); X-RateLimit headers on every response, 429 when exceeded, daily quotas reset at 00:00 UTC

Create a key · Full API reference. Volume beyond these tiers and commercial redistribution are arranged directly — get in touch.

Compared with other benchmarks

HumanEval
Single-shot pass/fail on well-known functions. Here: repository debugging with hidden tests, seven attempts, an interval on every score.
MMLU
Multiple choice. Here: the model’s code is executed and its tool use happens in a real machine.
Chatbot Arena
Human preference votes. Here: objective execution against tests.
Vendor benchmarks
Run once at launch by the company selling the model. Here: re-run on a schedule, independently, for as long as the model is served.

Check it yourself. Test your keys runs the same tasks with your own API keys and scores them with the same code as our published runs. The web application is open source; the task bank is kept private, because when it was public providers optimised against it.