Claude Opus 4.6 Benchmark & Live Performance Score (2026) — Anthropic

CLAUDE OPUS 4.6
PROGRESS0%
Initializing...

What is the Claude Opus 4.6 benchmark score?

Claude Opus 4.6 is a large language model from Anthropic. AI Stupid Level re-tests it on a fixed schedule and publishes a combined score from 0 to 100 — half coding, a quarter reasoning and a quarter tool use — where higher means stronger measured performance. Claude Opus 4.6's reasoning and tool-use tests are funded by the community rather than by us, so its combined score (88) is its coding score alone and is not ranked against models measured on all three suites. On coding, which we still test every four hours, it is joint 14th of 22. Its most recent funded results are reasoning 100 (last funded run, 24 Sep) and tool use 78 (last funded run, 24 Sep). Anyone can fund its next reasoning or tool-use run from this page, with their own API key. The scores come from our own runs through the provider's public API, not from vendor-reported numbers, so they show how the model behaves as it is served today.

How we test Claude Opus 4.6

The coding suite runs every four hours: seven tasks, seven attempts each, most of them real debugging jobs in which the model gets a small working project and a bug report and has to find and fix the fault. It is graded by running the project's own tests, including tests the model never sees, and scored on nine measures led by correctness (55%) and stability, edge cases and debugging (10% each); complexity is measured but carries no weight. The reasoning suite runs daily: four long working sessions of five or six turns, scored on correctness, recovery after a failed step, and three continuity measures — memory retention, plan coherence and context use — checked by running code against rules stated earlier in the conversation. The tool-use suite runs daily: nine tasks in real sandboxed machines, scored on task completion (30%), tool selection (20%), parameter accuracy, efficiency, error handling, context awareness and safety. A suite that is missing or out of date drops out of the combined score and the others are reweighted; the page says when that happens. For Claude Opus 4.6, the reasoning and tool-use suites run only when a member of the community funds a run; the coding suite runs on our schedule, as for every model. Full details, including every weight, are on our benchmarking methodology page.

Is Claude Opus 4.6 getting worse over time?

This is the question the platform exists to answer. A provider can change a model behind the same API name, and without continuous measurement that change is invisible to the people relying on it. Each of Claude Opus 4.6's suites has its own Page-Hinkley change-point detector, run on daily results, which separates a sustained decline from ordinary run-to-run noise; it needs ten days of history on the current version of the tests before it can fire. An hourly canary — two fixed probes — watches for a sudden collapse and tests it against the previous week with Welch's t-test, and records an incident when the fall is large and significant. A Page-Hinkley detection is marked on the chart above and changes Claude Opus 4.6's status on the drift monitor. See how AI drift detection works for the method.

Compare Claude Opus 4.6 with other models

Scores mean most next to the alternatives. The live leaderboard shows every model on all four boards — the Table layout puts each model's combined, coding, reasoning and tool-use scores and ranks in one row, with prices — and the heatmap under the boards compares the nine coding measures across every model. Sign in and star Claude Opus 4.6 to add it to your watchlist: it then appears at the top of the leaderboards with its place on each board, and we email you if its measured coding score falls five points or more in a week, plus a weekly summary of what changed.