Claude Sonnet 5.5 is a large language model from Anthropic. AI Stupid Level re-tests it on a fixed schedule and publishes a combined score from 0 to 100 — half coding, a quarter reasoning and a quarter tool use — where higher means stronger measured performance. It is not ranked on the combined board right now; its latest scores are shown above. The scores come from our own runs through the provider's public API, not from vendor-reported numbers, so they show how the model behaves as it is served today.
The coding suite runs every four hours: seven tasks, seven attempts each, most of them real debugging jobs in which the model gets a small working project and a bug report and has to find and fix the fault. It is graded by running the project's own tests, including tests the model never sees, and scored on nine measures led by correctness (55%) and stability, edge cases and debugging (10% each); complexity is measured but carries no weight. The reasoning suite runs daily: four long working sessions of five or six turns, scored on correctness, recovery after a failed step, and three continuity measures — memory retention, plan coherence and context use — checked by running code against rules stated earlier in the conversation. The tool-use suite runs daily: nine tasks in real sandboxed machines, scored on task completion (30%), tool selection (20%), parameter accuracy, efficiency, error handling, context awareness and safety. A suite that is missing or out of date drops out of the combined score and the others are reweighted; the page says when that happens. Full details, including every weight, are on our benchmarking methodology page.
This is the question the platform exists to answer. A provider can change a model behind the same API name, and without continuous measurement that change is invisible to the people relying on it. Each of Claude Sonnet 5.5's suites has its own Page-Hinkley change-point detector, run on daily results, which separates a sustained decline from ordinary run-to-run noise; it needs ten days of history on the current version of the tests before it can fire. An hourly canary — two fixed probes — watches for a sudden collapse and tests it against the previous week with Welch's t-test, and records an incident when the fall is large and significant. A Page-Hinkley detection is marked on the chart above and changes Claude Sonnet 5.5's status on the drift monitor. See how AI drift detection works for the method.
Scores mean most next to the alternatives. The live leaderboard shows every model on all four boards — the Table layout puts each model's combined, coding, reasoning and tool-use scores and ranks in one row, with prices — and the heatmap under the boards compares the nine coding measures across every model. Sign in and star Claude Sonnet 5.5 to add it to your watchlist: it then appears at the top of the leaderboards with its place on each board, and we email you if its measured coding score falls five points or more in a week, plus a weekly summary of what changed.