claude-sonnet-4-5-20250929 Benchmark & Live Performance Score (2026) — Anthropic

CLAUDE SONNET 4 5 20250929
PROGRESS0%
Initializing...

What is the claude-sonnet-4-5-20250929 benchmark score?

claude-sonnet-4-5-20250929 is a large language model from Anthropic. AI Stupid Level benchmarks it continuously and publishes a combined score from 0 to 100, where a higher score means stronger measured performance. As of the latest run, claude-sonnet-4-5-20250929 scores 67/100, ranking #14 of 22 models we track, and its recent trend is improving. The score is recomputed from fresh benchmark runs rather than self-reported vendor numbers, so it reflects how the model behaves in production right now — not how it performed at launch.

How we test claude-sonnet-4-5-20250929

Every model runs the same three benchmark suites. A 7-axis code suite scores correctness, adherence to spec, code quality, efficiency, stability, refusal behaviour and error recovery. A deep reasoning suite measures multi-step problem solving, plan coherence, long-context retention and hallucination rate. A tooling suite measures tool selection, argument accuracy and recovery from failed calls. The headline score weights the code suite at 50% and the reasoning and tooling suites at 25% each. Full details are on our benchmarking methodology page.

Is claude-sonnet-4-5-20250929 getting worse over time?

This is the question the platform exists to answer. Model quality can shift after a provider updates a model behind a stable API name, and without continuous measurement that change is invisible to the people relying on it. We track claude-sonnet-4-5-20250929 for performance drift using CUSUM change-point detection, which separates a sustained decline from ordinary run-to-run noise. When claude-sonnet-4-5-20250929 degrades in a statistically meaningful way, it shows up on its chart above and in our drift alerts. See how AI drift detection works for the method behind it.

Compare claude-sonnet-4-5-20250929 with other models

Scores are only meaningful next to alternatives. Use the AI model comparison tool to put claude-sonnet-4-5-20250929 side by side with other current models on coding, reasoning, tool use, price and measured latency, or browse the live AI model leaderboard for the full ranking.