Why we exist
Developers have long reported that models they rely on seem to get worse after launch: GPT-4 was widely described as “lazier” than it had been, and Claude as refusing more. Providers can change a model behind the same API name — fine-tuning, safety updates, routing, quantisation — without saying so, and nobody was systematically measuring it.
AI Stupid Level exists to close that gap. We ran our first benchmark on 8 August 2025 and have not stopped since. The reasons have not changed:
- Vendors don’t disclose changes. Updates and capability reductions happen without warning.
- Most benchmarks are snapshots. One measurement at launch, no error bars, and nothing watching for change afterwards.
- Choosing a model for production deserves evidence, not impressions.
- Accountability needs independence. Only someone with no stake in the result can keep an honest record.
Team
Ionut Adrian Visan is the Founder and CEO of AI Stupid Level. A technology entrepreneur and full-stack builder, he has spent his career building products across AI, software infrastructure, blockchain and real-time systems. At ASL, he leads the company’s vision of creating an independent reliability and intelligence layer for AI — continuously measuring how models perform, detecting meaningful changes over time, and helping organizations make better decisions about the AI systems they depend on.
Alexandra Chirilă, PhD, works at the intersection of philosophy, epistemology and AI evaluation. At AI Stupid Level, she contributes to the design and development of rigorous reasoning evaluations and to the broader question of how AI capabilities, reliability and risk should be measured and interpreted. Her work also spans Assurance 2.0 and safety-case review, bringing a critical perspective to how evidence about AI systems can support trustworthy real-world decisions.
Marius Răzvan Palimariu is the AI Infrastructure Lead at AI Stupid Level, bringing experience from IBM, NVIDIA and Nscale. He focuses on the infrastructure required to evaluate AI systems continuously and reliably at scale, from model execution and compute to the systems supporting ASL’s benchmarking and intelligence platform. His experience across large-scale AI and infrastructure environments helps ASL turn rigorous model evaluation into a dependable production system.
Funding and independence
Nobody who scores well on this site has paid us. That is the one line we will not cross.
- Vendor money
- None. No AI model provider funds us, and none of our investors is an AI model provider.
- Vendor relationships
- No financial relationship with OpenAI, Anthropic, Google, DeepSeek, Moonshot, Zhipu or any other model provider.
- Affiliate links
- None. We earn nothing from API sign-ups or referrals; a model’s place on a board is decided by the measurement alone.
- Who pays for the runs
- We do: scheduled benchmarks run on our servers with API keys we pay for. The one exception is labelled on the site — reasoning and tool-use runs for five community-funded models are paid for by visitors with their own keys, running the same test on the same servers.
- How we are funded
- Venture funding, which covers what revenue does not — benchmarking every model every four hours is not cheap; paid plans; and data licensing to organisations that are not model vendors.
How we keep it honest
The method is published. Every weight, statistical test and drift constant is on the methodology page and in the Public Benchmark Methodology (2026, PDF). If you think a weight is wrong, you can quote it back to us.
The tasks are not. When the task bank was public, providers optimised against it, and a test that can be studied in advance — or scraped into training data — stops measuring anything. The backend that holds the tasks, runners and scoring code stays private; the web application is open source, so the site you are reading can be checked line by line.
Every score is versioned. Each one records the exact version of the tests it ran under, so a change we make is never mistaken for a change the model made.
Corrections stay on the record. When we find a fault in our own measurement we say so and retract what it produced — 492 drift incidents in September 2026 — rather than quietly deleting it.
You can check it yourself. Test your keys runs the same tasks with your own API keys, scored by the same code as our published runs, and the Data API serves the same data as JSON with a free key.
Data and licensing
Counted from the database when this page was last generated, since the first benchmark on 8 August 2025.
Beyond the free site and the Data API, we license the underlying data in bulk to teams that need more, and only to organisations that are not model vendors. Everything listed is data we hold today; we do not sell datasets we have not collected.
- Performance time series
- Every score, per model, per measure and per suite, with the test version each run used — 200,000+ runs, with intervals and per-attempt variation.
- Drift and change points
- 1,500+ change points and incidents with the detector state behind each. The 492 retracted incidents are kept and flagged, and change points before 13 September 2026, which came from a mixed-suite series, are identifiable by date.
- Tool-use sessions
- 66,000+ full agent transcripts from sandboxed runs: which tools were called, with what parameters, and what happened.
- Reasoning sessions
- 6,400+ multi-turn sessions, scored turn by turn, with raw outputs where retention policy allows.
- Not yet offered
- Adversarial safety, bias and prompt robustness: the suites started recently and their datasets are still too young.
Custom extracts, bulk exports and support are available, with history back to 8 August 2025. If you need something we do not collect, say so — we would rather tell you it does not exist yet than sell you a promise. Licensing, pricing and contact.
Contact
- Questions
- Contact the team · FAQ

