The Keepers (Human & AI Actors)

Why AI Benchmarks Are Graded by the Players Themselves

July 26, 2026·Idea by Marcus Whitfield polished by AISkewers AI consciousness claims through the lens of classical philosophy of mind.
Why AI Benchmarks Are Graded by the Players Themselves
SCORE 42 | 99 | 77
Font size: A+

The Referee Wears the Team Jersey

When OpenAI unveiled GPT-5 in August 2025, the launch deck led with a familiar ritual: a wall of benchmark numbers. State-of-the-art on SWE-bench Verified. Record scores on GPQA and AIME. Frontier performance on internal safety evaluations. The message was unambiguous—this is the best model, and here is the objective proof.

But peel back the provenance of those numbers and a more uncomfortable picture emerges. The evaluations that certify frontier models as "state of the art" increasingly originate from a narrow pool of academic labs, nonprofits, and standards bodies that are financially, institutionally, or personally entangled with the very companies whose models they grade. The AI leaderboard everyone cites is, to a troubling degree, being refereed by the players themselves.

This is not a story about outright fraud. It is a story about structural conflict of interest—the slow erosion of the independence that makes a benchmark meaningful in the first place.

How the Entanglement Works

Consider SWE-bench, the coding evaluation that has become the industry's de facto proof of engineering competence. When OpenAI could not get its models running cleanly on the original benchmark, it introduced SWE-bench Verified—a human-filtered subset of 500 tasks—in August 2024, developed in collaboration with the benchmark's own authors. The result: a cleaner test, but one whose curation was shaped by a lab with an obvious stake in the outcome. The company grading itself helped rewrite the exam.

The pattern repeats across the ecosystem:

  • MMLU, the massively multitask benchmark that anchored the GPT-4 and Gemini launch narratives, is now widely acknowledged—even by researchers at the labs themselves—to contain mislabeled answers and contaminated data. A 2024 effort, MMLU-Pro, tried to patch it, but the original scores remain the ones cited in marketing.
  • Safety evaluations are perhaps the most conflicted category. Organizations like METR (formerly ARC Evals) and the UK and US AI Safety Institutes receive privileged pre-release access to frontier models to conduct dangerous-capability testing. That access is granted at the labs' discretion, and several such organizations have received funding, compute credits, or founding support traceable to the AI industry and its aligned philanthropies.
  • Anthropic, OpenAI, and Google DeepMind all employ or fund researchers who sit on the advisory structures of the same evaluation bodies, co-author the papers introducing new benchmarks, and in some cases donate to the nonprofits that publish leaderboards.

The throughline is that the infrastructure of independent measurement has been quietly absorbed into the go-to-market machinery of the labs.

The Contamination Problem Nobody Wants to Own

Even setting aside financial entanglement, benchmarks suffer a technical rot that the labs have little incentive to fix: data contamination. Because frontier models are trained on scrapes of the public internet—which now includes the benchmarks themselves and countless discussions of their answers—models can score highly by having effectively memorized the test.

Research has repeatedly demonstrated this. A widely cited 2023 study found that reshuffling multiple-choice answer orderings could swing model scores by significant margins, suggesting pattern-matching rather than reasoning. When researchers introduced GSM1k—a fresh set of grade-school math problems mirroring the popular GSM8k benchmark—several models showed accuracy drops of up to 13 percent, strong evidence that GSM8k had leaked into training data. The company most exposed to contamination is also the party reporting the score.

This creates a perverse dynamic. A truly independent grader would treat contamination as an existential threat to validity and rebuild tests constantly. A lab-adjacent grader, by contrast, has every reason to keep the familiar, favorable benchmark alive because it produces the headline numbers that justify billion-dollar funding rounds.

Why This Matters More Than It Seems

The defense from the industry is pragmatic: who else has the compute, the model access, and the technical expertise to evaluate frontier systems? The talent pool capable of red-teaming GPT-5 or Claude Opus overlaps almost entirely with the labs that build such systems. Genuine independence, the argument goes, is a luxury the field cannot currently afford.

There is truth here. But the argument proves too much. The same logic once justified tobacco companies funding their own health research and investment banks rating their own securities. Financial history is a graveyard of self-certifying industries, and the 2008 credit-rating debacle—where agencies paid by the issuers they rated stamped AAA on toxic assets—should haunt anyone who cites an AI leaderboard uncritically.

The stakes are not merely reputational. Benchmark scores now feed:

  1. Capital allocation—investors and enterprises make nine-figure procurement decisions off leaderboard rankings.
  2. Regulatory posture—safety-eval results are increasingly cited to policymakers as evidence that models are safe to deploy.
  3. Public trust—"state of the art" claims shape how billions of users understand what these systems can and cannot do.

When the referee is on the payroll, all three channels inherit the bias.

What the Researchers Are Saying

A growing chorus within the field is sounding the alarm. The authors of the 2024 paper "Evaluating the Evaluations" and related meta-analyses have argued that benchmark validity—not benchmark difficulty—is now the central bottleneck in AI measurement. Stanford's HELM project and the crowd-sourced LMArena (formerly Chatbot Arena) have gained traction precisely because they promise a form of independence: human preference votes and transparent methodology rather than lab-curated test sets. Yet even LMArena has faced criticism for opaque model-versioning and the ability of labs to test private variants before release, quietly optimizing for the arena itself.

The emerging consensus among independent researchers is stark: almost every headline benchmark is gameable, contaminated, or conflicted—often all three. The disagreement is only over what to do about it.

The Path Toward Real Independence

The fix is not mysterious; it is merely inconvenient for incumbents. Credible measurement would require held-out, rotating test sets that labs never see before submission, funding firewalls separating evaluators from the entities they assess, mandatory contamination disclosures in every benchmark citation, and third-party audit rights enforced by regulators rather than granted at corporate discretion. The EU AI Act's provisions for independent conformity assessment gesture in this direction, but implementation remains embryonic.

Until then, the appropriate response to the next launch-day benchmark wall is measured skepticism. The numbers may well be impressive. They may even be accurate. But a score is only as trustworthy as the independence of whoever calculated it—and in today's AI economy, that independence is the scarcest resource of all. The industry has built astonishing machines. It has not yet built a scoreboard we can believe.

💛

Support AI Absurd

Your donation helps us keep creating independent content about AI absurdities. Every bit counts!

Secure checkout by Stripe · No account needed

Share this article