Every time a new model drops, everyone watches it top the leaderboards and races for the "world's strongest" headline. But almost no one stops to ask: who gave those scores, and why should we trust them? This piece argues that "which model is strongest" is the wrong question. The real one is who grades the AI, and is that judge trustworthy. When the exam gets saturated, leaked, gamed, and even the judge gets bought, the score loses its meaning — and the part that's actually scarce and valuable is quietly migrating somewhere else.
Let's throw out a question that's been asked to death: "Which model is strongest? Who's SOTA?"
That question is the wrong one, because its answer flips every few weeks and means less and less — frontier models have long been fighting over decimal points. What's actually worth asking, and reveals something real, is a different question:
When the "exam" itself can be gamed, leaked, and bought — can the score still be trusted? Who grades the AI?This isn't tech gossip. It's a question about power. Whoever controls the trusted standard for judgment holds the hidden throne of the AI era.
I'll walk you through four cracks splitting open — leaderboards getting saturated, leaked, gamed by the test-takers, and bought by the judges — then show you where the real value is migrating as public benchmarks lose credibility en masse. By the end you'll see this is the same logic as the AI "trust layer" I've written about before.
Why is this worth your fifteen minutes? Because it decides a very practical judgment: in the AI game, where does long-term, durable value actually settle? Most people put all their attention on the "model" layer — who has more parameters, who scores higher. But if models ultimately converge, commoditize, and depreciate like utilities, then betting your fortune on "guessing which model wins" is betting on something that gets rapidly commoditized. What's worth watching are the positions that don't depreciate as models iterate — and "the right to judge" is one of them: no matter how the models change, "deciding whether it's good and can be trusted" always needs someone to do it, and it only gets harder. See this layer, and your view of AI upgrades from "chasing new models" to "finding the real moat."
01Crack one: the leaderboards got maxed out (saturation)
First, what is a "leaderboard"? The AI field grades models with a pile of benchmarks — essentially sets of "standardized exams." For example, MMLU (multiple-choice across 57 subjects) and GPQA (graduate-level science problems). A model takes the exam, gets a score, and everyone ranks them by it.
The problem is that these exams are getting maxed out. By early 2026, the top frontier models all cluster around 90% on MMLU-Pro — Gemini 3 Pro ~90.1%, Claude Opus 4.5 ~89.5%. When every top contender scores above 88%, the gaps between them are statistically almost meaningless and useless for procurement decisions. Even wilder is GPQA Diamond — a set of "graduate-level" hard problems just two years ago — where frontier models now hit about 94.3%.
That's why frontier labs have quietly moved MMLU and HumanEval from the front page of launches to the appendix, switching to newer, harder exams: HLE (Humanity's Last Exam), FrontierMath, ARC-AGI 2, SWE-Bench Verified… leaderboards churn like an arms race. But it's the same medicine in a new bottle — as long as "the score" is the core marketing pitch, the next three cracks can't be plugged.
For you, the "user," saturation brings a very practical headache: you can no longer pick a model by leaderboard alone. Two years ago, MMLU rank could roughly screen out the weaker models; today, with frontier models all scoring above 90, the leaderboard tells you "they're all smart" but not "which one is more reliable for the specific thing you need to do." It's like every grad from a top school having a 3.99 GPA — once everyone has that number, it loses its discriminating power, and you have to find another way to judge who actually fits you. Remember this "loss of discriminating power" — it's the root driver of why value migrates to private evals later.
02Crack two: the exam leaked (contamination)
Even if the exam isn't saturated, there's a sneakier problem — the exam leaked. The term is data contamination.
How common is this? Studies comparing against a clean reference dataset found lots of "dirty data" in mainstream leaderboards: about 27–29% of MMLU questions appear contaminated, ARC-Challenge as high as 32.6%; and the Chinese-subject benchmark C-Eval is even worse, around 45.8% — nearly half the questions the model may have "already seen."
The scary part of contamination is that it requires no one to actively cheat — questions leak online, the model reads them, and the score quietly inflates. You think you're testing "reasoning," but a big chunk is testing "recall." That's why two models both scoring 89 on MMLU can differ wildly in real ability: one truly understands, the other may just have memorized more.
And contamination is a vicious cycle that worsens the more popular a benchmark gets: the more authoritative and cited an exam, the more its questions get compiled, discussed, and reposted online, and thus the more likely they enter the next generation's training data — the more famous it gets, the faster it contaminates itself. That's why benchmarks have ever-shorter shelf lives, and frontier labs must push a brand-new, not-yet-leaked exam every six months. But that raises a new problem: who writes the new exam, and who grades it? Concentrate the power to set and score exams in a few hands, and the third and fourth cracks are seeded.
03Crack three: the test-takers cheat (leaderboard gaming)
The first two cracks are "honest mistakes." The third is deliberate gaming. The classic case is the Llama 4 leaderboard scandal of April 2025.
First, a special "exam hall" — LMArena (formerly Chatbot Arena). It isn't a written test; it has real humans vote blind: the same question, two anonymous models each answer, you pick which is better, and a "popular vote" ranking emerges from mass voting. Because it tests real conversation quality and is hard to pre-memorize, it was considered one of the hardest-to-game, most trustworthy leaderboards.
Then Meta tripped. When it released Llama 4, the version it uploaded to LMArena was a custom build tuned to "please human preference", not the same thing as the version it open-sourced to the public. With that custom build, Llama 4 Maverick briefly beat GPT-4o and Gemini on the Arena. After it blew up, LMArena publicly released 2,000 raw battle records for audit; Meta denied "cheating" but admitted experimenting with chat-optimized builds. More awkwardly, per reports, Meta's departing chief AI scientist Yann LeCun later conceded the results were "fudged a little bit" and that "different model versions were used for different benchmarks to give better scores."
Why does this matter? Because LMArena was supposed to be "the hardest leaderboard to game" — real-time blind human votes, in theory unmemorizable. Yet Meta still found a way to game it: not by memorizing questions, but by training a version that's "especially likeable, emoji-loving, with long enthusiastic answers" to take the test, targeting the "human vote" mechanism itself. This exposes a deeper truth — as long as a leaderboard's ranking is valuable enough, test-takers will always find a technique aimed at it; plug one hole and another appears. It's not a flaw of one leaderboard; it's the inevitable result of "public leaderboard + huge marketing payoff."
04Crack four: the judge got bought (conflict of interest)
If the first three cracks are "flaws in the exam system," the fourth strikes at the root — the judge writing and grading the exam isn't neutral. The case is FrontierMath and OpenAI.
FrontierMath is an extremely hard math exam, assembled by a nonprofit called Epoch AI with a group of mathematicians, billed as a rigorous standard for AI math ability. OpenAI used it to showcase how strong its o3 model is. Sounds authoritative, right? Until people learned one thing: the exam was paid for by OpenAI — and OpenAI could see the questions.
Worse was how it was disclosed. OpenAI's funding was only quietly written in when the final paper was published (December 20, 2024); per reports, several contributing mathematicians didn't know OpenAI was the funder, let alone that OpenAI would have access to the questions. Epoch acknowledged the transparency issue, explaining "legal constraints" delayed disclosure, and stressed there was an informal agreement barring OpenAI from using the questions for training.
Note the subtlety: even if OpenAI truly never used a single question for training, the mere fact that "the funder can see the exam" is enough to discount the score — because once trust has to rest on "believe I didn't cheat," it's no longer objective evidence, it's a request. And that "no training" agreement was only informal — no contract, no audit, no third-party escrow. In a field where this swings hundred-billion-dollar valuation narratives, resting fairness on a gentleman's promise is far too light. FrontierMath's real lesson isn't "OpenAI is bad," but that it made everyone seriously think for the first time: next time I see an "authoritative leaderboard," I should first ask — who funded building this?
when the people paying can see the exam — how do you know the test was fair?Behind "who grades the AI" is another question: whose money is that judge taking? This is the neutrality problem of the right to judge.
05Capital eyes the right to judge: LMArena turns into a company
Since "who grades" is such power, capital won't leave it alone. The best example is the fate of that "popular-vote arena," LMArena itself.
It started as a UC Berkeley academic project (LMSYS) — pure research, pure non-profit. But because it became the industry's default "kingmaker" — whoever ranks first on the Arena gets to write "voted #1 by global users" in their launch — its value got noticed. So: in May 2025, LMArena raised a $100M seed at a $600M valuation, co-led by a16z and UC Investments; then in January 2026 it raised another $150M, with valuation soaring to $1.7B, total funding over $250M. An academic scoring project became a $1.7B unicorn in a little over a year.
Here a sharp tension appears: when the "kingmaker" itself becomes a VC-backed company chasing commercial returns, can it stay neutral? It took a16z's money, and a16z has invested in a slew of AI companies; when a portfolio company is ranked on its leaderboard, it's hard for outsiders not to wonder "are the judge and the contestant family?" This doesn't mean LMArena will favor anyone — it means it now sits in a structurally "conflicted" position, and neutrality that has to be maintained by "trust me, I won't cheat" is already shaky.
Connect sections three, four, and five and you find an unsettling pattern: who grades is a form of power, and all power gets contested and capitalized. Model companies want to game the leaderboard (Llama 4), funders want to influence the questions (FrontierMath), VCs want to buy the judge's bench (LMArena). When the right to judge becomes this valuable, "staying neutral" becomes the most counter-human, hardest-to-maintain option. That's why "public leaderboards" as trusted judges are destined to erode — not because any one person is bad, but because the incentive structure is just sitting right there.
06Where the real value is: migrating from "public leaderboards" to "private evals"
By now you might feel a bit grim: public leaderboards are saturated, contaminated, gamed, bought — what can you trust? Good news: precisely because public leaderboards are losing credibility en masse, the real value is being forced out. It's migrating to "private evals."
The logic is simple: a bank deploying an AI agent doesn't actually care "who has the highest MMLU." It cares "in my own customer-service conversations, my own compliance scripts, which model is least likely to hallucinate or cause trouble." No public leaderboard can answer that — because it needs a closed-door evaluation tailored to its own real business scenario.
This is growing into a real industry. A batch of companies build enterprise-grade eval tooling: Braintrust (making evals part of the engineering workflow), Patronus AI (founded in 2023 by ex-Meta researchers, detecting model hallucinations and safety risks), plus LangSmith, Galileo, DeepEval, and more. In March 2026, Patronus and Braintrust each shipped suites specifically for testing "AI agents." Their moats are clear too: data residency (your questions never leave your servers), game-proofing, and fit to real tasks — exactly what public leaderboards can't give.
Note this connects to my earlier AI-bubble piece: MIT's report said enterprise AI's real returns are in the back office, in vertical scenarios, not the chatbox. To use AI right and stably in vertical scenarios, step one is having your own trusted eval to judge whether this model actually works for you. So private evals aren't a nice-to-have — they're the "quality-control gate" for enterprises actually deploying AI, the precondition for end-user value to materialize. In other words, evals are shifting from "a researcher's small tool" to "infrastructure no company serious about AI can avoid." That's why a batch of eval startups can raise funding and tell an independent business story — they don't sell models, they sell the ability to "judge models." In a world flooded with models, "helping you pick right, and proving you picked right" is itself a hard requirement.
07Why evals are the AI "trust layer"
Now I can connect this back to the through-line I keep returning to.
When I wrote about the machine economy's "battle for the money layer," I made one call: models and protocols will converge, commoditize, and become free plumbing like HTTP; what's truly scarce and valuable is the layer above that "judges whether the thing can be trusted" — the trust layer. Evals are the core of that trust layer in the AI world.
but the power to judge "is this output good, can it be trusted, do I dare use it" is forever scarce.Whoever controls trusted, game-proof, task-fit evals holds the AI era's hidden right to judge — the same problem agents must solve: "can this behavior be trusted."
That's why I say "who grades the AI" is a badly underrated real question. In a world where model capability is highly convergent, what creates separation is no longer "who's stronger" but "who can credibly prove who's stronger, and a better fit for you." Once the right to judge is scarce, it gets contested, capitalized, and gamed like all scarce power — LMArena's $1.7B valuation, FrontierMath's conflict of interest, Llama 4's gaming are all footnotes to that contest.
One layer deeper, this also shapes how AI itself evolves. Models evolve "toward what they're tested on" — test it on X and it grows toward X. If the whole industry's "exams" are held by a few players with an agenda, AI's capabilities get quietly steered toward what those exams want, not necessarily what the real world needs most. So "who writes the questions" isn't just a fairness issue — it actually defines what "good AI" means, a far more profound power than "whose model is stronger." Whoever holds the eval standard quietly sets the North Star for the whole field.
One extension I'm personally most interested in — and will keep digging into: if "judge neutrality" is the core pain, this is exactly what crypto / mechanism design is best at. There's a term on-chain: "credible neutrality" — using cryptography and economic incentives so a set of rules doesn't depend on "trusting some institution to be fair," but is structurally unable to favor anyone. Port that thinking to AI evals — lock questions with commitment schemes, put grading results on-chain for audit, bind cheating by staking from both the question-setter and the tested party — and you might have the next stop for a "game-proof right to judge." That's another underrated intersection of AI and Crypto, which I'll write up separately.
08Conclusion: don't ask "who's strongest," ask "who scored it, and can that judge be trusted"
Let's close the board.
"Which model is strongest" is the wrong question, because it flips every few weeks and the frontier is bunched too tightly to separate. The real question is one of power: who grades the AI, and can that judge be trusted. Public leaderboards are being hollowed out by four simultaneous cracks — saturation (maxed ceiling), contamination (questions leaked into training data), gaming (submitting a custom build), conflict of interest (the funder can see the exam); meanwhile the "kingmaker" LMArena itself became a VC-backed $1.7B company, making neutrality even harder to self-certify. So the real value is migrating from "public leaderboards" to "closed-door private evals" — the quality-control gate for enterprises deploying AI, and the AI trust layer.
So next time you see "Model X tops the global leaderboard," don't get excited — watch three things far more useful than "the score":
Three judgment indicators more reliable than "the benchmark score"
- Watch which exam it uses: does the model card report old, maxed/contaminated benchmarks (MMLU, HumanEval), or newer, more game-proof ones (SWE-Bench Verified, HLE, ARC-AGI 2)? Be wary of anyone still headlining the old ones.
- Watch where the judge's money comes from: does this leaderboard / eval org have conflict disclosures? Is the funder one of the tested parties? Can the payer see the exam? FrontierMath already taught the lesson.
- Watch whether you have your own "private eval": if you're deploying AI into your business, trust no public leaderboard's ranking — use your own real data and your own standards to build a closed-door eval. That's the right to judge that truly belongs to you.
This is why I care so much about this dull-looking corner called "evals": the spotlight always lands on the highest-scoring model, but the ones who actually decide the rules of this game are the people standing in the shadow of the light, writing the questions, grading the answers, setting the standards. When you watch AI, don't only watch who's performing center stage — glance once more at who is grading the performance.
One line to close: in an era of converging models, the end of "building models" is "building the standard." Whoever controls the trusted right to judge is the one who, off the spotlight, actually calls the shots.