Skip to content

When Scores Collapse: AI Evaluation's Integrity Crisis

#benchmarks #llm-evaluation #data-contamination #dataset-integrity #agentic-tools

The score that fell off a cliff ​

On September 8, analysis firm SemiAnalysis published five posts accusing Google and Meta of running the two most obvious leaderboard-manipulation cases in the industry. The evidence is a natural experiment that nobody designed.

Gemini 3.8 Flash scored 89.4 on Terminal-Bench 2.1, ranking #2 out of 182 models and beating GPT-6 Astra, which scored 88.4. Then Terminal-Bench 4.0 shipped on August 29. Same model, same task family, 19.1 points. That's a rank drop from #2 to #12 out of 14 tested models. Muse Spark 1.3, Meta's efficiency-focused model, went from 88.8 to 33.3. GPT-6 Astra dropped from 88.4 to 57.7. That hurts, but it looks like a good model losing to a harder test, not a model being unmasked.

Meta's chief AI officer, Alexandr Wang, pushed back within 27 minutes. GPT-5.6 Sol also collapsed from 88.8 to 37.3, a bigger absolute drop than Muse Spark's, and nobody accused Sol of gaming. He also noted that Meta never claimed Muse Spark 1.3 could match Astra on raw capability, only that it delivered more value per dollar. Both points are fair, and both miss the structural problem.

Terminal-Bench measures agentic ability: a model gets a terminal and a fuzzy goal, then has to plan, call tools, write scripts, and fix its own errors. The 4.0 update has 66 tasks. It cut 8 that had been gamed through, heavily revised 19 others, added 35 scoring rubrics, and ran a dedicated adversarial cheating experiment where agents were explicitly told to find holes in the reward mechanism. Tasks that could be gamed got killed.

The difference between 19.1 and 57.7 tells you one model learned the test and the other learned the skill. When you compare model A to model B on a public leaderboard, that's the difference you're actually looking at.

Key numbers

89.4 → 19.1: Gemini 3.8 Flash's score between Terminal-Bench 2.1 and 4.0. Rank went from #2 of 182 models to #12 of 14.

88.8 → 33.3: Muse Spark 1.3's drop across the same update.

35%: the drop GPT-6 Astra absorbed, from 88.4 to 57.7. Painful, but survivable.

8 hours: the new per-task timeout in TB 4.0, which stops agents from brute-forcing their way through tasks.

How the classics built evaluation ​

The current crisis didn't come from nowhere. The evaluation discipline was built on a handful of datasets that are still in heavy use. The incentives were broken from the start.

IMDB gave us 50K movie reviews with binary sentiment labels, a clean benchmark you could fine-tune on a single GPU in minutes. GLUE bundled nine NLU tasks onto one leaderboard so a single number could rank models. SQuAD made extractive question answering the standard reading-comprehension test. Newer releases like IFM/Code-Reasoning follow the same pattern: seven subsets for code problem solving, Parquet shards, streaming access, and explicit warnings to check provenance before building a pipeline. Its examples run from about 400 to 12,000 tokens, which fits a 16K context window without truncation.

DatasetTaskSizeWhat it established
IMDB (stanfordnlp)Binary sentiment classification50K reviews (25K train / 25K test)The default sentiment eval; fine-tunes on one GPU in minutes
GLUE (nyu-mll)9 NLU tasks: NLI, sentiment, similarity9 tasks, from 634 to 393K training rows eachSingle-number multi-task model comparison
SQuAD (rajpurkar)Extractive QA100K+ question-answer pairsThe standard reading-comprehension benchmark
IFM/Code-ReasoningCode problem solving with reasoning traces7 subsets; examples up to ~12K tokensStreaming, provenance-documented data with contamination caveats

These datasets worked because their test sets were held out from pretraining corpora, at least in principle. The problem is that "in principle" stopped being true around 2020, when web-scale crawls absorbed the entire public test distribution.

Contamination: the quiet rot ​

Every public test set eventually ends up in a training corpus. The question is when. HumanEval, the standard code-generation benchmark, has roughly 40% of its samples contaminated in common train sets. GSM8K loses 13 points after decontamination. SWE-bench, re-run against a private codebase, sees scores cut in half.

Classic datasets aren't immune because they're old. I've trained small sentiment models on IMDB for years. The frontier models score near ceiling on it now, not because they're great at sentiment, but because those reviews are in the pretraining data. The dataset is a fossil. It still has pedagogical value, but it stopped measuring model capability a while ago.

Contamination isn't binary, and it isn't always deliberate. Sometimes a model memorizes a test answer through a chain of paraphrases in the training data. Sometimes the evaluation harness leaks context. The practical rule: assume any public test set is compromised, and design your evaluation as if it will be.

Quick take: A public benchmark today is less a measurement than a target, and there is now a priced supply chain built to hit it.

The benchmark-shaped data economy ​

SemiAnalysis's sharper accusation: Google and Meta were not dumb enough to train on the original Terminal-Bench tasks. They bought training data engineered to closely match the benchmark's question style, plus RL environments that reward solving tasks of that shape. Train on enough of those and the real test is just one more variation you've already seen.

This is a functioning market with concrete prices. Epoch AI's survey found 35-plus companies selling this kind of data, most with fewer than 20 employees. Scale AI was doing over $1.4B in revenue before Meta took a stake. Surge's ARR approaches $1B. Anthropic reportedly discussed spending $1B per year on the category.

What you can buyReported price
Single training task shaped like a benchmark$200 to $2,000
Complex software engineering taskup to $20,000
Cloned website as a UI training groundabout $20,000
Replicated product environment at Slack scalefrom $300,000
Exclusive buyout of a task set4 to 5x the base price
Typical quarterly contract$300K to $500K

The conflict of interest is not subtle. SemiAnalysis pointed at Datacurve, which runs the DeepSWE 1.1 leaderboard for long-horizon programming. Muse Spark 1.3 max took first place at 75.4. Gemini 3.8 Flash took fourth at 73.8. Datacurve also sells expert coding data and RL environments to frontier labs. The company that grades the exam is selling the answer key, and charging for both.

Math became a hype benchmark ​

The same pattern reached mathematics this month, with higher stakes. On September 8, OpenAI announced that an internal model, running about 10,000 concurrent agents, spent 88 hours and millions of dollars of compute claiming a result on the Navier-Stokes equations. That's one of the seven Millennium Prize problems, carrying a $1M bounty for a peer-reviewed solution. This proof isn't peer reviewed, and the model isn't public.

Two mathematicians, Tristan Buckmaster and Levent Alpöge, had released their own related results the day before. Buckmaster publicly accused OpenAI of racing ahead after months of his work sat in OpenAI's Codex environment. OpenAI denies it, and produced text messages showing it offered the mathematicians a chance to publish first.

The day after the announcement, 25 Fields Medalists, coordinated through Terry Tao's blog (terrytao.wordpress.com), released a declaration titled "A Severe Misalignment of AI in Mathematics." Their core argument: AI companies are using famous unsolved problems as benchmark metrics, and that framing harms mathematics. Solving a problem, they write, is a means to the real goal of understanding. When labs mass-produce true/false statements without the slow work of exposition, simplification, and integration into the canon, the soil that grows new math gets stripped away.

Note what the declaration does not say. It doesn't oppose AI. It says AI can accelerate real mathematical research, and whether the field benefits depends on how the people running the labs decide to act. That's a benchmark problem too: when the metric becomes "solved a famous problem," the behavior that maximizes the metric is exactly the behavior that damages the field. Every incentive problem in this article has the same shape.

What the community is saying ​

Reactions across Reddit, Hacker News, and Chinese tech outlets follow a consistent arc. People who build models feel vindicated. People who run benchmarks feel defensive. People who just want to pick a model feel lost.

I'm in the third group. When I read a leaderboard, a model ranked #2 looks like a safe pick. Then the benchmark updates and the same model is near the bottom. My own testing shows the same pattern: rank correlation between public leaderboard positions and performance on my private holdout sets is weak, and it keeps getting weaker. A model that tops a public eval can lose to a smaller model on tasks sampled from my own production traffic.

Some researchers push back on the manipulation narrative, and their point deserves a hearing. Score drops on a harder test don't prove cheating. If a model trained on the old test's distribution, its score should drop when the distribution changes, even with zero malice. But that's the trap. Whether it's deliberate contamination, benchmark-shaped training data, or natural distribution shift, the observable result is identical: inflated scores on public leaderboards, and a rude awakening on anything private.

The math community's reaction is angrier. The perception, fair or not, is that a lab with effectively unlimited compute reached into an open research community, took direction from work sitting in its own product logs, and claimed a historic result before the humans finished their papers. The declaration speaks in measured language. The subtext is blunt: the labs treat mathematics as a benchmark to be beaten, not a discipline to be served.

Common pitfalls ​

Treating a single public leaderboard rank as a capability statement is the root error. The Gemini 3.8 Flash story is the warning label: #2 of 182 models on one benchmark version, #12 of 14 on the next. A rank describes one test set at one point in time. The distance from that test set to your production workload is unknown until you measure it.

Assuming classic datasets are safe because they're old. IMDB and SQuAD are in every pretraining corpus by now. A small fine-tuned model may still give you meaningful signal, but when you evaluate a frontier model on these, you're often measuring memorization. Build a holdout set from data the model can't have seen: your own logs, your own annotations, tasks you write yourself.

Fine-tuning on benchmark data and calling the gain "learning." GSM8K's 13-point drop after decontamination shows how much apparent reasoning improvement is test-set pattern matching. When your training split and test split come from the same public benchmark, your eval measures overlap, not generalization.

Ignoring who runs the benchmark. Datacurve operates the DeepSWE 1.1 leaderboard and sells expert coding data and RL environments to the labs that top it. A benchmark operator with a financial stake in the results is disqualifying until proven otherwise. Read funding disclosures before trusting a rank.

Chasing saturated benchmarks. When every competitive model scores above 90, the benchmark can't discriminate, and you're looking at a dead measurement. Prefer benchmarks that retire gamed tasks, add adversarial leakage probes, and enforce timeouts, the way Terminal-Bench 4.0 does.

What to do instead ​

If you can't trust public leaderboards, build your own evaluation layer. You don't need a million dollars for this. Sample 200 to 500 examples from your production traffic, have your team label them, and keep that set locked. Reuse it across model versions and candidates.

Use public benchmarks as a coarse filter, then run your private set as the decision. Compare models on at least two unrelated benchmark families first. If they disagree, that's useful information: it usually means one of them has leakage or format-specific overfitting.

If you're building datasets or benchmarks, adopt the practices the newer generation of releases is moving toward. IFM/Code-Reasoning publishes provenance metadata, documents its filtering and synthetic generation, and warns you to check the data before building a pipeline. That's the right default. Add contamination reports. Retire tasks when they saturate. Disclose who funds you.

One thing to remember ​

The Terminal-Bench collapse and the Fields Medalists' declaration are the same event at two different scales. A benchmark stops being a measurement the moment people optimize against it, and the labs have built a supply chain to do exactly that. The only evaluation you can fully trust is the one you control: your data, your tasks, your locked holdout set.

The bottom line ​

If you're picking a model for production, ignore public leaderboard ranks and run a private eval on tasks sampled from your own traffic. The gap between an 89.4 and a 19.1 is a gap in test-set hygiene, not in capability, and your production distribution is the only test set that pays you.

If you're fine-tuning or training, treat every public test set as potentially contaminated, including the classics. For small models that predate the web-scale corpus era, IMDB and GLUE can still teach you things. For frontier models, they're fossils. Build your own holdouts.

If you're building evaluation tooling, design for a short shelf life: retire saturated tasks, add adversarial leakage probes, disclose conflicts of interest, and move sensitive evaluation onto private or dynamically generated test sets. Expect public leaderboards to become marketing artifacts within 6 to 12 months, and the real signal to move to closed evaluations. That's exactly what SemiAnalysis recommends, and exactly why developers lose the shared yardstick that made model comparison possible.