Skip to content

Aggregate Scores Are Hiding Regressions: What Five New Eval Papers Found

#llm-evaluation #benchmarking #agents #multimodal #regression-testing

Here's a scenario you've probably lived. A vendor deprecates the model you're on. You migrate to the successor. The aggregate benchmark score goes up a few points, so you ship it. A week later, a customer files a bug about something that used to work. The benchmark said things improved. What broke?

Five papers released this month suggest the benchmark wasn't wrong, exactly. It was compressing. The studies cover item-level regression analysis, bilingual document reasoning, market-validated agent tasks, credit-assignment validity, and latent-factor structure. Different methods, different domains, same conclusion: standard evaluation practice systematically hides failures.

StudyWhat it evaluatesScaleKey finding
BEAR-Bench [1]Bilingual (EN/RU) professional document reasoning1,000 questions, 16 MLLMsStrongest systems still leave clear headroom
API migration audit [3]Item-level change across GPT-5.4 to GPT-5.6900 items × 50 runs × 3 benchmarksImprovements and regressions coexist in all 9 cells
StartupBench [4]Market-validated E2E agent workflowsReal startup product workflowsBest agent completes about 30%
Cross-view correspondence [2]Validity of agent credit assignment1,586 trajectory pairs55.9% temporal localization disagreement
Latent factor analysis [5]Human vs. LLM cognitive structure6 LLMs, 2 assessmentsExperts can't interpret LLM-derived factors

What the migration audit found ​

The item-level regression paper [3] is the most immediately useful thing in this cluster. The authors took three real API migrations in the GPT-5.4 to GPT-5.6 Sol product sequence, ran 900 public benchmark items 50 times per item per model, and classified each item as reliably improved, reliably regressed, practically equivalent, or inconclusive. Classification ran under false-discovery-rate control with a practical-significance threshold, calibrated against a label-permutation null. That's 45,000 responses per model per benchmark. Enough data that sampling noise isn't an excuse.

The headline: in all nine migration-benchmark cells, reliable improvements and reliable regressions coexist. Edges with aggregate gains of up to 7.3 percentage points still contain up to 8.3% reliably regressed items. Edges with aggregate losses contain up to 10.7% reliably improved items. Aggregate scores hide movement in both directions at once. A 7.3-point gain with 8.3% of items regressing means roughly 1 in 12 questions got worse on an upgrade that looked like a win.

Key Numbers: 9 migration-benchmark cells, all with coexisting improvements and regressions. Up to 8.3% of items regress on edges that look like gains. Up to 10.7% improve on edges that look like losses. 50 runs per item is what separates real change from noise.

The strict versus loose scoring finding is the one that should scare you. On the instruction-following benchmark, the latest migration shows a 3.9-point regression under strict scoring and a 0.04-point change under loose scoring. Same model pair, same items, same responses. The only difference is how strictly the output is graded. If your eval harness uses loose scoring, regressions that exist will not show up in your numbers.

I've made the loose-scoring mistake myself. When I first built an eval harness for a product feature, I graded for "did the model mention the right entity" instead of "did the model do what the user asked." The scores looked great. The product felt broken. When I re-ran the same items under strict scoring, the regression appeared immediately. This paper is the formal version of that experience: scoring strictness isn't a detail, it's the measurement.

BEAR-Bench: bilingual reasoning with real headroom ​

BEAR-Bench [1] targets a gap most multimodal benchmarks ignore. Existing evals emphasize information extraction, require external domain knowledge, or treat professional documents as one setting among many. They're also mostly English-centric or Chinese-centric. Russian, in particular, is substantially underrepresented. BEAR-Bench is a self-contained benchmark of 1,000 human-annotated questions built on text-rich business and scientific documents in English and Russian. That's small enough to run a full evaluation in an afternoon, but large enough that the 16-model comparison holds up.

The authors evaluated 16 proprietary and open-weight MLLMs, including Gemini 3.1 Pro and Qwen3.5-397B. The 397B-parameter model needs multi-GPU serving, not a single card, and it still leaves headroom on the benchmark. That's the finding: even the strongest systems have clear room to improve on professional document reasoning.

The paper also does something smart with the outputs. It compares hallucination detection methods, measuring not just how often models fail on BEAR-Bench but how reliably those failures can be caught. That's the question that actually matters in production. A model that hallucinates is a problem. A model that hallucinates and evades detection is a liability.

Quick Take: if your eval only reports aggregate scores, you're not measuring regressions, you're averaging them away.

StartupBench: the 30% ceiling ​

StartupBench [4] takes a completely different approach to task selection. Instead of researcher-selected tasks, the authors studied AI startup products with demonstrated adoption, extracted their product workflows and users, and translated those workflows into deliverable-oriented tasks with fine-grained rubrics. The tasks are market-validated: real products people pay for, doing work people actually need done.

The result is humbling. Under a unified agent harness, the strongest model completes about 30% of StartupBench. It makes substantial partial progress on many tasks, but end-to-end completion is where it falls apart. The failure analysis points at complex instruction following and domain-specific expertise as the main culprits.

30% is the number to sit with. It means 7 out of 10 real-world workflows fail end-to-end, even for the best agent available. Partial progress is not progress that ships.

My team ran into this exact gap when we tried to use our internal agent benchmark to predict production performance. Our researcher-selected tasks were too easy, and we had no idea until a pilot with real users fell apart. StartupBench is the control group we were missing.

Cross-view correspondence: the preprocessing problem ​

The cross-view correspondence paper [2] is the most abstract work in this cluster and possibly the most dangerous finding. Agent evaluations and trace-based learning often compare outputs across transformed views: different code versions, different SQL dialects, different prompt templates. A post-response correspondence map is used to align them, and it's treated as neutral preprocessing.

It isn't. The paper shows this correspondence is a measurement intervention. Omit it and you can manufacture sensitivity. Make it too aggressive and you can manufacture invariance. When multiple optimal correspondences exist, mechanism labels and signed learning credit become unidentified. You can't tell which step in the trace deserves credit for the outcome, because the answer depends on which map you chose.

The empirical numbers back up the theory. Two deterministic optimal tracebacks disagree on temporal localization for 55.9% of 1,586 nonzero trajectory pairs. For more than half of the pairs, which step gets credit flips depending on which optimal map you pick. Two frozen 800-rollout tool-use audits, including a task-and-seed-disjoint replication, exposed exact-optimum reversals of intended turn-level credit.

AuditScaleFinding
Code/SQL tracebacks1,586 nonzero trajectory pairs55.9% disagree on temporal localization
Tool-use audits (frozen)800 rollouts × 2, task-and-seed-disjointExact-optimum reversals of turn-level credit
Transport gate (pre-registered)Natural responsesFailed; benign-only calibration erased harmful responses

The transport gate result has the sharpest practical edge. A map calibrated only on benign examples erased every retained harmful response. Two-sided validation selected response-preserving alternatives instead. When I've built agent eval harnesses, I treated the correspondence between traces as a plumbing detail. This paper is the argument for why that's wrong. The map between views is a measurement intervention, and it needs to be declared, validated, and propagated into uncertainty before any point conclusion is supported.

Interpretable humans, alien LLMs ​

The latent-factor paper [5] is the one that keeps me up at night. The authors ran Exploratory Factor Analysis on responses from humans and six LLMs across quantitative reasoning and chemistry assessments. Subject-matter experts then blindly evaluated the resulting factor graphs, trying to ascribe pedagogical meaning to the latent constructs.

The experts successfully interpreted most human-derived factors. They could not ascribe meaning to any LLM-derived factors in quantitative reasoning, and they could interpret only half of the LLM factors in chemistry. The LLMs perform at or above human level on these assessments while the latent structure of their responses is statistically opaque to human experts.

This is the strongest evidence I've seen for the claim that LLMs don't reason like humans. If you're using LLM outputs to model student cognition, or building evals that assume a shared cognitive construct between AI and humans, this paper is a warning. The assumption that AI and humans employ similar underlying constructs is not holding up under direct test.

Common Pitfalls ​

Five mistakes show up across these papers, and I've hit most of them myself.

Reporting only the aggregate. A single number compresses bidirectional item-level change into nothing. Report the distribution, or at least the reliable regression rate alongside the mean. The migration audit shows why: 7.3 points up can hide 8.3% of items going reliably backward.

Treating correspondence maps as neutral preprocessing. In agent evals, the map between transformed views is a measurement intervention. Validate it with nuisance-removal and response-preservation checks, or your credit assignment is unidentified.

Scoring loosely and calling it done. The 3.9 versus 0.04 gap shows scoring strictness can erase real regressions. Pick strict scoring, audit your rubric, and check what changes when you loosen it.

Benchmarking on researcher-selected tasks. StartupBench's 30% ceiling suggests researcher intuition overestimates agent capability. Ground at least some of your eval tasks in real workflows with demonstrated demand.

Assuming human-interpretable constructs. If your eval assumes LLMs reason with the same latent factors as humans, the EFA results say you're wrong. Design evals that don't depend on that assumption, and treat statistically opaque mechanisms as a feature of the system, not a bug in your analysis.

One thing to remember: every number in this article came from a paper that released its data or its method. The migration audit released the complete response-level archive and per-item scoring outputs. BEAR-Bench released its 1,000 questions. If you're building an eval pipeline, you can run these checks yourself instead of trusting the aggregate.

The Bottom Line ​

If you're migrating between model API versions, don't ship on aggregate scores alone. Run item-level regression analysis with repeated sampling and a practical-significance threshold, because a 7.3-point gain can hide 8.3% of items getting reliably worse.

If you're building agent evals or trace-based learning, declare and validate your cross-view correspondence before drawing conclusions. Two optimal maps disagreed on temporal localization for 55.9% of trajectory pairs, so credit assignment flips depending on which one you pick.

If you're evaluating on English-only, extraction-heavy, researcher-selected tasks, you're overestimating your system. Add bilingual professional-document reasoning and market-validated end-to-end workflows to your suite, because the strongest models still fail 70% of real-world tasks.

Sources ​

[1] BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models. http://arxiv.org/abs/2608.17895v1

[2] Cross-View Correspondence Is a Measurement Intervention: Two-Sided Validation for Agent Evaluation and Credit Assignment. http://arxiv.org/abs/2608.17713v1

[3] What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations. http://arxiv.org/abs/2608.17719v1

[4] StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows. http://arxiv.org/abs/2608.17800v1

[5] Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assessment Responses. http://arxiv.org/abs/2608.17810v1