Appearance
The illusion of the average
A model can top a leaderboard and still hand you a recommendation you shouldn't act on. Five recent works in LLM evaluation and auditing take that gap seriously. The five don't propose better averages. They change what gets measured.
One paper gives you a formal test for when a single recommendation deserves reliance. Another turns repeated-query auditing into a statistical protocol with power calculations. Two more show that standard metrics mislead badly in specific domains, from SVG generation to flight prediction. And one audits the benchmarks themselves, showing that a single score routinely mixes capabilities the label never mentions.
| Work | Target | Unit of analysis | Core method | Headline result |
|---|---|---|---|---|
| Epistemic warrant | Reliance on one recommendation | A single decision | Four-tier reliance certificate | Warrant is distinct from verbalized confidence |
| Dice Roll Method | Brand recommendation stability | Repeated identical queries | Variance decomposition + generalizability theory | Iteration tiers: n=5/10/15 for G=0.58/0.74/0.81 |
| SVG-Score | Text-to-SVG faithfulness | One generated SVG | Human-aligned CLIP + trained VLM judge | CLIPScore misses wrong colors, counts, spatial relations |
| FLY-EVAL++ | Flight trajectory prediction | Per-prediction constraint check | Deterministic verification + rubric aggregation | 28+ point safety gap among models with similar accuracy |
| BenchMIRT | Benchmark contents | Individual prompt | Multidimensional item response theory | Recovers safety and reasoning dimensions without labels |
The pattern across all five: evaluation needs to happen at the unit where the decision actually occurs. For a procurement officer relying on a model's vendor pick, that's one recommendation. For an auditor checking brand answers, that's a batch of repeated queries. For a benchmark being used to claim safety, that's the individual prompt. Aggregate accuracy papers over all of it.
Epistemic warrant: when ground truth is unavailable
Most organizational decisions don't have a ground truth to check against. You can't verify whether the model's recommended supplier was the right one, so the basis for reliance has to come from somewhere else. The epistemic warrant paper adapts an idea from epistemology, where warrant means the justification for holding a belief, and applies it to models: a recommendation deserves reliance when the model's preference is stable and the scope over which that preference holds is understood.
The contribution is operational. The authors define a four-tier reliance certificate for pairwise recommendations, ranging from unstable (the preference flips under perturbation) through context-dependent and locally supported up to broadly supported. Each tier tells a decision-maker what they're allowed to conclude.
The validation matters as much as the construct. Known-groups tests recovered expert-prespecified warrant orderings, and stronger warrants systematically aligned with independent consensus from crowd workers. The key result: warrant carries information that verbalized confidence doesn't. A model can say it's 99% sure and still flip its preference when you rephrase the question. That's the gap this paper is after.
Practical implication: before acting on a high-stakes recommendation, perturb it. Change the phrasing, swap the ordering, vary the context. If the preference survives, you have evidence for a higher warrant tier. If it flips, you've just dodged a confident wrong answer.
Repeated queries are a measurement problem
If you've ever run the same prompt twice and gotten different answers, you know the brand-recommendation audit problem firsthand. The Dice Roll Method formalizes what's going on. Under a generative model of temperature-scaled nucleus sampling, total response variance decomposes into sampling, prompt-phrasing, run-to-run, and model-version components. The statistical stack on top: a negative-binomial mixed model with iterations as repeated measures, Cliff's delta for distribution-free effect sizes, dependence-preserving bootstrap, simulation-based power, and a generalizability-theory decomposition. Drift diagnostics on pinned snapshots catch model-version changes over time.
The tiers come from reanalyzing five brand-recommendation audits: roughly 190,000 observations, 270+ brands, six languages, iteration counts ranging from 5 to 40. The protocol also distinguishes four metric families: count, set, embedding, and fairness-adjusted PASOR. The authors' point is that these measure different things, and a compact battery beats any single indicator.
The external validation is the strong part. Pre-registered re-runs on three independent corpora, including Motoki et al.'s 100-round study and Rozado's 24-model audit, reproduced the D-study reliability prediction in 37 of 39 cells with no failures. The n=5 power value reproduced to two decimals. Then comes the twist: the fixed tiers don't transfer. The same validation that confirms the statistical machinery shows you can't copy the iteration counts to a new model version. The authors recommend a pilot-then-solve reading: estimate reliability in a small run, then choose your iteration count.
| Tier | Iterations | G-coefficient | What it supports |
|---|---|---|---|
| Exploratory | 5 | 0.58 | Detecting that variation exists |
| Confirmatory | 10 | 0.74 | Stable effect directions, coarse rankings |
| Rigorous | 15 | 0.81 | Fine-grained ranking claims |
For an audit with money or reputation on the line, treat 15 as a floor, not a target. G=0.81 still leaves a fifth of the variance unaccounted for, and your pilot estimate is the only reliable guide to how many queries your specific setup needs.
Quick take: aggregate scores hide exactly the information you need when a single model output feeds a decision, and what you perturb, where you look, and how many times you ask matter as much as the model itself.
Benchmarks measure more than their labels
BenchMIRT, from Ai2, attacks the benchmarks themselves. Building on item response theory, a psychometrics technique for estimating what each test question measures, it applies multidimensional IRT at both the model level and the question level. For each question it estimates difficulty and discrimination. For each model it estimates strength on whichever latent capabilities the questions draw on. Trained on 100 open-weight LLMs across 16 benchmarks and more than 34,000 questions, enough data for stable IRT estimation, it wasn't told which benchmark measured what. It recovered two dominant dimensions anyway: safety and general reasoning. Re-run from scratch, the same two dimensions emerged.
The per-benchmark findings get uncomfortable.
| Benchmark | Intended construct | What BenchMIRT found |
|---|---|---|
| MMLU-Pro, GPQA, MATH, BBH | General reasoning (one of six reasoning suites) | Tracked reasoning, as labeled |
| HarmBench (standard, contextual) | Safety | Aligned with safety, as labeled |
| HarmBench (copyright) | Safety | Aligned with general reasoning |
| BBQ | Social bias / safety | Aligned with general reasoning |
| WMDP | Dangerous knowledge | Aligned with reasoning, but negatively: higher reasoning, lower score, because refusing is scored as correct |
| WildJailbreak (harmful prompts) | Safety | Aligned with safety |
| WildJailbreak (benign prompts) | Refusal calibration | Aligned with general reasoning |
BBQ, grouped with safety benchmarks everywhere, aligned more strongly with general reasoning. A low BBQ score can reflect comprehension difficulty rather than stereotype reliance. WMDP behaves even stranger: scores track reasoning negatively, because the benchmark rewards refusing to answer, and a model that reasons better knows it should refuse. HarmBench splits internally, with copyright questions behaving more like reasoning tasks than safety tasks. None of this shows the benchmarks are broken. It shows a single averaged score can combine several distinct signals.
BenchMIRT also has a practical payoff. Keeping only 10% of the best questions preserved nearly the same picture of which models are stronger or weaker on safety and reasoning; 50% often matched the full set even more closely. And it predicts held-out question performance correctly 79% of the time, versus 70% for a benchmark-average baseline. That's enough to shrink evaluation suites substantially.
The caveats are honest. The training models all predate March 2025, so newer generations could shift the recovered dimensions. The dimensions themselves depend on the benchmark set you feed in; a different mix could surface different capabilities. And for the narrow task of ranking models on randomly held-out items, the plain average still does slightly better. BenchMIRT's advantage is resolution, not ranking accuracy.
The transparency also cuts both ways. The same question-level estimates that identify a benchmark's most informative safety items could be used to delete them, producing a safety eval an unsafe model passes. The authors accept the trade-off. The cost only shows up later, when someone games an eval with the map BenchMIRT provides.
Domain metrics quit working outside their domain
Two papers show what happens when you evaluate with a metric built for a different world.
SVG-Score applies to text-to-SVG generation, where the default evaluation is CLIPScore. CLIP was trained on natural images and never saw vector graphics. In controlled perturbation experiments, CLIP-based scores barely reacted to the mistakes SVG generators actually make: wrong colors, wrong counts, broken spatial relations. The errors that should tank a score don't move it. Off-the-shelf VLM judges did better but responded unevenly across error types and SVG styles. The fix is a human-annotated Semantic Alignment dataset plus two evaluators built on it: an adapted CLIP scorer aligned to human preferences for fast large-scale runs, and a VLM judge trained with supervised fine-tuning and reward-shaped reinforcement learning for more interpretable assessment. CLIP isn't weak here. It's being applied outside its trained domain, and the numbers show what that costs.
FLY-EVAL++ makes the same point where the stakes are higher. Flight trajectory prediction can be numerically close to ground truth and still violate operational constraints, combine fields in physically inconsistent ways, or fail to produce a usable structured output. The instantiation extends PilotBench with history-conditioned and multi-step prediction tasks. Across 66 LLMs, safety compliance was the most discriminative dimension of model behavior. Models with comparable predictive performance differed by more than 28 points in safety score. The protocol combines deterministic verification of protocol compliance, physical feasibility, and safety constraints with rubric-based aggregation into multi-dimensional scores. The recurrent failures it surfaced, safety violations under physically plausible predictions and instability in multi-step rollouts, would be invisible in an accuracy-only report.
Key numbers
- G = 0.81: the generalizability coefficient at 15 iterations, the rigorous tier for repeated-query audits. Below that, rankings carry noise.
- 28 points: the safety-score gap between FLY-EVAL++ models with comparable predictive accuracy. Accuracy reporting alone misses it.
- 79% vs 70%: BenchMIRT's held-out question prediction accuracy versus the benchmark-average baseline.
- 2: the dominant capability dimensions BenchMIRT recovered from 16 benchmarks, without being told what any of them measured.
- 190,000: observations reanalyzed in the Dice Roll corpus, covering 270+ brands and six languages.
What the community is saying
I've watched a model flip between two brands across ten identical queries and assumed the model was being perverse. The Dice Roll corpus suggests otherwise: the variation is structural, and most of it isn't irreducible randomness. It decomposes into sampling, phrasing, run-to-run, and model-version components. My real mistake was underpowered auditing. At n=5, a G-coefficient of 0.58 means a chunk of what looked like a ranking was sampling noise. The fix was more queries, chosen with a power analysis.
BenchMIRT's results landed in a familiar place too. I've debugged "unsafe" model behavior only to find the prompt was a comprehension test in disguise. BBQ aligning with reasoning rather than safety matches that experience: low scores get blamed on the model when the benchmark measures a different capability. WMDP's negative reasoning correlation explains why stronger models sometimes score worse. The benchmark rewards refusal, and a model that reasons harder knows it should refuse.
The sharpest debate is the transparency trade-off. The same estimates that reveal a benchmark's most informative safety questions can be used to strip them out. I've seen teams trim failing items to push a model through an evaluation gate; a principled map of which questions carry the signal makes that easier, not harder. The authors are right that transparency is worth it, but it's a narrow bet, and the attack surface is the part people don't want to talk about.
Common pitfalls
Averaging a benchmark that measures two things. HarmBench's copyright questions track reasoning, not safety. Averaging them with jailbreak prompts produces a number that means neither. Read the item-level signal before trusting the headline.
Treating CLIPScore as a general visual metric. It was never trained on vector graphics. In the SVG-Score experiments it barely reacted to wrong colors, counts, and spatial relations. Use a domain-adapted scorer or a VLM judge trained on human preference data for the domain.
Confusing verbalized confidence with reliability. The epistemic warrant results show confidence and warrant are distinct. A confident model can still flip its preference under perturbation. When the decision matters, perturb before you act.
Underpowering repeated-query audits. n=5 gives a G-coefficient of 0.58, so rankings derived from it are partly noise. The fixed tiers don't transfer across model versions either. Pilot first, estimate reliability from your own data, then pick your iteration count.
Reporting accuracy alone in safety-critical domains. In FLY-EVAL++, predictive accuracy barely distinguished models that differed by more than 28 points in safety compliance. Report constraint satisfaction and structured validity explicitly.
One thing to remember
Evaluation is a measurement task with its own statistical and epistemological rules. The five works agree on the diagnosis: a better average won't fix a misaligned unit of analysis. Decide whether you're measuring a single recommendation, a repeated query, a constraint violation, or a benchmark item, then pick the tool designed for that unit.
The bottom line
If you're building an evaluation pipeline for organizational decisions, adopt a warrant-style check: perturb each high-stakes recommendation, verify the preference is stable, and report the scope of that stability. Confidence scores alone are not a basis for reliance.
If you're auditing model outputs for publication or procurement, standardize with the Dice Roll stack, but run your own pilot first. Estimate reliability in a small run, then target n=10 for confirmatory claims or n=15 for rigorous ones, and report the variance decomposition next to your point estimates.
One thing to watch: prompt-level benchmark auditing is about to change how safety evals get built and attacked. The 10% question subset already preserves rankings, so expect leaner benchmarks within the year, and expect attempts to game them by trimming the very items BenchMIRT flags as most informative. Whether the transparency survives that attack surface is the open question.