Skip to content

Clinical AI Scores Are Built on Choices Nobody Reports

#clinical-ai #model-evaluation #bias-auditing #medical-imaging #llm-reasoning #protein-design

The score hides the pipeline ​

Every medical AI paper ends with a number. A voxel AUROC. An any-shift rate. A Jaccard agreement score. Every one of those numbers was produced by someone making choices: how the reference was aligned, where the threshold was set, which false-positive budget was allowed, which explanation method was trusted, which lesion definition counted as a hit. The number hides the choices.

Five recent preprints, plus one dataset page that fails with a schema cast error, converge on the same lesson from different directions. The evaluation apparatus determines what you conclude about the model. Often it determines more than the model does.

The axis bug that sank a diffusion model ​

MIRTO is an evaluation protocol for unsupervised anomaly detection (UAD) in brain MRI, and it is the sharpest example of the problem. UAD methods rank on a single score, and that score rests on choices that are rarely reported: how each anomaly map is aligned with the reference, how and on which data the threshold is set, and which false-positive budget, metric, aggregation, and lesion definition are used. MIRTO makes every one of those choices explicit and measures their effect.

The first casualty was an axis-order mismatch. One diffusion model's stored anomaly maps had their axes ordered differently from the reference. Voxel AUROC collapsed from 0.873 to 0.583. That is the difference between a deployable model and a coin flip, since 0.5 is chance. Slice-level AUROC barely moved, so any evaluation that reported only slice-level scores would have shipped that model without noticing.

Fifteen thousand pipelines, one honest answer ​

Applied to four UAD methods trained on the same healthy data and tested on 312 BraTS 2020 subjects, the analysis produced a striking pattern. MIRTO repeats every comparison over 15,552 defensible evaluation pipelines. That number is a confession: there is no canonical pipeline. Your one-pipeline leaderboard is a single sample from a space of thousands.

Within each metric, the choice of evaluation method explained at least 0.95 of the variance in voxel AUROC and AUPRC, and 0.77 in Dice. For lesion sensitivity it explained only 0.14. The lesion definition and hit criterion dominate that metric, so choosing your lesion definition is choosing your result.

A Dice advantage that was statistically significant at validation thresholds vanished when methods were compared at equal realized false-positive burden. That is threshold transfer: the ranking reflected where you cut, not what the method could do. A training-free change to REFLECT's latent aggregation raised Dice at equal burden by 0.052. No retraining, no architecture change, a bigger gain than many published method improvements.

MIRTO states its own limits up front. All nine hypotheses were tested against explicit criteria, and because the same cohort served to develop the protocol, all inference is exploratory. That level of self-disclosure should be the default.

Size doesn't buy fairness ​

A counterfactual audit of ten open-source LLMs asks a direct question about pediatric Emergency Severity Index (ESI) prediction: change only the patient's demographics, socioeconomic status, healthcare access, behavior, social context, or system context, and does the acuity assignment shift? The clinical presentation stays fixed. The model families span Qwen2.5-7B, Qwen2.5-14B-Instruct, a QLoRA fine-tuned Qwen2.5-7B, MedGemma variants, MedLLaMA2-7B, GPT-OSS-20B, and GPT-OSS-120B. These are the models people actually consider for local, privacy-preserving clinical decision support, because 7B and 14B weights run without a cloud GPU.

Key Numbers10 open-source LLMs audited on pediatric ESI triage 5.27% any-shift rate for the fine-tuned Qwen2.5-7B, versus 16.02% for the base model 0.0534 vs 0.1706 mean absolute ESI shift, same pair Larger and medical-domain models did not shift less. Several shifted more.

Counterfactual sensitivity varied substantially across the ten models. It did not consistently decrease with larger size or medical-domain pretraining. The QLoRA fine-tune cut the any-shift rate more than threefold, from 16.02% to 5.27%, and mean absolute shift from 0.1706 to 0.0534. Several larger or medical-domain models shifted more than the base Qwen. Domain fine-tuning on triage-adjacent data is the cheapest fairness mitigation on the table.

Aggregate rates hid the clinically interesting part. Stratified and correlation analyses revealed directionality: undertriage and overtriage behave differently, and the axes cluster into shared failure patterns that a single any-shift number flattens away.

Quick Take: Evaluation choices dictate the conclusions, from voxel AUROC to ESI any-shift rates.

When explanation methods disagree ​

Explainable AI has a disagreement problem. SHAP says the model looked at "sepsis," another method says it looked at "history," and both can be describing different aspects of the same forward pass. For medical text, that ambiguity carries clinical weight.

The DeBERTa-v3 framework audits zero-shot classification of medical abstracts using a natural language inference engine, five enriched hypotheses per diagnostic category, and a balanced sample of 1,000 texts per class. Five explanation methods are compared: SHAP and LIME as model-agnostic approaches, occlusion and Input x Gradient as deep-learning-specific approaches, and Attention x Gradient as a transformer-specific approach. Explanations are standardized through top-token attribution, and pairwise agreement is quantified with the Jaccard index.

Two clear patterns emerge. Predictive accuracy is high in well-defined clinical domains and degrades under high semantic ambiguity. Explanatory stability mirrors predictive certainty: methods converge in univalent categories and diverge under diagnostic uncertainty. Qualitative error auditing adds three systemic failure mechanisms: lexical hypersensitivity, semantic overlap, and loss of attribution coherence.

The defensible workflow treats explanation as a collection of methods, with quantitative agreement metrics, and prefers specific clinical ontologies over broad diagnostic labels as classification targets.

The clinical reasoning rubric gap ​

Exam-style accuracy does not establish that an LLM reasons over clinical records. The rubric review defines clinical reasoning as integrating and updating evidence across time and sources to form, revise, and justify a patient's problem representation and a defensible plan.

It maps three literatures: medical education assessment instruments, clinical LLM benchmarks published from 2023 onward, and general-domain methods for evaluating long-form generation. Six dimensions get examined:

DimensionCoverageWhat exists
Problem representationReasonableMedical education instruments; reliability varies by setting
Differential and management reasoningReasonableBenchmarks; reliability varies by setting
Temporal synthesisTargetedTIMER-Eval
Sequential diagnostic belief updatingTargetedER-Reason
Counterfactual reasoningEmergingDedicated evaluations; limited on longitudinal free-text
Calibrated uncertaintyEmergingEarly evaluations; not applicable to long-form reasoning
Reasoning faithfulnessWeakestOne clinical causal-ablation study on MCQs

No single instrument covers all six dimensions. Problem representation and differential or management reasoning are reasonably covered. TIMER-Eval targets temporal synthesis, and ER-Reason assesses sequential diagnostic belief updating. Counterfactual reasoning and calibrated uncertainty evaluations are emerging, but they don't handle longitudinal free-text yet. Factual completeness is well theorized in general-domain evaluation, with early clinical evidence of important omissions. Faithfulness is the weakest dimension: one identified clinical causal-ablation study on multiple-choice questions.

The review's combination strategy is concrete: binary rubric items, separate completeness and correctness scores, case-specific importance weighting with non-compensable safety caps, temporal order-consistency checks, and chance-corrected reliability reporting. The review provides a design rationale, with a validated instrument still to come. Read it as the spec sheet for what to build next.

IDiom: interpretable control for disordered proteins ​

One paper in this cluster builds instead of evaluates. IDiom is an autoregressive protein language model for intrinsically disordered regions (IDRs), the protein segments that don't fold into stable structures yet run transcriptional regulation, signal transduction, and subcellular localization. Structure-based design methods don't apply to IDRs, and existing protein language models train on full-length sequences, so their prior is biased toward folded domains.

IDiom trains on IDiom-DB, a dataset of 54 million predicted IDRs curated from the AlphaFold Database. That scale lets the model recapitulate the composition, patterning, motifs, and predicted disorder of natural IDRs without inheriting the folded-domain bias of full-sequence models.

RL-SAE, reinforcement learning with sparse autoencoder features, is the post-training method that matters here. It rewards the model for generating sequences that activate specified SAE feature sets. Across eight IDR design tasks, RL-SAE sequences activated 90% of 30 targeted features on average, versus 24% for activation steering. Roughly four times the feature control, and the control is interpretable by construction.

RL-SAE sequences also improved predicted subcellular localization and transcriptional activity compared with steering and supervised fine-tuning. Different biological functions can be combined within a single sequence by composing their feature sets. That is the evaluation lesson applied to generation: specify what you want in feature space, optimize for it directly. Code is on GitHub under rotskoff-group/idiom.

What the community is saying. The practical side of this cluster is messier. The opus5 dataset of doctor-patient conversations aims to cover all human diseases, with clinician personas, patient scenarios, differential diagnoses, and common mistakes. When I tried to load it, the viewer died with a schema cast error: a nested struct field carried a dmid column that the declared feature schema didn't expect. One field, one type mismatch, and the dataset became unviewable. It is the same story MIRTO tells at a different scale: the infrastructure choice you didn't document determines whether anyone can use your work.

Common pitfalls ​

Five mistakes show up across this cluster repeatedly.

  1. Publish the pipeline, not just the score. Method choice explains 0.95 of the variance in voxel AUROC. Reporting AUROC without registration, threshold, false-positive budget, and lesion definition makes the number unreproducible. The axis-order mismatch moved a diffusion model from 0.873 to 0.583 while slice-level scores looked fine.

  2. Set thresholds on validation data only. The Dice advantage that vanished at equal false-positive burden came from threshold transfer. If your threshold touched test data, your ranking is an artifact of your choice.

  3. Assume fairness scales with size or domain pretraining. The audit found larger and medical-domain models with more counterfactual shift than a fine-tuned 7B. Run your own paired counterfactual vignettes before deployment.

  4. Trust a single explanation method. Attribution methods disagree most when predictions are genuinely ambiguous. Use at least two methods, a quantitative agreement metric, and report the disagreement.

  5. Report bias as one aggregated number. Any-shift rates hide directionality. Disaggregate undertriage from overtriage and stratify by demographic axis before claiming a model is fair.

One thing to remember ​

Every result in this cluster is a statement about evaluation choices. The axis-order bug, the threshold transfer, the explanation disagreements, the missing rubric dimensions, the SAE features that turn generation into controllable optimization. When you read the next clinical AI score, ask which choices produced it before you ask what the score means.

The Bottom Line ​

If you're deploying an open-source LLM for clinical triage, run a counterfactual audit with stratified directionality analysis before you commit. A QLoRA fine-tune on domain data cut any-shift from 16.02% to 5.27% in the audit, which is the cheapest bias mitigation currently on the table.

If you're ranking unsupervised anomaly detection methods, adopt a MIRTO-style protocol with registration gating, validation-only thresholds, and equal false-positive burden comparisons. The 15,552-pipeline approach is the only way to know your leaderboard reflects method quality rather than evaluation choices.

One thing to watch: reasoning faithfulness evaluation for clinical free-text is the weakest dimension with the most regulatory weight. Expect dedicated faithfulness benchmarks within a year, likely built on causal ablations, because faithfulness is the gap every other evaluation dimension runs into eventually.