Skip to content

AI Is Proposing the Hypotheses Now: Four Frontiers in Scientific Discovery

#reinforcement-learning #drug-discovery #genomics #scientific-discovery #foundation-models #formal-math

The pattern shift ​

Until a few years ago, machine learning in science mostly played the role of an interpreter. You handed it data, it found patterns, and the hypothesis still came from you.

That division of labor is cracking. Right now, four separate frontiers show models moving upstream, into the stage where the next question gets chosen.

A reinforcement learning agent is searching the SMEFT operator space for the new physics behind experimental anomalies. Generative models are designing drug candidates from scratch. LLMs and a genomics foundation model are mining DNA, living and extinct, for antimicrobial peptides and regulatory effects. And an AI system claims to have produced a proof of a Millennium Problem.

The through-line is a shift in who proposes the next step. The model proposes, the human verifies. Each frontier runs on that same move.

RL hunts for new physics ​

The SMEFT idea fits in one paragraph. You don't know what the new physics is, but whatever it is, its low-energy effects can be written as a long list of higher-dimensional operators built from the Standard Model fields you already have. Find the operator set that matches an anomaly, and you have a model-independent description of what's going on.

The catch is the search space. The simplest counting gives thousands of dimension-six operators, and with flavor indices the space grows to tens of thousands. At loop level, operators correlate in complicated ways; picking one operator in isolation will mislead you. Human analyses usually lean on phenomenological intuition to shrink the space, and that intuition is biased toward the operators people already know.

The RL framing in the paper turns the search into a decision problem. An agent proposes a set of SMEFT operators. The system computes the loop-level prediction, compares it with the experimental observable, and returns a reward. Each round, the policy learns which operator combinations actually explain the data, not which ones a textbook would suggest first.

The test case is the CDF W-mass anomaly, and the numbers show why it's a fair stress test. CDF measured the W boson mass at 80,433.5 MeV/c². The Standard Model prediction sits near 80,357 MeV/c². That's a gap of roughly 76 MeV/c², which the CDF collaboration estimated as about a 7 sigma tension with theory. For context, particle physics treats 5 sigma as the discovery threshold.

The agent reproduced the known result and, per the paper, improved on it. Then it took a harder case: multiple anomalies at once, where operator sets overlap and interact, and it found explanations there too. Nobody should read this as a discovery. The contribution is a search method that removes the intuition bottleneck, and that's the part that matters.

Generating molecules, not just screening them ​

Traditional virtual screening searches a finite library of known compounds. The 82 methods in this review work differently: they generate chemistry that doesn't exist in any database yet. The review organizes them into five families, and each family has a different failure mode.

FamilyCore mechanismWhere it winsWhere it hurts
RNN / TransformerAutoregressive generation over SMILES strings or fragment vocabulariesFast, cheap to train, fluent in SMILESWeak at 3D geometry and stereochemistry
VAEEncode molecules into a latent space, then decode samplesSmooth latent space suits property optimizationSamples often invalid; tends toward familiar chemotypes
GANGenerator and discriminator competeSharp distributions, fast samplingUnstable training, mode collapse; fading from practice
Flow-basedInvertible transforms with exact likelihoodTractable density, good for conditional generationLimited expressiveness, slower sampling
DiffusionIterative denoising on 2D or 3D representationsBest reported results on pocket-conditioned and 3D designSlow sampling, heavy compute

Five families, 82 methods, and the honest summary: no family dominates. The review's comparative analysis suggests that the headline metrics, validity and uniqueness, no longer separate the families as sharply as they once did. The differences that matter show up on target-aware tasks, 3D pocket-conditioned generation, and property control. For a practitioner, the family choice is a task choice, not a quality ladder. If you're designing a ligand for a known pocket, diffusion models are where the best 3D results come from. For fast, cheap exploration of chemical space, an RNN or a VAE gets you there at a fraction of the compute.

The review also collects experimentally validated case studies, which is the part teams actually need. A generated molecule with a perfect validity score and a clean docking pose still has to be synthesizable, stable, and non-toxic, and benchmarks don't tell you any of that. The public repository that ships with the review is the practical takeaway: 82 method references, benchmarks, and evaluation metrics in one place. The review also flags where the field is heading: standardized 3D data, receptor flexibility, interaction-aware generation, and multi-objective design. Those are the directions that will decide whether generated molecules become drugs.

Quick Take: the models now propose the next step of the science, and the human's job is verification, which moves the bottleneck from search to validation.

LLMs as research assistants ​

The antibiotic search is the most practical of the four frontiers, because it's already running as a lab workflow. César de la Fuente's lab uses Codex and ChatGPT to scan genomes, living and extinct, for antimicrobial candidates against drug-resistant infections.

What impressed me about the write-up is how unglamorous the workflow is. The LLM isn't inventing chemistry from scratch. It's writing the code that searches genome databases, extracting candidate peptides, prioritizing them by predicted antimicrobial properties, and handing a shortlist to the lab for wet validation.

The search space matters. These are genomes that have never been screened for antimicrobials, including extinct ones. Nobody has ever tested a Neanderthal peptide, so there's no training signal from real experiments, only predictions. The LLM makes that kind of exploration cheap enough to attempt.

Drug-resistant infections kill roughly a million people a year, and the antibiotic pipeline is thin. The economics of this workflow are the point: candidates that once took months of manual bioinformatics now take a prompt and a validation budget. Whether a candidate survives the lab is a separate question. The search cost just collapsed.

Reading the regulatory code ​

Genomics got its foundation-model moment. AlphaGenome, from Google DeepMind and described in Nature this year, models the regulatory code directly: given a DNA sequence, it predicts gene expression, splicing patterns, chromatin features, and contact maps, most outputs at single base-pair resolution.

The input window is up to 1 million base pairs, and that length is the point. Gene regulation is long-range. An enhancer can sit hundreds of kilobases away from the gene it controls, and earlier models couldn't see both ends of that relationship. At 1M bp, you feed it a gene plus its full regulatory neighborhood in one shot. At 1 bp resolution, you can see whether a single nucleotide change, not a whole gene, perturbs expression.

The practical setup is an API, free for non-commercial use, plus a precomputed resource called the Atlas that covers the entire human genome with variant effect scores and feature importances.

python
from alphagenome.data import genome
from alphagenome.models import dna_client

model = dna_client.create("MyAPIKey")
interval = genome.Interval(chromosome="chr22", start=35677410, end=36725986)
variant = genome.Variant(
    chromosome="chr22", position=36201698,
    reference_bases="A", alternate_bases="C",
)
outputs = model.predict_variant(
    interval=interval, variant=variant,
    requested_outputs=[dna_client.OutputType.RNA_SEQ],
)

That call returns expression tracks for both the reference and alternate bases, overlaid as two signals. You get the variant's regulatory effect rendered as a figure rather than a black-box score.

Key numbers from AlphaGenome:

  • 1,000,000 bp: maximum input sequence, enough for a gene and its full regulatory neighborhood
  • 1 bp: resolution of most outputs, so variant effects localize to a single nucleotide
  • 4 output families: gene expression, splicing, chromatin features, contact maps
  • 0 cost: non-commercial API access, with the whole-genome Atlas precomputed

Two limitations sit in the docs, and they change how you plan a study. The API is sized for thousands of predictions, not millions, so whole-genome variant scoring belongs in the Atlas. And outputs are for research, not clinical decision-making. Both constraints are sensible for a model this new, but they'll surprise you if you read the README as a marketing page.

The math proof question ​

In September, OpenAI said it had produced a proof of the Navier-Stokes existence and smoothness problem, one of the Clay Millennium Problems, and the New York Times covered it. OpenAI's announcement followed the same week. If verified, this is the kind of result that changes a field. The operative word is verified.

What the community is saying: I spent the morning after the announcement in the r/MachineLearning thread, and the tone was far more careful than the headline. The first question was about process, not mathematics. Who has checked the proof? What refereed venue will publish it? Does a machine-generated argument even satisfy the Clay prize rules? Several commenters walked through the requirements: publication in a major refereed journal and acceptance by the mathematical community, a process measured in years. Others pointed at history, where similar announcements for famous problems collapsed under scrutiny. The most useful comment I saw cut through both sides: even a correct proof is a formal-verification project, and that infrastructure is still being built.

That last point is the durable one. AI systems have produced correct proofs before, but never with a million-dollar prize attached. The claim forces the question of how a community verifies an argument that no single human produced. Whatever happens to this proof, the verification pipeline it drags into existence is the real development.

What trips people up ​

Common mistakes in this area, in no particular order.

  • Treating the RL output as a discovery. The agent returns operator sets that fit the data. It does not tell you the anomaly is real, and degenerate operator sets can fit equally well. Read it as a prioritization for human analysis, not a result for the abstract.

  • Sizing the AlphaGenome API wrong. The docs are explicit: thousands of predictions, not millions. If your study needs whole-genome variant scores, use the Atlas. When I tested the API on a few regions it was responsive, but looping a cohort through it would grind to a halt.

  • Trusting molecule generation benchmarks over experimental reality. Validity and uniqueness scores are nearly saturated, but they don't tell you whether a molecule is synthesizable, stable, or active. The review's validated case studies exist precisely because benchmark scores mislead.

  • Reading the Navier-Stokes announcement as a finished proof. The honest status right now is claim under verification. Citing an unrefereed proof as established fact poisons the conversation. Ask what a human referee has signed off on.

  • Ignoring loop-level correlations in SMEFT fits. A tree-level operator choice can be killed by radiative corrections. The RL paper exists partly because humans underestimate these correlations, so if you're doing SMEFT analysis by hand, check operator mixing before you publish.

One thing to remember ​

The common thread is where the models sit in the pipeline. They join at the hypothesis stage now: proposing candidates, while the human verifies. That makes validation the scarce resource. Build your validation pipeline before you build your search, because search is the cheap part.

The bottom line ​

If you're a particle physicist sitting on an anomaly, use the RL search as a systematic second opinion. It reproduces and extends the known SMEFT analysis on the W-mass case, and it will propose operator combinations your phenomenological intuition would skip. Run it before you finalize your own fit, not after.

If you're building a drug discovery pipeline, stop comparing benchmark scores across the five generative families and pick by task. Diffusion models for pocket-conditioned 3D design. RNN or VAE families when you need speed and property control. Budget most of your time for synthesis and assay, because that's where generated molecules fail.

One thing to watch: the Navier-Stokes claim is the first test of how the community verifies a machine-produced proof. Expect formal verification tooling and referee processes to get much more serious over the next year, and treat any "AI solved X" announcement as a status update, not a conclusion.