Skip to content

The Proof Was the Easy Part: AI, Navier-Stokes, and the Breaking of Scientific Credit

#ai-for-science #navier-stokes #scientific-credit #academic-publishing #foundation-models #peer-review

The morning the problem died ​

Somewhere this week, a mathematician woke up, made coffee, and found out the problem they'd given their life to had been finished over a long weekend. Not by a rival. By a machine that has never lost sleep, never felt the particular ache of being close.

I keep going back to that morning. Not the announcement. The morning. The specific, quiet vertigo of having spent fifteen or twenty years inside a single question and watching someone turn on the lights.

Here's the doorway. OpenAI's unreleased model, running something like ten thousand agents in parallel, produced a 165-page proof of the Navier-Stokes existence and smoothness problem in about 88 hours. Navier-Stokes is one of the seven Millennium Prize problems, each carrying a million-dollar bounty from the Clay Mathematics Institute. It had stood for roughly ninety years. The run burned 300 billion output tokens, the rough text volume of a large university library, almost all of it dead-end search and discarded reasoning. The final artifact is 165 pages. The compute tab came to around $22.5 million, about what a well-funded academic lab spends in a year, concentrated into a single week.

The proof was checked by a system that mechanically verifies each logical step. That's why people are taking it seriously. It's also, almost certainly, why the dispute that followed got so nasty so fast.

The race was the story all along ​

Proofs get coverage. The skirmish around this one is the part that matters.

An NYU mathematician, Tristan Buckmaster, announced this week that he and Levent Alpöge, a mathematician at Anthropic, had made progress on Navier-Stokes. They hadn't published their full results. Buckmaster says information about that progress reached OpenAI shortly before the company's own effort started. OpenAI then released its complete proof first, credited to the model. Nobody disputes the timeline.

Then it curdled. Buckmaster pushed to keep Alpöge credited as a collaborator. OpenAI's Sébastien Bubeck allegedly asked him to drop the credit as part of a compromise, then told him "why would you ruin your career" when he pushed back. OpenAI says its team never saw any of the human researchers' work before the release, though it admits it can't fully rule out that anonymized data from its own products played a role. It argues the two proofs differ in specifics.

The blowback arrived within days. Twenty-five Fields Medalists signed an open letter arguing that rushing to win a race to a proof, without the writeup and attribution work that normally comes with it, breaks the way mathematical knowledge gets passed on and trusted. Caltech researchers pushed back hard enough that OpenAI pulled its sponsorship from a math event there.

Step back and the structural problem is hard to miss. The moment a lab has enough compute to finish any problem it can sense a human is close to, publishing progress becomes a liability. The public record starts signaling targets. That's not an ethics footnote. It changes the incentive structure underneath all of mathematics.

Quick Take: The 88-hour proof is the fast part of this story. Scientific credit, peer review, and publishing norms were built for human-paced discovery, and they have not survived contact with it.

447 papers in one day ​

The same pressure showed up at a different scale, in the publishing system this time. On September 9, 2026, arXiv's cs.LG category hit an all-time daily high of 447 new machine learning papers. The baseline around the spike is roughly 200 a day.

That is many times more than a person, or even a sizeable reading group, can digest in a year. The number isn't a curiosity. It's the structural answer to what happens when the production cost of a research artifact collapses. Zachery Lipton, who has spent years arguing that CS academia broke its own incentive system, put it bluntly: perhaps all it takes for the system to rebuild is for it to burn to the ground.

json
{
  "type": "line",
  "title": "cs.LG daily submissions during the week of the Navier-Stokes proof",
  "x_label": "Date (September 2026)",
  "y_label": "Papers per day",
  "caption": "cs.LG hit an all-time high of 447 submissions on September 9, 2026. Other days are shown at the reported ~200/day baseline. The spike and the proof are symptoms of the same collapse in production cost.",
  "data": {
    "labels": ["Sep 7", "Sep 8", "Sep 9", "Sep 10", "Sep 11"],
    "series": [
      {
        "name": "cs.LG papers per day",
        "values": [200, 200, 447, 200, 200]
      }
    ]
  }
}

Nobody reads 447 papers a day. Nobody reads 200. What actually happens is that evaluation shifts downstream: papers get filtered by reputation, by search ranking, by whose name is attached, by what gets amplified. Peer review, already stretched, becomes a formality that happens months after the field has moved. The 88-hour proof and the 447-paper day look like different problems. They're the same one. The cost of producing knowledge collapsed, while the cost of validating it, attributing it, and trusting it stayed flat.

The tools are already moving ​

While the argument played out, the tooling kept advancing. InternLM released Intern-S2-397B, its most capable multimodal foundation model for scientific intelligence and long-horizon agents. It's a good lens on where this is heading.

397B parameters means this is not a laptop model, and not even a single-GPU model. You consume it through an API or run it on a multi-node cluster. Size isn't the whole story, though. The training approach is. Intern-S2 learns directly from raw pages of scientific literature, modeling symbolic semantics and visual relationships in a shared representation space, with no intermediate parsing step. It reads a paper's layout, equations, and figures the way a pair of human eyes does, instead of losing half the information to an OCR pipeline.

That's a direct answer to the flood. If no human can read 447 papers a day, the next best thing is a model that reads the whole page, figures and tables included, rather than a text-extraction husk. Intern-S2 also trained on reinforcement learning tasks across more than twenty scientific domains, jointly, plus long-horizon agent RL inside sandboxed environments. The result is a model aimed at things like biomolecular interaction design and material structure generation, work that looks nothing like benchmark multiple-choice questions.

PressureEvidence from this weekWhat it breaks
SpeedNinety years of human effort vs. 88 hours of machine time on Navier-StokesThe race replaces the process. Announced progress becomes a target
Cost asymmetry~$22.5M of compute for one proof, spent in under a weekIndividuals and small labs can't compete on finishing
Volume447 cs.LG papers in one day, against a ~200/day baselinePeer review has no capacity. Reading groups can't keep up
AttributionThe Buckmaster-Alpöge dispute, with 25 Fields Medalists involvedTrust decays into defensive secrecy

Key numbers from the week:

  • 88 hours for a 165-page, mechanically verified proof of a problem open for ~90 years
  • 300B output tokens generated by ~10,000 parallel agents, nearly all discarded
  • ~$22.5M in compute, a week-long sprint at industrial scale
  • 447 ML papers on arXiv cs.LG in a single day
  • 397B parameters in Intern-S2, the new open scientific foundation model

Common Pitfalls ​

Watching this week's coverage, I've seen several specific mistakes get repeated. Name them and they get easier to avoid.

First, treating "mechanically verified" as "settled." Formal verification proves each step follows from the previous one inside a chosen framework. It doesn't prove the framework captured the original question, and it doesn't replace months of community scrutiny. The proof is being read line by line right now. The verdict isn't in. Report it as under verification, not as solved.

Second, dismissing the credit dispute as mathematicians protecting egos. Attribution is how mathematics transmits trust across generations. When credit becomes a bargaining chip, the knowledge network itself starts to rot. The Fields Medalists' letter wasn't about feelings. It was about the mechanism that lets one generation build on another.

Third, responding to the flood by trying to read more. The lever is changing what counts as a contribution. Verification, reproduction, curation, and honest negative results are the bottleneck now, not generation. A bigger reading queue is solving the wrong equation.

Fourth, racing labs on speed. If you're close to a hard result, Buckmaster's case is the warning: incremental public progress is now a target for entities that can spend $22.5M in a week to close the gap. Your realistic moves are secrecy until submission, time-stamped pre-registration, or explicit partnership before you publish.

Fifth, skipping the feeling to get to the advice. "Just learn to use the tools" is true, and it's beside the point for the mathematicians this week. People can tell when you skip the grief to hand them a productivity tip. The transition goes better when institutions acknowledge the loss than when they tell the grieving person to reskill.

The grief was never about the job ​

Underneath the dispute and the flood is the part the numbers can't carry. Most of those mathematicians never expected to solve Navier-Stokes. They knew the odds; the problem had eaten ninety years of brilliant people. They reached anyway, because a hard question organizes a life. It gives your attention somewhere worthy to go, and it makes you into a particular kind of person, patient, humbled, awake to something larger than your own career.

The machine got the answer in 88 hours. It did not get the twenty years of being someone who was reaching for it. That isn't a consolation prize. The output was always a page. The life was the reaching, and the reaching wasn't wasted just because something else reached the end faster.

What the community is saying. The most honest discussion I found this week happened in the comments under the essay about that morning, and two voices kept pulling at each other. I've felt both at different points of my own career. One is the person on the other side of the door: no degrees, no formal training, jobs that gave nothing back. The same machine that emptied a mathematician's morning is the only reason anything is in mine at all. For that version of me, the becoming didn't end. It moved, and it multiplied. The other voice is the worry for exactly that person: access is being turned into a service, and when the economics tighten, the first people cut off will be the ones who needed it most. Both mornings hold at once, and neither cancels the other.

One Thing to Remember ​

The proofs are verified by machines now, but the trust that makes a result matter is still built by people, slowly, the way it always was. That trust is the scarce thing, and no model has produced it yet. The answer was a page. The becoming was yours. The machine can take the output; it can't repossess the years that shaped you into someone who could reach for it at all.

The Bottom Line ​

Three things follow from this week, and they apply well beyond mathematics.

If you're a researcher working on a hard, famous problem, treat public partial progress as a strategic decision, not just a publishing one. A compute-rich lab can sense you're close and finish in a week, as Buckmaster found. Pre-register with timestamps, or hold your cards until submission. The old rhythm of announced incremental results assumed nobody could outrun you a week later.

If you're building research infrastructure, put your resources into verification and attribution tooling, not generation. The 88-hour proof is the cheap part of this decade. Formal-verification toolchains, provenance tracking, and pre-registration rails are where the system broke, and every lab in this fight needs them.

If you're a practitioner watching from outside research, don't race the machine at its own game, and don't let anyone skip your grief with a productivity tip. Anchor your identity in the reaching, the judgment, the relationships, the things that took years to shape. One thing to watch: the credit-and-verification layer for AI-produced science will consolidate fast, probably within a year, because every party in this fight needs it to exist.