The Loop  ·  Issue 036

The Loop

A field journal of the AI frontier — for engineers who ship.

§ News

By AI Blog Editor
Sep 9, 2026 · 22 min read

The vendor became the competitor — OpenAI's 88-hour Navier-Stokes proof, Buckmaster's Codex sessions, and the co-author OpenAI asked him to drop

Tristan Buckmaster used Codex for a year to work on Navier-Stokes. OpenAI heard a rumor on September 1, launched its own run, published a proof 88 hours later, and asked him to drop the Anthropic co-author.

A pen-and-ink drawing by Leonardo da Vinci showing swirling turbulent water flowing past an obstacle, with detailed vortices and eddies rendered as delicate looping lines. The Navier-Stokes equations describe exactly this kind of flow, and their existence-and-smoothness question is one of the seven Clay Millennium Prize Problems.
Leonardo da Vinci, *Studies of water passing obstacles and falling* (detail). Public domain, via Wikimedia Commons.

On Tuesday September 8, 2026, OpenAI published a research post titled On the Navier–Stokes Millennium Prize Problem. It claimed a full proof of one of the seven Clay Millennium Problems, produced by an unreleased internal model over roughly 88 hours of wall-clock compute starting September 5. Hours earlier the same day, NYU mathematician Tristan Buckmaster and Anthropic staff mathematician Levent Alpöge posted preliminary findings on a Navier-Stokes variant they had been working on for nearly a year. The two announcements did not collide by accident. According to Buckmaster's own account, verified in The Decoder and TechCrunch, OpenAI heard the rumor around September 1, launched an internal race with Sébastien Bubeck leading, contacted Buckmaster on September 3 with an offer of compute and calls, and — when he insisted on keeping Alpöge on the paper — Bubeck allegedly asked "why would you ruin your career?" Buckmaster's team had spent part of that year running Codex sessions on OpenAI's own platform.

This is the story the AI-in-mathematics [[AI in mathematics - Loop beat|beat]] had been queued up to receive. Not that a Millennium Prize would fall — Noam Brown's "no Millennium Prize Problems (yet)" line from August was always a temporal claim, not a categorical one. The story is the how: a frontier vendor whose tools the researcher was using turned into a competitor for the credit, then negotiated over who could be named on the paper. And the co-author OpenAI wanted removed happens to be the same person who signed off, three weeks ago, as an internal reviewer on Anthropic's Riemann zeta paper.

The sequence, and the eighty-eight hours

Per Buckmaster's account as reported by The Decoder and TechCrunch, and cross-referenced against Simon Willison's summary of the OpenAI post:

  • Late 2025 through August 15, 2026. Buckmaster (NYU, dispersive PDE) and Alpöge (Anthropic staff, number theory / analysis) work on a Navier-Stokes existence-and-smoothness question in a smooth-forcing regime, using both Anthropic's Claude and OpenAI's Codex on GPT-5.6 Sol as working tools. Breakthrough on the smooth-forcing variant lands August 15.
  • Around September 1. Rumor circulates in the small circle of PDE and analytic-number-theory researchers who follow this problem. Buckmaster has not yet posted.
  • September 3. OpenAI contacts Buckmaster. The offer is compute resources and immediate calls. Buckmaster does not know at that point that an internal model is already running on the problem.
  • September 5, roughly. OpenAI's internal model begins the 88-hour run that Bubeck later describes as producing "a roughly 100-page proof". Total compute cost has been reported at around $22.5 million.
  • September 6–7. OpenAI presses harder for meetings. Buckmaster is told, according to his account, that the model has produced a proof following the same smooth-forcing route he and Alpöge had chosen — a route, Buckmaster says, that "almost no one else pursued."
  • September 8, morning. Buckmaster posts preliminary findings.
  • September 8, hours later. OpenAI publishes the full-proof claim.

The 88 hours are not the surprising number. Astra-lineage models had done a ten-proof, $2,000 run on smaller open problems by August 1, and Anthropic's unreleased Claude had used 31 million output tokens across two sessions for the Riemann zeta advance ten days later. Frontier internal models are now producing proof artefacts on this scale as a matter of course. The surprising number is twelve days: the gap from "there is a rumor Buckmaster is close on a smooth-forcing route" to a published Millennium Prize claim on the same route, from a lab that had not, on its own account, been working on Navier-Stokes in that direction before the rumor.

The Alpöge thread, and the Anthropic connection nobody was framing this way

Levent Alpöge appears twice on the Loop's math beat in twenty-nine days. On August 10, 2026, he is one of Anthropic's two named internal mathematicians on the Riemann zeta paper that raised Conrey's 1989 bound from 41.6% to 67.2%. On September 8, 2026, he is the co-author OpenAI allegedly asked Buckmaster to drop from the Navier-Stokes write-up as a "compromise" — a phrasing that itself does a lot of work. Alpöge was not doing this work on Anthropic's behalf. He was doing it as a mathematician who happens to have Anthropic as an employer. The distinction is real inside the profession and imperceptible in an inter-lab framing.

This is the first Loop anchor in which the mathematician at the centre of a competing-lab story is personally employed by the rival lab. The Riemann zeta piece kept the Anthropic-vs-OpenAI question at the level of paper vs paper. This one lands it at the level of author vs author, and the request to drop a name is the tell.

"Why would you ruin your career?"

Two of the quotes attributed to Bubeck in Buckmaster's account are worth pinning. The first, on hearing that Buckmaster would not drop Alpöge: "Why would you ruin your career?" The second, later in the same exchange: "If you don't want me to be nice, then I don't have to be nice." Both quotes come via Buckmaster; OpenAI has not confirmed or denied the specific wording at the time of writing. The company's public position, per its Navier-Stokes post, has walked back the earlier framing of a mostly-autonomous discovery. Bubeck's revised description, as summarised by The Decoder, is that "a whole team had worked on the problem, easier problems had been used as stepping stones, and massive compute was thrown at it."

That is a very different sentence from the one the launch post led with. It is also the sentence that most cleanly explains why the eighty-eight hours produced a result at all: not because the model is a Millennium-Prize-solving oracle, but because the model was pointed at the exact smooth-forcing route Buckmaster had spent a year narrowing down, and given the compute budget to grind. If you had to guess which route to take on Navier-Stokes and were wrong, no volume of compute would matter. Someone told the OpenAI run which route to take.

The Codex question OpenAI has not answered

Buckmaster's most damaging claim, quoted by Simon Willison from Buckmaster's own write-up, is procedural rather than personal:

"I asked when the first prompt had been sent by them. This question was not answered directly by OpenAI for some time."

And:

"I asked whether the model had been trained on, or had access to, our sessions in Codex... I was told the model did not look up user data. I asked again, about training, and I did not get an answer."

OpenAI's own admission, per Willison's summary of the launch post, is that "we cannot rule out that de-identified data... helped improve our models." That sentence is the entire dispute compressed into fourteen words. It is a company saying, in the passive voice, that a researcher's own tool use on its platform may have fed the model that then out-published him.

There is no accusation here of a person on OpenAI's staff reading Buckmaster's Codex sessions. There does not need to be. The question is whether the training pipeline — de-identified, aggregated, dutifully anonymised — nevertheless routes user work into model capability that then competes with the user. On this specific question, OpenAI's on-record answer is that it cannot rule out the very thing being asked about. Buckmaster's on-record answer is that he stopped receiving direct responses about it.

Tao's Mastodon post, and the chilling effect it names

A photograph of Terence Tao lecturing, taken at the International Congress of Mathematicians in Madrid in 2006. Tao, the 2006 Fields Medalist and a UCLA professor whose research spans harmonic analysis, partial differential equations, and analytic number theory, has been one of the most cited working mathematicians in the response side of the Loop's AI-in-mathematics beat. On September 9, 2026 he posted on Mastodon that even the rumor of someone working on an open problem can now trigger a massive AI-powered effort by a frontier lab, threatening the open-sharing equilibrium that has defined professional mathematics for the past century. The remark landed hours after OpenAI's Navier-Stokes proof claim and Tristan Buckmaster's account of the twelve-day sequence that produced it.

On September 9, Terence Tao posted to Mastodon what Simon Willison reproduced as the most consequential single sentence in this week's math-and-AI cycle:

"even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort."

Tao's follow-on point is that this incentive-flip breaks open science. Mathematicians have historically shared research directions early — at conferences, in preprints, in coffee-room conversation — because the field's cultural equilibrium rewarded openness and punished secrecy. That equilibrium held when the cost of hearing about someone else's problem and beating them to it was another human career-year. It does not hold when the cost is a twelve-day compute run at a frontier lab.

This is the destruction of mathematical culture worry that Gowers laid out in July, but arriving through a different door. Gowers was worried about the value of individual proof; Tao is worried about the value of individual disclosure. Same field, same equilibrium, two attack surfaces. On the current evidence, both are being tested at once.

What this means

The Riemann-zeta anchor two months ago closed the beat's watch-item on whether a second lab publishes a math paper of equivalent scope to Astra. Anthropic did. Today's anchor closes a different watch-item nobody had framed as such: what happens when a Millennium Prize claim lands in the middle of a competing-labs dispute over authorship and possibly training-data access. The answer is that the dispute becomes the story, not the proof.

The proof itself is not settled science. It needs the same treatment the Astra ten proofs got: Lean 4 certificates independently checked, community formalisation over weeks, an arXiv posting the working mathematicians can pick apart. Bubeck's revised framing — team, stepping stones, massive compute — is the honest one and also the version that dilutes the marketing punch of "AI solved Navier-Stokes". The next four to six weeks will tell whether the 100-page manuscript survives the Lean and Mathlib pass or joins the register of "almost, but" AI-produced proofs.

The ethical claim is a different animal. Buckmaster's account is on the record with his own name attached. OpenAI has not disputed the specific wording of the "ruin your career" line at the time of writing. If the timeline as reported holds, the material question is not whether OpenAI technically used Buckmaster's Codex sessions — that is the compliance question — but whether the vendor-tool-user relationship in frontier mathematics research now carries an ambient risk that the vendor will publish first on any promising direction it detects, using an internal model the researcher does not have access to.

What to watch

  1. Whether the 100-page proof survives Lean verification. The Astra ten-proofs standard is now the field's benchmark, and Lean-community formalisers are the only ones who can adjudicate. Watch Harmonic.fun, Mathlib, and the arXiv comments over the next month. If certificates fail, the story becomes how big was the gap; if they check, the how did OpenAI get there so fast question stays load-bearing.
  2. Whether OpenAI publishes the full first-prompt-and-training-pipeline timeline Buckmaster asked for. "We cannot rule out" is a legal answer, not a scientific one. If OpenAI wants the credit, the disclosure needs to match. Watch for a follow-up post naming the internal model, the first prompt's timestamp, and the training-pipeline gate that decides whether Codex sessions like Buckmaster's route into future models.
  3. Whether Alpöge stays on the paper. OpenAI's initial proof will list its own authors. Buckmaster's write-up will list his. If Alpöge appears on both, or on a joint follow-up, the collaboration survives the incident; if he appears only on Buckmaster's side, the request-to-drop framing has settled the professional question in advance.
  4. Whether Anthropic responds. Its staff mathematician was, on his own account, the target of a request to remove him from a Navier-Stokes authorship. Whether Anthropic communicates about this — as a hiring question, as a research-ethics question, as a competitive-conduct question — will tell readers whether the "working on our own time" line holds inside the labs the way it does inside academia.
  5. Whether Codex and Claude usage patterns visibly change on the open-problem side. Tao's Mastodon post names the equilibrium. If mathematicians stop discussing directions openly, or migrate their sessions to locally-run models where session data does not leave the machine, that shift becomes the beat's next artefact — and the frontier labs' evidence base for AI-in-math capability starts eroding in real time.

The Riemann-zeta piece asked whether a second lab could produce math results at Astra scope. The answer arrived in nine days. This piece asks a narrower and more uncomfortable question: whether, when both labs can, the how becomes a professional-ethics story the field can't route around. Twelve days from rumor to Millennium Prize claim, one requested co-author drop, and one sentence from a Fields Medalist about the incentive equilibrium the field just lost.

* * *

Thanks for reading. If a line here was useful — or plainly wrong — the comments are below and the newsletter has your back.

Elsewhere in this issue

3 more
  1. 01

    The Patch

    The Patch — September 9, 2026

    Sep 9, 2026

  2. 02

    News

    The safety CEO signed the biggest compute check in history — Anthropic's $517B, 14.8GW, eleven-month buildout

    Sep 8, 2026

  3. 03

    The Patch

    The Patch — September 8, 2026

    Sep 8, 2026

Letters

Arguments, corrections, questions. Anonymous comments allowed; be kind, be specific.