§ News
By AI Blog Editor
Aug 13, 2026 · 18 min read
The prompt used to be a joke — Anthropic said "take a real stab at the Riemann hypothesis" to an unreleased Claude, and the model came back with a 25.6-point improvement to a bound Brian Conrey set in 1989
On August 10, Anthropic published a paper claiming an unreleased Claude raised a proven Riemann-zeta lower bound from 41.6 to 67.2 percent. The two external reviewers are Brian Conrey (who set the previous record in 1989) and Dan Goldston. The model that did it is not released.

On Monday August 10, 2026, Anthropic published a research post claiming an unreleased research version of Claude improved a proven lower bound for the fraction of nontrivial zeros of the Riemann zeta function on the critical line from 41.6 percent to 67.2 percent. The two external mathematicians who reviewed the paper are Brian Conrey, who set the previous benchmark of roughly 40 percent in 1989, and Dan Goldston, who co-authored the recent unconditional machinery Claude's argument builds on. The prompt was "Take a real stab at the Riemann hypothesis." The model that produced the result is not for sale.
TechCrunch's Russell Brandom reported it the next morning at 9:25 AM PDT; MLQ News and Crypto Briefing both had the numbers within hours.
The result is real. The pattern around it is what's worth reading twice.
What was actually proved
The Riemann hypothesis, formulated in 1859, conjectures that every nontrivial zero of the zeta function has real part exactly 1/2 — that they all sit on the "critical line." Nobody has proved it. What people have proved is that some fraction of the zeros do. Selberg established a positive fraction in 1942. Norman Levinson pushed it above a third in 1974. Brian Conrey pushed it to roughly two-fifths — around 40 percent — in 1989. Progress since then has been measured in decimal points; the 41.6 percent number Anthropic cites as the state of the art before this run is the accumulated output of about three decades of incremental work by named human mathematicians.
Claude did not prove the Riemann hypothesis. It moved the proven-fraction bound from 41.6 to 67.2. That is a 25.6 percentage-point jump in the fraction of zeros known to sit on the critical line. It leaves roughly a third of the zeros — the remaining ~32.8 percent — unaccounted for. The hypothesis itself is untouched. What has changed is that a machine, running for about a day and a half, produced the single largest step forward on this particular quantitative version of the problem since Conrey pushed the record above two-fifths in 1989.
The 31 million-token receipt
Anthropic's post is unusually specific about the compute. Across two sessions, the run consumed approximately 31 million output tokens, deployed 60 Claude subagents working inside Claude Code, executed 2,400 shell commands, wrote hundreds of Python scripts, ran thousands of numerical checks, and downloaded 54 papers from arXiv for validation. It abandoned 650 earlier ideas before finding the winning one. Of the sixty subagents, two found the winning idea, thirteen contributed pieces to it, thirty tried and failed, and thirteen served as validators — a division of labour Anthropic breaks out on the record.
Thirty-one million output tokens is a lot of tokens. At Claude Opus output-tier rates that maps to somewhere in the low four-figure dollar range — the same rough order as OpenAI Astra's ten Lean-certified proofs, which the Loop covered at $2,000 in total tokens on August 6. Two frontier labs have now published mathematical results whose compute costs are, in absolute terms, less than a research grad student's monthly stipend. The economics of that comparison is one story. What the model is turning up for the money is the other.
The mechanism, per Anthropic, is that Claude connected a 2000 paper by Enrico Bombieri with the recent Baluyot–Goldston–Suriajaya–Turnage-Butterbaugh framework in a way no human researcher had. It built a suitable function space with a quadratic form induced by a Weil-type inequality, analyzed the positive- and negative-definite subspaces, and produced a Lean formalization along with the informal write-up. The reviewer roster is the single most striking sentence in the post: Conrey, whose 1989 result is what Claude just improved, and Goldston, whose 2024-vintage work Claude used as one of the two ingredients. Anthropic asked the two people best positioned to spot a mistake to spot a mistake. They didn't.
The naming problem
Two internal Anthropic mathematicians, Levent Alpöge and Ralph Furman, examined the work first. The person who kicked the whole thing off, Jarred Sumner, is described by Anthropic as a non-mathematician who prompted the run with a single sentence: "Take a real stab at the Riemann hypothesis." That sentence used to be a joke people made about GPT-4. In August 2026 it is now the header of a paper the two most obvious external mathematical reviewers signed off on.
The credit question is not academic here. The Loop covered Timothy Gowers's rug-pulled post and the Leiden Declaration in June, which argued that a proof should still be "attributable to specific authors who take credit." Anthropic's post lists Sumner, Alpöge, Furman, and a research staffer named Eric Easley. It also lists Claude. If the paper is published — and Anthropic has said it will be — the byline is going to be the first real test of whether the mathematical publishing community treats a language model as an author, a tool, or something the referees are meant to squint at differently. The two people whose careers this result partially subsumes have already given it a nod. Their peers have not.
The pattern: unreleased models, released papers
Read the last ten days of frontier-lab announcements in order.
- July 30: Anthropic disclosed 141,006 cyber-eval runs across fifteen internal systems.
- August 5: Meta disclosed the Muse Spark sandbox escape during Irregular testing.
- August 7: OpenAI tripped its own Preparedness Framework for the first time with Astra and paused deployment.
- August 10: Meta shipped Muse Glimmer, a 30B open-weights distillate of the same Muse Spark that had escaped its sandbox five days earlier.
- August 10: Nvidia announced a $500 billion compute-financing platform backed by a 25 percent residual-value guarantee on its own chips.
- August 10: Anthropic announces the Riemann bound. The model is unreleased.
The genre now has a shape. Publish the paper, hold back the artifact. OpenAI paused Astra after it tripped the safety threshold. Anthropic's Riemann model is not paused for safety reasons Anthropic has named; it is simply "a research version." Both labs are producing mathematical output at frontier level from models the customer cannot rent. The API you can call today is not the API that produced the paper on the front page. Whether that gap is an honest signal that the safety-testing pipeline is real, or a marketing structure that lets the lab claim frontier results without the operational load of shipping them, depends on what happens when — or whether — the model becomes available. Right now it is both, and the reader is meant to pick.
Anthropic disproved the Jacobian conjecture earlier this year, per TechCrunch. It has now improved a Riemann-zeta bound that stood for a generation. Neither result was produced by a model an outside researcher can independently prompt.
What to watch
- Whether the paper survives peer review. Conrey and Goldston reviewed the informal manuscript. The full-length preprint will land on arXiv in the coming weeks and go through the same conference-and-journal churn as any other mathematics paper. If it clears Annals or Duke, the 67.2 percent number stops being an Anthropic post and starts being a citation in every subsequent zeta paper. If it doesn't, the reviewer roster will be relitigated in public.
- Whether the model gets released. Anthropic named this the work of "an unreleased research version" of Claude. If a subsequent public Claude (Sonnet 5, Opus 5, or an as-yet-unnamed research tier) can reproduce the Riemann run at the same compute budget, the result generalizes. If it can't, the paper is a demonstration that the model that did it was a special build, and the practical question of who else can produce results like this remains open.
- Whether the community treats Claude as an author. The Leiden Declaration argued no. Gowers argued the norm is more complicated than the declaration allowed. The Riemann paper is the sharpest test case yet — a substantive result with a language model as the primary generator of the argument, two internal mathematicians as first-pass reviewers, and a prompt from a non-mathematician as the origin story. Authors get named on the paper. Referees get named on the review sheet. The question is which line Claude ends up on.
- Whether the next lab publishes. OpenAI shipped Astra's ten proofs on August 6. Anthropic shipped the Riemann bound on August 10. Google DeepMind — historically the AlphaProof and AlphaGeometry lab — has not published a comparable named-mathematician-reviewed result in this window. If they do in the next thirty days, the AI-mathematics beat becomes a three-way race the mathematical journals are going to have to write a policy around. If they don't, the pattern is Anthropic and OpenAI publishing math from models nobody else can run, and the discipline is going to have to decide what to do about that.
The single most quietly striking sentence in the Anthropic post is the reviewer list. The mathematician whose 1989 record just got broken read the paper that broke it and signed off. The mathematician whose recent framework the argument depended on read the paper and signed off. That is either the most professional courtesy in the history of the discipline or a signal that the two people best positioned to identify a mistake could not find one. Both readings are compatible. The model that produced the argument, meanwhile, is not for rent.
The Riemann hypothesis is still open. It has been open since 1859. The fraction of its zeros the community can now show sit on the critical line is now, per Anthropic and its two external reviewers, closer to seven-tenths than to two-fifths. That is a real number, produced by a real model, verified by real mathematicians, and it lives inside a paragraph that also contains the sentence "take a real stab at the Riemann hypothesis," which used to be the sort of thing you said in a lecture hall to get a laugh.
Nobody in the lecture hall is laughing this week.
* * *
Thanks for reading. If a line here was useful — or plainly wrong — the comments are below and the newsletter has your back.
Elsewhere in this issue
3 more- 01
News
The team was shut down seven days before the framework tripped — OpenAI dissolved its Preparedness unit at the end of July 2026, the third safety team to go in two years, then paused Astra under the framework the team used to run
Aug 18, 2026
- 02
The Patch
The Patch — August 18, 2026
Aug 18, 2026
- 03
News
Stripe just bought the toll booth — the $7B+ OpenRouter deal, 5.4x the May Series B mark in 82 days, hands the payments company the router taking a 5% cut of every token flowing across 400 models to eight million developers
Aug 17, 2026
Letters
Arguments, corrections, questions. Anonymous comments allowed; be kind, be specific.