OpenAI releases 722 math papers from an internal AI model
OpenAI says an internal model produced 722 math papers with modest compute, raising the prospect that more chips could speed mathematical discovery.

OpenAI released 722 math manuscripts on October 6, grouped into 372 families of results that it says solve or advance major open problems. As The Rundown reported, an unreleased internal model produced the work.
OpenAI puts the average compute at the equivalent of roughly three hours of ChatGPT Pro thinking per result. That combination of volume and compute suggests a sharp acceleration in mathematical discovery, if the claims hold up.
What OpenAI released
OpenAI’s repository says it posed about 4,000 problems. The 372 families group related work, including companion arguments, consequences and alternative proofs. Several manuscripts can therefore belong to the same advance.
The company also published a catalogue of papers with main results formalized in Lean, software that checks a proof’s logical steps. Formalizing a result gives researchers a way to check the reasoning encoded in the proof.
Among the boldest claims is a proof of the “quasi-Riemann hypothesis,” a weaker version of the Riemann hypothesis, the famous problem about how prime numbers are spread out. OpenAI’s scope note for that work says the formalization excludes the paper’s later applications.
The reported computing effort differs sharply from OpenAI’s earlier math announcement. In its September 8 account of a claimed Navier–Stokes solution, the company described roughly 10,000 concurrent agents reaching the result after about 88 hours. Lean formalization and verification took another 17 hours. The different problems and workflows limit any direct comparison of efficiency.
The release has also drawn complaints from mathematicians. WIRED reported on October 6 that meeting attendees understood OpenAI to have promised to space out its releases. A company spokesperson said OpenAI was unaware of that assurance.
Why it matters
If these results hold up, math research could scale much more closely with computing capacity. Hundreds of results at a few hours of reported compute each suggest a way to expand discovery through repeated model runs. Research groups could pursue more questions at once as they add computing power.
For mathematicians, that could shift more effort toward choosing worthwhile problems, judging the significance of results and developing their consequences. Expert judgment would help direct the growing volume of machine work toward questions that matter to the field.
Verification would have to grow alongside generation. Lean can check the logical steps of a formalized theorem, but the quasi-Riemann example shows that a paper can extend beyond what has been checked. OpenAI itself warns that unformalized results may contain issues. As output grows, checking the reach and limits of each proof could become a larger part of mathematical work.
Access also shapes how quickly this change can spread. OpenAI says it is working toward releasing the model. The ChatGPT Pro figure describes computing effort, while the system that produced the results remains internal. Researchers can assess the public manuscripts and proofs now. Access to the model would let them test whether the approach transfers to their own questions.
The release leaves open how reliably extra compute produces useful new results. Judging the full cost also requires knowing whether OpenAI’s average includes failed attempts, formalization and human input. Those factors will help determine how quickly the promise of hundreds of results can grow into a repeatable research process.
Sources & further reading
This story builds on reporting from The Rundown newsletter on October 7, 2026.