OpenAI Released Hundreds of Math Papers Written by an AI Nobody Can Test. Mathematicians Want Proof.
On October 6, 2026, OpenAI published a public GitHub repository of mathematical research produced by an internal AI model it has not released. The first release contained 722 manuscripts grouped into 372 "families" of results on open research problems; after early revisions, OpenAI's own catalogue now lists 719. The company says roughly 42% of the headline results come with Lean formal proofs, computer-checked arguments that can be verified mechanically.
It may be the largest single release of AI-generated mathematics yet. It is also one of the most contested: researchers say they cannot properly judge results from a model they cannot run, and a group of 25 Fields Medal winners had already warned last month that AI labs are treating famous problems as leaderboard targets.
OpenAI has released hundreds of mathematical research papers written by an artificial intelligence model that no one outside the company can use. On October 6, 2026, it published a public GitHub repository containing 722 manuscripts, grouped into 372 families of results on open mathematical problems, all produced by an internal model it has not released. After early revisions, the company's catalogue now lists 719 manuscripts.
The scale is unprecedented. OpenAI says the model was given roughly 4,000 problems, spending about three hours of ChatGPT Pro-level thinking compute on each result, and that nearly every paper came from a single prompt to a single AI agent. The results cover areas from number theory to mathematical physics, and OpenAI published summaries of the model's reasoning for ten of them, including work on the irrationality exponent of π and the Mahler conjectures.
Not all of it has been checked. OpenAI says about 42% of the headline results come with Lean formal proofs, computer-verified arguments that leave no room for error, and warns that "some of the unformalized results could have issues." It has promised to correct mistakes quickly while keeping earlier versions public.
The reaction from mathematicians has been sharply divided. Some see a remarkable demonstration of what AI can now do. Others say they cannot properly assess claims from a model they cannot test, and object to hundreds of results arriving at once, without named authors or the usual process of write-ups and peer review. NYU's Tristan Buckmaster said he doesn't think OpenAI did its "due diligence," while MIT's Andrew Sutherland said the claims should be treated as unverified until others can reproduce them.
The release came weeks after 25 Fields Medal winners signed a declaration titled "A Severe Misalignment of AI in Mathematics," criticising AI labs for treating famous problems as benchmarks to beat. OpenAI says it will fund workshops and conferences on AI-produced mathematics and is working to release the model responsibly. Whether it does, and how many of the 719 papers survive independent checking, will determine whether this goes down as a breakthrough or a cautionary tale.
Why It Matters
Mathematics is one of the few fields where an AI's claims can, in principle, be checked with certainty: a proof is either correct or it isn't, and formal tools like Lean can confirm it line by line. That makes it a genuine test of whether AI systems can do original research rather than recombine what they've read. If even a fraction of these results hold up, it is a real step in AI's ability to produce new knowledge.
But how the results were released matters as much as what they claim. Mathematical progress normally comes with named authors, careful write-ups, credit to prior work, and slow checking by peers. Hundreds of papers arriving at once, from a model outsiders can't test, puts the burden of checking on human mathematicians and raises questions about credit that the field hasn't settled.
Key Details
- Released by: OpenAI, public GitHub repository openai/math
- Release date: October 6, 2026
- Manuscripts: 722 at launch; 719 in OpenAI's current catalogue after revisions
- Result families: 372
- Formal verification: About 42% of top-line results have Lean formal proofs, per OpenAI
- Model: Unreleased internal OpenAI model, not available to outside researchers
- Compute per result: About three hours of ChatGPT Pro thinking compute on average
- Problems attempted: Approximately 4,000
Technical Analysis
According to OpenAI, the model was posed roughly 4,000 problems during the evaluation, with each result using about three hours of ChatGPT Pro-level thinking compute on average. Outputs were grouped into families, a principal result plus companion arguments, consequences, or alternative proofs, and filtered for significance. OpenAI says nearly every paper came from a single prompt to a single AI agent, and that it ran these evaluations because the model had "saturated" its existing math benchmarks.
The topics span pure and applied mathematics. OpenAI released abridged reasoning summaries for ten results, including the irrationality exponent of π, the Mahler conjectures, Kaplansky's direct-finiteness conjecture in characteristic two, and the three-dimensional relativistic Vlasov–Maxwell system. Two results didn't follow the standard procedure: work on a zero-free region for the Riemann zeta function, whose write-up a human edited for readability, and a proof of the Hodge conjecture for CM abelian varieties.
OpenAI is candid about the limits. "Some of the unformalized results could have issues," the repository says, adding that it will fix problems quickly and keep earlier versions public. The drop from 722 to 719 manuscripts shows that process has already begun.
Expert Analysis
The strongest reactions have been about verification and process. Tristan Buckmaster of New York University, who had been working on the Navier-Stokes equations before OpenAI's claimed result there, doubted the results had been properly checked: "I don't think they've done their sort of due diligence at all." Andrew Sutherland of MIT said single-agent claims should be treated as unverified until others can run the model and reproduce them.
The Advisory Group on Mathematics and Artificial Intelligence (AGMAI), hosted by the Institute for Advanced Study, which OpenAI consulted, issued guidelines on September 29 asking labs to stop testing hard problems on private models and to publish prompts, time, and compute for each result. It warned that "the use of proprietary internal models by AI labs to do mathematical research risks creating a two-tier system." After the release, AGMAI said its advice was not an endorsement, adding: "This release is the beginning, not the completion, of the process of human understanding."
Earlier, on September 11, 25 Fields Medalists, including Terence Tao, Manjul Bhargava, Maryna Viazovska, and Peter Scholze according to reports, signed a declaration titled "A Severe Misalignment of AI in Mathematics." It responded to OpenAI's September claim about the Navier-Stokes equations, and argued that AI companies treat famous problems as benchmarks to beat rather than milestones of understanding, and that rushed announcements skip the write-ups, credit, and community scrutiny real progress needs. The signatories also acknowledged AI's potential to accelerate genuine understanding if guided responsibly.
Industry Impact
For mathematicians, the immediate effect is a large checking burden and an unresolved question of credit: how to cite, build on, or referee results produced by a system that has no named human author and can't be inspected.
For the AI industry, it raises the bar on how capability claims are made. OpenAI's own research lead, Dan Roberts, said the proofs were a byproduct of testing internal models. Rival labs will be under pressure both to match the results and to answer the verification and transparency questions this release has made unavoidable.
For everyone else, it is a preview of how AI may soon contribute to science more broadly, and of the friction that follows when machine-generated research arrives faster than humans can check it.
Future Outlook
The following is analysis and prediction, not confirmed fact.
The next few months will decide how this is remembered. Expect independent mathematicians to work through the formalized results first, since Lean proofs can be checked quickly, and to report errors in the rest. OpenAI says it will fund workshops and conferences on AI-produced results and is working to release the model responsibly; whether outsiders can actually run it will be the key test of the claims. If a substantial number of results survive scrutiny, this will be seen as a turning point for AI in research, even as arguments about credit and process continue.
Key Takeaways
- OpenAI published 722 AI-generated math manuscripts on October 6, 2026, now 719 after revisions, grouped into 372 families of results.
- They were produced by an unreleased internal model, using about three hours of compute per result across roughly 4,000 problems.
- About 42% of the top-line results have Lean formal proofs; OpenAI says unformalized results could contain errors.
- Mathematicians including NYU's Tristan Buckmaster and MIT's Andrew Sutherland say the results remain unverified until others can reproduce them.
- The release follows a September declaration by 25 Fields Medalists criticising how AI labs approach famous math problems.
Frequently Asked Questions
What did OpenAI release in October 2026?
On October 6, 2026, OpenAI published a GitHub repository (openai/math) of mathematical manuscripts produced by an unreleased internal AI model: 722 at launch, grouped into 372 families of results on open problems, now listed as 719 after revisions.
Are the AI math proofs verified?
Partly. OpenAI says about 42% of the headline results have Lean formal proofs, which can be checked by computer. The rest have not been formally verified, and OpenAI says some could have issues.
Can researchers use the AI model that wrote the papers?
No. The model is internal and unreleased, which is a central criticism: outside mathematicians cannot test it or reproduce its results.
What did the Fields Medalists say about AI and maths?
In September 2026, 25 Fields Medalists signed "A Severe Misalignment of AI in Mathematics," warning that AI labs treat famous problems as benchmarks and that rushed announcements skip proper write-ups, credit, and peer scrutiny.
Source: OpenAI (GitHub: openai/math), The Next Web