Four thousand problems. About three hours of ChatGPT Pro compute apiece. 722 manuscripts in 372 families. On Tuesday OpenAI put the lot on GitHub, all of it produced by an internal model it hasn't released, and the Hacker News thread passed 900 comments by morning.

The first question was whether any of it is true, and OpenAI answered it more thoroughly than it had to. It consulted an independent advisory group of mathematicians at the Institute for Advanced Study and followed much of its checklist: reasoning summaries, compute estimates, the number of problems attempted. Its formalization catalogue lists 162 papers whose main result is written out in Lean, a language in which a computer checks every step. One of them is what the manuscript map calls the quasi-Riemann hypothesis: the zeta function has no zeros with real part above 7/8. For those 162 papers, checking is now a compiler's job. For the other 560 or so, the README warns that some unformalized results could have issues and promises quick fixes. That reads like a changelog, not a journal.

So for a growing share of these papers, a machine can settle whether the result is true. What it can't supply is understanding, and mathematics depends on understanding more than on any single result.

That group's published norms draw on more than 600 responses from working mathematicians, and they start with the field's oldest rule:

the authors of a paper should understand the mathematical argument of that paper
Advisory Group on Mathematics and AI

The rule's other two clauses are that authors verify the work themselves and take responsibility for it. Lean now handles verification, and more reliably than a tired referee. OpenAI, the only author listed on every paper, can take responsibility as an institution. Neither of them can give the talk, the hour at a whiteboard when someone asks why the proof works.

Every engineer has inherited this artifact: a module that passes every test and that nobody on the team can explain. It works. You can call it. You can't extend it, because extending it means knowing why it works, and the tests don't tell you. Mathematics is mostly extension. The next result usually comes from the last proof's method rather than its statement. A formalized theorem gives the field the statement with high confidence and gives the method to no one.

A formalized theorem gives the field the statement with high confidence and gives the method to no one.

Mathematics has absorbed machine proofs before. In 1976 Appel and Haken proved the four color theorem with a computer checking cases no person could, and the field eventually accepted it. But they designed the strategy and the computer only did the casework, so the reason the proof works stayed with people. Here the model supplied the strategy as well. The README also admits that some outputs build on the model's own earlier results. The compounding is already happening, just inside a model nobody outside OpenAI can use.

The advisory group saw this coming. It asked labs to stop testing open problems on proprietary models at all, and it warned of a two-tier field in which the labs outrun everyone else. OpenAI's response is to fund workshops on understanding AI-produced results and to work toward releasing the model. The workshops pay to move understanding back to people after the fact. Releasing the model is the only thing that would let a mathematician question the author directly.

722 papers. A compiler vouches for 162 of them. The one author that could explain why they work is a model nobody outside OpenAI can ask.