166 pages. Ten thousand agents, 88 hours, roughly 130 billion output tokens. The Lean code compiles. Two weeks after OpenAI announced it had settled the Navier-Stokes problem, NPR went and asked the mathematicians what they had learned from it, and the answer was close to nothing. Nobody thinks the proof is wrong. Nobody can read it.
The paper is not written for humans.NPR
That is Javier Gómez-Serrano at Brown, who uses these models in his own work and is not the skeptic in this story. His complaint is specific: the paper does not say which parts are important, which parts are routine, or how the idea feeds anything else. James Maynard at Oxford put the stakes in one line: it was never only about answering the problem, it was about the human understanding behind it. Correct, complete, and useless to the people it was supposedly for.
The same day, in a different world, Colin Breck published an essay called I don't want to read what you didn't write, and it went to the top of Hacker News on the strength of a complaint every engineer now recognizes. People build something with an agent, then have the agent write the design document afterward. Pull request summaries written by machines for machines: this was changed to that, these were split, tests were added. Rich in detail, empty of the questions a reader actually has. Why are we doing this? How risky is it? Where do you want my input? Breck's verdict on the documents is two words. Unreadable. Inhumane.
Put the two side by side and they are the same document. Gómez-Serrano's questions about the proof are Breck's questions about the PR, nearly word for word. The pattern is not that AI writes badly. It is that generation stopped being the expensive step, and the thing that replaced it, a human who has to understand the output well enough to carry it forward, did not get any cheaper. The proof is finished for the machine. It is unfinished for mathematics, because in mathematics the artifact is the understanding, and the understanding is still sitting in 166 pages that nobody has extracted.
Breck names the mechanism. When you are the one who prompted the model, its output is useful, because you already hold the context: the constraints, the source, the logs, the call and answer that produced it. Send that same text to someone else and they have none of that. They are forced, in his phrase, to read exhaustively, peering into the internals of a machine hoping to reconstruct why any of it matters. For the Navier-Stokes paper there was no prompter to ask. Ten thousand agents ran the search. No human held the context, so no human can supply the introduction.
The proof is finished for the machine. It is unfinished for mathematics.
The skeptic's answer is that correctness is the point and prose is a detail. Lean compiled; the theorem is true; if the paper is ugly, have a model rewrite it. Concede the first part in full. The theorem is true, and a true theorem is a real asset that did not exist three weeks ago. But rewriting for a reader means knowing which of the 166 pages are load-bearing and which are routine, and that judgment is exactly what Breck says the output lacks and the prompter supplies. He tried the reverse direction on his own paper, asking the model to write a paragraph from the source code it had been given, and reports it was never valuable. Not once. The one section it wrote well, and he shipped unchanged, was the abstract, the most mechanical part of the paper. Summary the models have. Perspective they still borrow from whoever is in the loop, and here nobody was.
Look at what the mathematicians did about it. On Sunday Terence Tao's blog carried the announcement of an Advisory Group on Mathematics and Artificial Intelligence: Gowers, Hairer, Witten, Vakil and five others, hosted at the IAS, unpaid, with no decision-making power, whose stated purpose is to advise the labs on the responsible presentation and release of results. Their first task is written down. OpenAI has a large number of further results from the same internal model, and the group is to recommend how to coordinate releasing them. Nine of the best mathematicians alive have volunteered to be the rate limiter on a queue of proofs, because the constraint is no longer producing theorems. It is producing readers.
Tao's post carries one dry note at the top: it was drafted in another format and converted using AI. Even the committee formed to manage machine-written mathematics could not quite avoid it, and that is the honest picture. The text will keep coming. A proof compiled, a PR merged, a design document exists, and until a person has read it and can say what it means, nothing has been learned. Written for no one is not the same as written.