Eleven days. 13 million lines of Lean. 29,500 intermediate theorems. On Thursday Anthropic published the first complete computer-checked proof of Fermat's Last Theorem, written largely autonomously by a swarm of Claude agents on a platform called Prove2Me. Kevin Buzzard, who has spent two years and a £1M EPSRC grant leading the human effort to do the same thing, compiled the repo, ran the comparator, and confirmed it checks out.
Then he wrote the sentence that matters more than the announcement.
Note that mathematically this work of anthropic tells us essentially nothing: I am on record as saying that I am 99.9% sure that the proof of FLT is OK, and most people in the number theory community are 100% sure.Xena Project
That is not sour grapes. It is the precise location of what changed. Nobody doubted Wiles. What was expensive was the certainty, and certainty just got cheap: a proof the community expected to take years to formalize, with an 86-page blueprint for the first phase alone, now exists as an artifact Lean can check from three axioms. The bottleneck in mathematics was never producing claims. It was the years of refereeing between a claim and a result people build on. Helfgott's 2013 proof of weak Goldbach is still under review. That is the queue this collapses.
But look at what the artifact is. The proof is over 5x the size of Mathlib, the entire community library it stands on. Buzzard reports it takes nearly 20 times as long to compile as Mathlib on a 96-core machine, and that Lean gets sluggish navigating it even with 500 GB of RAM. Anthropic's own footnote concedes it is likely much longer than it needs to be. Roughly 7% of the non-boilerplate lines are residue from failed early attempts that lost track of the project's state. This is a proof in the sense that a core dump is a description of a program. It certifies. It does not explain.
Which is why Buzzard says his job is not done. The part of his grant Anthropic did not touch is the one he calls most important: a dynamic document that lets humans explore the modern proof. Anthropic formalized the 1995 Darmon–Diamond–Taylor exposition, not the current Khare–Taylor route, and it produced no such document. His guess is that it never will. The trust layer got automated. The understanding layer is still a person in Wales going through a thousand unread emails.
This is a proof in the sense that a core dump is a description of a program. It certifies. It does not explain.
The honest objection: so what. Verification was the actual problem, and verification is what Wiles needed in 1993 when a reviewer's question opened a gap it took him a year to close. Mathematicians can keep writing expositions for each other; the machine handles the drudgery. Anthropic says as much, that a formal proof should sit alongside a human write-up, not replace it. Fair. But the economics of the two artifacts just diverged. A verified proof now costs 6 billion output tokens, which David Jao priced at about $300,000 on Buzzard's comment thread, or a third of that at Anthropic's margins. A readable exposition still costs a Buzzard: five years, one million pounds, and someone who cares whether a 21-year-old can follow it. When one output gets a thousand times cheaper and the other does not, the cheap one is what gets produced. Sylvain Kalache used the phrase comprehension debt this week for what happens to engineers when agents resolve every incident. Mathematics just took out its first loan.
Buzzard's next task is the right one. He wants the machines pointed at the Langlands program to check whether his paranoia about it is justified, ruthlessly flagging every argument that leans on something "known to the experts." That is what autoformalization is for: not proving Fermat again, but finding out what the literature has been assuming. The theorem was never in doubt. The corpus is. Now there is a tool that can tell us which parts.