On October 6, OpenAI put 722 mathematical manuscripts on GitHub and called it a release. They are grouped into 372 families. An internal model produced them. The company did not name the model. Scott Aaronson, writing the next day, reported that the same model had been tried on about 8,000 problems and had hit about 5 percent of them, after a single attempt on the order of three hours of Pro-level compute. OpenAI’s own note uses a smaller count of problems posed, about 4,000, and the same three-hour average for a typical result. Either way, the public was shown the hits. The misses stayed in the building. What landed on the field was a pile of machine drafts.
The argument that followed has been filed under the wrong heading. It is not a revolt against AI. Terence Tao wrote, the afternoon the drop was being reported, that AI can contribute to exposition, to the life of a community, and to the opening of new directions. Dana Moshkovitz, the complexity theorist whose career has run toward the Unique Games Conjecture, one of the claims in the pile, texted that a good future is one in which a person with an idea can ask a model to check it and carry it out. She also texted that the manuscript in front of her was so badly written she could not read it without asking another model for help. Those two sentences are the week. The tool is welcome. The invoice is not.
What five percent actually delivered
Five percent of 8,000 is 400. The catalogue OpenAI published is 372 families, later counted at 719 manuscripts after three were withdrawn. The arithmetic is close enough that Aaronson’s “about 5 percent” and the size of the public pile describe the same event: a wide sweep, a thin set of keepers, and then the keepers mailed as papers. OpenAI says it consulted the Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study, published protocols for revision and citation, and included Lean formalizations for many of the proofs, with more to come. A typical result, on the company’s account, used about three hours of ChatGPT Pro thinking. It promised workshops. It said it is still working on a responsible release of the model. It also said that future releases will need better citations, better exposition, and a presentation a reader can follow. That last sentence is a description of what this release was not.
The Verge’s count, as of October 8, is the quality-control record that arrived with the pile. OpenAI had said 300 of 719 manuscripts were formalized, about 42 percent. The public history already listed revisions to more than a dozen manuscripts, and the removal of three papers over a sign error that invalidated an argument. Formalization is a real check. It is a check on a minority of the files, performed in part after the files were already in the world. A sign error is a small mistake with a large meaning. The company shipped an argument, other arguments depended on it, and the correction was a deletion. That is what a draft does. It is not what a finished paper is supposed to do to the people downstream of it.
The manuscript a specialist could not read
Moshkovitz’s notes, which Aaronson published, are the closest thing this week has to a referee report, and they match the complaint exactly. The Unique Games manuscript felt, she wrote, like something written by someone on psychedelics. Much of it was unclear. It piled up citations to earlier work without explaining why those results could be used, including in the face of impossibility results. The proof invents what she called a completely new bizarre code, with a noise test, a crazy recursive construction, neither the long code nor the short code. “Some alien craziness.” The citations were often irrelevant. The paper was so horribly written that reading it required AI help. She asked Astra to assemble reasonable completeness and soundness claims for the noise gadget by combining sentences scattered through the manuscript.
Read that as a labor diagram. The model produced a claim. The person who has spent years on the conjecture could not tell, from the draft, which prior theorems it was actually applying. She had to hire another model, in the only sense that matters, to reconstruct the argument the first model had refused to write down in order. Aaronson added the sentence that should embarrass the lab more than any boycott: it appears that no human has understood just about any of these proofs yet. The race to do so has just started. He described the OpenAI method as a crazy race among people to digest and explain a messy machine proof, work that can be thankless, barely credited, competitive, and unfun. That is the job of a referee, done without a journal, without a decision, and without pay.
She is not asking for the machines to stop. In the same thread of texts she sketched a heavenly version of the field, if the vision stays human and the model helps check and implement it, and she wrote that there is a lot to learn from the aliens. The day after, on Aaronson’s blog, she pushed back on the metaphor of a lonely teleport to a mountaintop. For each of these problems there was already a community, a framework, and existing results, and the model had clearly built on them. The results go much farther. They do not finish mathematics. A researcher who talks that way is not an opponent of the tool. She is refusing to pretend that a pile of name-dropped priors, with no account of where they apply, is a proof she can sign.
Tao’s objection is about the soil
Tao’s own words that day belong in the same column. On the afternoon of October 6, as reports of the release spread, he described what he called Math 1.0: a breakthrough on an old conjecture sets off talks, workshops, and new collaborations. What is happening now, he wrote, is different. Problems are solved autonomously by prompters who have no interest in the broader field once the target is marked solved. Promising open directions are withheld, because people are afraid their own research will be scooped. A solved problem cannot be reverted. Solutions are being harvested at large scale, in a way he called unsustainable, leaving entire fields less fertile. Math 2.0, in his phrase, has to decenter raw problem solving and value progress more holistically: exposition, the building of a community, the finding of new directions. Then the sentence that makes the “mathematicians versus AI” headline false. He believes AI can contribute positively in all of those directions as well.
That is a defense of the subject’s metabolism, not a ban on the instrument. Harvesting is what you do to a field when you take the fruit and leave the work of cultivation to whoever still lives there. The fruit, this week, was 372 claimed results. The cultivation is the part Moshkovitz was doing at midnight with a second model: does this citation apply, is this construction even the one we know, can a person stand in front of a board and say what the new idea was. Tao’s point is that if the only thing that counts is the solved line on a list, the field stops producing the next list. People hide the open questions. Seminars get thinner. The byproducts of a proof, which are the methods and the failed paths and the neighboring problems, never get written down, because the prompter has moved on.
One lab paid for a digested paper
There was a second method on the table, and it arrived the evening before. Virginia Vassilevska Williams and Josh Alman posted an arXiv preprint, Aaronson reported, taking 3SUM down to about n to the 1.9992 and all-pairs shortest paths to about n to the 2.9995, against conjectures that had stood for half a century. The crucial idea, on his account, came from an Anthropic model. Anthropic then did not post the undigested solution. It gave the two of them the chance to write and announce a digested version, and it paid them. Aaronson is frank that this has its own problem: a company chooses which mathematicians get to be the emissaries. He is equally frank about the other problem. OpenAI’s method sets up the race to explain the mess.
I think the contrast is the whole design question, and it does not require a verdict on which lab is virtuous. A digested paper has an author who is willing to say, in public, what the idea is and where the prior work fits. An undigested dump has a corporation and a file tree. The second asks every specialist in range to become a grader. Kevin Buzzard, at Imperial, told New Scientist that the release included about thirty papers in his part of number theory, that about seven looked impressive, and that one was formalized in Lean. The others wait until an expert is motivated to read the text, or until a formalization exists. “Journalists are going to have to wait while the mathematicians do their job.” Ben Allanach, at Cambridge, asked whether any of this was entering human knowledge, and said that just checking or understanding a machine’s working is probably not as attractive as the work mathematicians actually like. That is a polite description of unpaid verification.
The boycott is a different claim
On October 7 a statement from the Association for Human Mathematics appeared as a guest post on Tao’s blog, reposted from the association’s own page. It says mathematicians did not ask for this work to be done. It says OpenAI ignored the advisory group’s premise that frontier labs should not test advanced problems on internal models. It says that releasing over 700 files at once is not a demonstration of scholarship, but a demonstration of power. It urges mathematicians to stop working with OpenAI and return to a science centered on human understanding. Aaronson’s update on October 9 is the right footnote. The association speaks for itself, not for the field. He thinks the communities should hold the labs to account, on how they publish and on the larger question of pace. He also thinks “no one should spoil our fun by solving our problems” is untenable. Open problems are there for anyone to try, including a corporation people dislike. A ban is not enforceable once the capability sits behind a subscription.
The advisory group’s own note, which Aaronson quotes, refuses both the press release and the boycott. Making the work public is a first step. The release is the beginning of human understanding, not the completion of it. And the future of mathematical research cannot consist only of understanding results produced by AI labs. Mathematicians have to be able to ask their own questions. That is Tao’s soil, stated by the committee OpenAI says it consulted. It is not a demand that the model be switched off. It is a demand that the field’s agenda not be identical to a lab’s benchmark.
Power, in the association’s sentence, is still the right word for the file tree, once you detach it from the call to quit. A lab with an internal model can set a year’s reading list in an afternoon. It can choose which famous names appear in the announcement. It can leave the applicability, the exposition, and the corrections to whoever feels responsible for the subject. OpenAI’s email to New Scientist, from Lindsay McCallum Rémy, says the company wants to give the community time and space to assess the work, and that it is still looking for a home for the release that meets the committee’s guidelines. Time and space are what the format spent. Seven hundred files do not become easier to assess because the sender wishes the recipients patience.
Grade the hit, or refuse the pile
A five percent hit on problems that whole communities have worked for years would be a serious scientific event if each hit arrived in a form a specialist could read. Prior results would say why they apply. A strange construction would be introduced as a construction, with a reason to believe it, not as a recursive object the reader is dared to reconstruct. Failures would be visible, so the five percent could be interpreted. Some of that may still be true of particular files in this repository. Moshkovitz thinks the Unique Games claim may really be a proof, and she thinks the conjecture was true. Aaronson thinks several of the results, if understood, would have been the result of the year in their fields. The issue is not whether a machine is allowed to be right.
The issue is the delivery. OpenAI ran an internal model across thousands of problems, kept a thin layer of apparent successes, and dropped hundreds of drafts on the people most able to say whether the drafts make sense. Some of those drafts were already wrong by a sign. Many are not yet formalized. At least one, on a conjecture a leading theorist has lived with, cites a library of earlier theorems and does not stop to say which of them the new construction is allowed to use. The company has said the exposition will be better next time. The field is being asked to supply this time’s exposition for free, and to do it with a second model, because the first model’s prose does not hold together.
Tao is right that the soil matters, and that AI can be part of tending it. Moshkovitz is right that there is something to learn from a proof that does not look like ours, and right that a proof you cannot read is not yet a proof you have. The counterattack worth keeping is the narrower one. Do not mail the pile and call the recipients your collaborators. If the hit rate is five percent, publish the hits as mathematics: one argument, one author who will defend the citations, one account of the construction, and the list of attempts that missed. Until then the 722 files are not a literature. They are a grading queue.
References
- Sharing AI progress in mathematics. OpenAI. October 6, 2026.
- The Mathocalypse. Scott Aaronson, Shtetl-Optimized. October 7, 2026, updated October 9, 2026.
- Math 1.0 and Math 2.0. Terence Tao, Mathstodon. October 6, 2026.
- AHM statement on OpenAI's October 6 release of mathematical documents. Association for Human Mathematics. October 7, 2026.
- OpenAI announces 722 mathematical discoveries in one go. New Scientist. October 7, 2026.
- Mathematicians will need years to make sense of OpenAI's latest drop. The Verge. October 9, 2026.