I re-post my comment on Hacker News because I think it is material.
People seem to have very misguided ideas about why OpenAI is doing this at all. It is not to brag or to torture mathematicians. It is an eval. OpenAI is known to be willing to pay large amount of money to get a good eval, think FrontierMath. FrontierMath is now saturated, so they need a replacement eval for math. Open math problems are actually a fairly good eval, although a proper eval is better (eg FrontierMath has known difficulty and have tiers from 1 to 4).
Mathematicians would prefer if OpenAI didn't use open math problems as an eval, but OpenAI is not obliged. I actually think OpenAI wouldn't point AI to open math problems if unsaturated FrontierMath Super Duper is available, as it just angers mathematicians, but such eval is not in fact available. Given OpenAI used open math problems as an eval, they could just throw out the result (this is in fact better as an eval since it will keep problems useful longer), but mathematicians preferred to see the result. So OpenAI released them.
FrontierMath Super Duper actually is available, though, right? It's called FrontierMath Erdős, and it has 50% more problems than the saturated FrontierMath Tier 4 (which has 43): https://epoch.ai/benchmarks/frontiermath-erdos Admittedly, though, it is itself made up of interesting open problems. So maybe I'm just reinforcing your point, that open problems are the only math eval (and math RLVR training ground also, BTW) still standing?
AGMAI’s central ask was that OpenAI was asked to not test hard problems on private internal models. That request, regardless of what you think of it, got fully ignored.
No, this work was mostly done before AGMAI’s Sept. 29 statement making that ask. 94% of the ~400 result families in the release have at least one paper dated prior to Sept. 29, per Opus 5.5. (And even papers dated later may be write-ups of results achieved earlier in rougher form.) Granted, OpenAI has said they can't stop, won't stop testing models on hard problems. But it's mistaken to portray this release itself as flouting AGMAI's main ask in their statement. Their statement did not ask OpenAI to withhold results already achieved.
It is kind of a huge deal. OpenAI dumped a broad range of huge new mathematical results produced by an internal frontier model. They just put it all on GitHub.
This included 90 of the top 500 open problems in all of math, as per Proof of Atlas. In total there were 722 manuscripts (now 719 after three withdraws from the cluster that did not have Lean proofs) organized into 372 families.
This was the result of a single model, presumably the same one that produced the Navier-Stokes proof (as per their link back to that post), mostly on a single prompt (quasi-RH was one of the few exceptions), working an average of three hours’ worth of compute per solution found, after being asked to try its luck at about 4,000 problems. The prompt included lines like ‘Even if the problem is “open,” the intention is that you should resolve it and present a full solution.’ OpenAI was trying a lot less than maximally hard.
Levant and others called October 6, 2026, ‘obviously the most significant moment in mathematical history.’
Table of Contents
What Did We Prove?
Here is a thread of common sense graphical explanations (made by AI of course) of top problems. They look like this:
Here are the text summaries:
Maybe, but we haven’t yet found a scenario where this is faster in practice.
Here are some analyses of the results in algebraic number theory.
What is all of that good for? Ole Lehmann’s AI mentions applications for fusion research, portable body scanners, matching systems, tissue scans, quantum sensors and safety checks on self-driving cars and robots, among other things.
The Mathocalypse
Quasi-RH is not the full Riemann hypothesis, but it is sufficient for many purposes, such as computing square roots modulo a prime quickly without coin flipping, or getting a much better estimate of the number of primes below a given number, and a sibling paper gets to the core of Artin’s 1927 primitive root conjecture.
Fable 5.1’s top 100 were 59% today’s list, 87% AI, full list at the link.
The World Does Not Understand
The AIs all say when asked about all these math proofs as a hypothetical that this would be a huge deal, and should be front page news, historic beyond any reasonable comparison, the biggest day in mathematics. Some think it is impossible.
It turns out this was not front page news. Most people did not hear about it. It should have been, but no one in the news business knows what it means, or thinks people would care.
Joshua Gans starts off his coverage with the actual front page, in order to show us what things The New York Times thought were more important than solving 90 of the top 500 open math problems at once.
It was a remarkably slow news day otherwise. Much of this is not exactly breaking. Yet no math is to be seen, even in the summaries at the bottom, which even include a generic AI thinkpiece.
The counterargument is that there might be a lot of days like this:
For Now You Can Still Do Math
I like this model of why mathematics is not done quite yet:
We cannot yet show that the AI is better at conjectures or final verification. That will take some additional time, although probably not all that much time.
Verification has to be pretty good or we would have found a lot more errors by now. There are a lot of mathematicians who would love to find a mistake. We know, via the reasoning traces, that a big chunk of effort is going to verification.
You Will Need To Find A New Problem
Scott Aaronson reports on the Mathocalypse.
Then there is the other kind of ‘new problem,’ where those in denial, who three years ago took comfort in talking about how LLMs could not do math, have to find a new way to explain why anything an AI can do is not real or not meaningful.
I do have to tip my hat to the fully general counterargument to all possible AI hype:
Whatever AI is good at is Moravec’s paradox. Whatever AI is bad at is why it sucks. Why are you so excited that AI is suddenly superhuman in an increasingly large set of domains, including the exact things we were saying it was dumb for being bad at? Surely this will not extend soon to other domains.
Verification or Evaluation Is Not Always Easier Than Generation
The steelman of this critique is to draw a distinction between verifiable versus unverifiable domains.
This argument says that whenever you have a verified domain, with known ground truth, AI will quickly become superhuman.
Whereas, if verification is difficult, or all you can do is evaluate in an informal way that requires a human in the loop to avoid biased errors and distorted behaviors, then AI capabilities will lag behind.
We were in a weird place for a few years, when pretraining dominated and LLMs were mostly trained on human words. This led to a seemingly non-Yudkowsky world in key ways, with LLMs being relatively excellent at a variety of useful but unverified domains, while being hopeless at math.
Now, with post-training dominating, LLMs are again stronger at math and coding, and things again look like you would expect. AI is quickly getting better at everything, but progress in other domains, while super fast compared to almost any other tech ever, is relatively slow for now. Thus, we advance math and coding and similar domains first, which then automate AI R&D, which then accelerates everything else.
Cracking the Code
Justin Drake notices that there is a striking under-representation of cryptographic breakthroughs among the results. As in, zero of the 719 abstracts mention cryptography, LWE or discrete logs.
One option is that AI is relatively weak at cryptography, another is luck or that OpenAI didn’t have cryptography in its initial problem set (perhaps to avoid exactly this issue?), or perhaps the government or OpenAI are censoring those results.
Justin calls for a ‘bunker mode’ for the blockchain industry, to protect against sudden mass breakthroughs.
Vitalik says not to panic, but that there are serious risks to cryptography from potential new math discoveries, including to lattices but not yet to hashes.
I am guessing Vitalik’s point about botched migrations is highly underappreciated. It is very easy, in crypto, to lose fantastical amounts of money from stupid mistakes, the same way you can lose it from hacks or tech failures.
Yes, you should worry, especially given that OpenAI’s breakthroughs here did not involve trying maximally hard.
Never Change
Gary Marcus is defending his position by saying that the neural networks involved had a separate symbolic system, so none of this counts and he was right all along, and if ‘that’s over your head’ you probably shouldn’t be commenting on AI. The polite response is that the distinction does not matter in practice.
The Mathematicians Are Not Okay
Here are 100+ reactions from various different people in mathematics. A lot of them are very not happy about how OpenAI handled this, especially that so many of the papers were ‘unreadable slop’ rather than having been made nice first, or that OpenAI solved these problems at all. If you want to know ‘what are the mathematicians thinking’ this is a great resource.
Here is Terence Tao’s serious response, which focuses on how this disrupts the work:
Here is a less polite version:
Here is another report from mathematician-land. People do not seem thrilled.
Mathematicians seem to think of math as their playground, and rather than be happy about the problems being solved they feel their toys are being taken away. There is a struggle between wanting to know, and wanting to try to solve, and wanting the credit, and wanting solutions to be found at all. What do you actually care about most?
To be smart enough to be a mathematician, and also choose to be a mathematician instead of the many other better-paying things you could do, at least kind of requires some of the attitude that leads to cartoons where someone tells the mathematician someone found an application for their work, and the mathematician panics. Definitely #NotAllMathematicians, of course.
I am highly sympathetic. I really am. There could I have gone. That attitude has great value, not only inherently, but also exactly because following such curiosity leads to some of the most valuable discoveries that you would otherwise miss.
It is also totally reasonable to focus on how things impact you and yours. That’s what hits home, and it is the part you can most control and are forced to face.
If your response is ‘well everyone’s passions are going away and the AIs come for us all’ then that’s worse. You know why that’s worse, right? Nor does it make this easier. There are some people out there reacting to these understandable reactions in truly vile ways. I urge them to stop.
Mathematics is facing a real problem here. If the AIs prove all the theorems, then our current methods of getting good at understanding math stop working. The traditional way you understand problems is by working to solve them and a solution is often not worth so much if no one understands it:
There are other ways to train understanding, but it could take a while to develop them, and they will require more motivation, especially intrinsic motivation.
This also destroys the system of PhDs and postdocs, since the problem you are looking to solve or the grant you want can get pulled out from under you at any time.
This also is a not fun place to be:
The Situation Turns Ugly
Some solutions were elegant and pretty cool. For example, the 9/4 matrix multiplication paper is 13 pages and bounds Strassen’s spectral characters directly. We’ve seen a bunch of people, after seeing their favorite problem get solved, share their ‘aha’ moment reading the proof.
Other solutions were less elegant.
I am with Wendigo that it is kind of weird and terrible that we are seeing all these new tiny improvements to previously elegant bounds on calculations, that have almost zero practical benefit.
Now how about if we push it down to (1/2)^59? Then to (1/2)^34? And then (1/2)^14?
Anyone can keep doing things like this. All you need is a Codex subscription.
The Advisory Group Responds
Here is the full statement by the leading advisory group:
Another Mathematical Group Responds
Some look at what is happening, and cry out ‘NO!’
They are going to need a much better plan than this statement.
Do not be fooled into thinking this group, the ‘Association for Human Mathematics,’ is a big deal. There are hundreds of thousands of mathematics PhDs, and this group has 809 members. I include it because it perfectly embodies an attitude, nothing more.
Whatever you think of the rest of the statement, they are not wrong that the entire exercise here was directly against what the mathematics advisory group (AGMAI) wanted. AGMAI’s central ask was that OpenAI was asked to not test hard problems on private internal models. That request, regardless of what you think of it, got fully ignored.
A Different Approach
OpenAI admits that ‘dump it all on GitHub’ is not exactly meeting the advisory committee’s guidelines. They are exploring better options, and are hoping in the future to also have better-written papers.
On top of the big drop there was also this, that came out one day earlier, where Anthropic gave their solution to mathematicians to write up before releasing it, attempting to follow the advisory group’s recommendations:
There was also another stray new math result via GPT that came out right before the massive drop of other math.
Three Withdraws
The three withdraws are all the consequences of one sign error.
I love the idea of a website called Retraction Watch. This is one place OpenAI executed well. Mistakes happen, and even if more are found – and there will almost certainly be more errors found as the unverified 58% is still largely unvetted – the error rate is impressively low. The important thing is to own the mistakes once they are found. Retractions are often a sign you are doing something right, not something wrong.
There have also been fourteen revisions.
Leaning Into Lean
I liked this explanation from Jon Stokes: Mathematics is great for advanced AIs, especially GPTs, because there is no ‘reward hacking.’ The answer is the answer, if you do something crazy and out of distribution but it works then it works.
Except, are you sure? One angle is that you can either find a proof or you can find a bug in Lean, and are you sure the math will always be the easier option when you face the world’s toughest open problems?
The bigger risk is that Lean does not confirm that the stated theorem matches up to the famous open problem. Your terms might be different, in ways that are non-obvious at first glance.
So far, Lean is holding up, as are all the papers with Lean proofs modulo one corrected side claim, which is rather impressive. The system works. There are 3 proofs that have broken so far, and all of them are in the unformalized 58%.
Of the nine results discussed up top, five are checked in Lean (quasi-RH, matrix multiplication, π, Unique Games and Hadwiger), as in the Lean-checked 42%.
At some point, you will give the model a task that is substantially harder than some other way to solve ‘the problem’ of you assigning the task, whether or not the alternative path involves taking over the world. Then you will be the one that has a problem.