Then on X, I noticed that today's news included an announcement by Anthropic that one of their employees had last week managed to radically improve a bound related to the Riemann hypothesis
Apparently that result is also not an instance of "truly creative math research", much like none of OpenAI's latest ones. There appears to be some intuitive distinction between math results that are just "recombinations of known ideas" (possibly very sophisticated recombinations!), and math results that feel "truly novel", and the mathematical consensus is apparently that none of AI math results so far are of the latter type.
It's possible this is just a particularly memetically fit/convergent cope. Perhaps not, though; perhaps this intuition is pointing at something real. Notably, this is a take apparently shared by Noam Brown, who only expects LLMs to produce "entirely new mathematical machinery" "within two years".
I don't claim to understand what's going on there, or why LLMs would be limited in this manner (whatever "this manner" even means). As bearish on LLMs as I am, I don't see why RLVR wouldn't suffice to let them "solve" formal mathematics. Yet, something's up with that, I think.
My guess is that Sol wasn't being "truly creative" in your case as well.
Apparently that result is also not an instance of "truly creative math research", much like none of OpenAI's latest ones.
This I tend to agree with.
"truly novel", and the mathematical consensus is apparently that none of AI math results so far are of the latter type
I'd mildly disagree, in that the recent counterexamples (Jacobian Conjecture and Period-Index conjecture, the former purely AI, the latter very much Human-AI collaboration) are a third category that AI seems surprisingly good at: finding examples of stuff. Not novel theory, but personally I find it hard to deny there's a novel capability there. At least useful.
For context; my PhD research was partly aimed at verifying cases of BSD, and I care(d?) a fair amount about period-index.
a third category that AI seems surprisingly good at: finding examples of stuff
Right. Non-sofics would also go there, I think? My loose understanding is that those are arguably the biggest AI-derived conceptual result to date...
Could well be true, but far from my field, so I can't judge.
There appears to be some intuitive distinction between math results that are just "recombinations of known ideas" (possibly very sophisticated recombinations!), and math results that feel "truly novel", and the mathematical consensus is apparently that none of AI math results so far are of the latter type.
I think the conclusion is correct here but the model is not quite right. Most (or all) ideas originate from previous concepts and information, passed through a variable amount of reprocessing and abstraction. People understood the concepts of slope and area before the invention of calculus. Of course calculus took plenty of "creativity" to find, which I believe should be judged by the degree of transformation. It appears that LLMs are pretty good at combining known mathematical techniques in straightforward "right-out-of-the-packaging" ways e.g. a property of object X satisfies the assumptions of theorem Y, so X is an example of conjecture Z (exactly what happened with non-sofic groups btw). These can still be surprises, since the LLM has access to such an incredibly large array of techniques! As far as I can tell, however, they haven't demonstrated the ability to abstract ideas further to the extent of being considered "new perspectives".
I don't claim to understand what's going on there, or why LLMs would be limited in this manner (whatever "this manner" even means). As bearish on LLMs as I am, I don't see why RLVR wouldn't suffice to let them "solve" formal mathematics. Yet, something's up with that, I think.
RLVR for mathematics is not amenable to the same kind of exploration that you can do for board games like Chess and Go. In those cases you can basically harvest endless data from MCTS self-play, and thereby empirically discover winning patterns and new ideas that can be used at runtime. For math, the analogue would require the training process to somehow harvest all of the ideas that the model would need at inference to solve whatever research problem it might be posed with - essentially it would need to do the operative part of that original research ahead of time. This does not seems remotely plausible to me. Indeed RLVR for LLMs appears to be limited to improving sampling of known ideative pathways that appear in the pretraining. This last point was discussed really nicely by both Steven Byrnes and Beren Millidge recently, with some links to the relevant papers.
My view is that LLMs seem to be restricted in certain important aspects of math and other formal subjects by a lack of "conceptual fluidity". I think this is the same blocker that you've described quite cogently in your expectations for LLM progress in hard-to-verify research. Actually I don't think that the formal character of math makes that much of a difference in this regard! Sure, it makes it easier for the models to practice, but the most important driver of progress in math is still the development of qualitative concepts, and we don't seem to have a way of generating such concepts in training (except for massive-scale empirics in the case of board games).
This is just a continuation of the “God of the gaps” argument: “If a computer can do it, it’s not intelligence.” Overtime, the difference between what a computer can do and “true intelligence” becomes exponentially smaller, then ephemeral, then truly esoteric. I think the conflict at its core, is a theological conflict. A direct struggle with the theological doctrine that all true intelligence and true creativity comes from a divine spark in our meat brains, from God.
Funnily enough the weekend after sol came out I spent basically the entire weekend conjecture with a couple of tabs open focused on the BSD, where it would just crank away and I would just give it the occasional encouragement and tell it to keep going whenever it stopped. I had exactly the same experience where every time it would come up with some promising-sounding reformulation or new perspective on the problem and I would encourage it to go do that and it would claim to have proven some results and come back with another formulation of the problem. At first this seems super exhilarating like the AI was making real progress but I got a bit suspicious after a day of this pattern recurring and eventually, I think, that was essentially just spinning its wheels going around in circles pretty much the entire time without making any nontrivial headway. It is just very good at spinning a compelling story to you (and likely to itself) about the amazing progress it is making while in fact it is kind of lost and without a deep background it is very hard to tell what is actually going on.
That isn't to say that all the AI math is 'not creative' etc. The current level of AIs can do amazing things for sure and make novel mathematical progress but I think they are also quite bad (relative to their being superhuman at so many other things) at the metacognition needed to avoid fooling themselves and going round in circles and so an expert prodding and guiding them is very helpful and I think just cheering them on as they go it at autonomously. I wonder how much selection is going on when the labs post about solving some conjecture -- like is this 1 in 10 attempts, 1 in 1000, 1 in 1 million? obviously in some sense it doesn't matter since we only need to solve a conjecture once but still it is interesting to know roughly where we stand.
This is also my experience. For a field that I'm well versed in, they will look like they're making progress while spinning in circles. That said, if you can find the pitfalls they're falling into and wall those off, they start to make real progress. My guess is that self-knowledge of one's limitations is a deceptively difficult problem. It frequently arises only after genuine capability. We have a series of extremely eloquent Dunning-Krueger machines with the ability to fire rapid shots with arcing trajectories at a posed problem.
If you look at what happened in the HuggingFace incident https://huggingface.co/blog/agent-intrusion-technical-timeline, there's a firehose of activity with exponential decay in the number of trajectories that access the final answers. With an enduring myopia in how appropriate the goal is with respect to what humanity would want. My guess is the circling and exponential decay behavior starts to disappear as capabilities scale. The myopia is the harder problem.
From personal experience, ChatGPT still likes to spin in circles when it's unable to completely one- or two-shot the task, all the while sounding like it's making continual progress.
I spent a week trying to vibe-prove an optimization bound, and it almost always reached the same conclusion using different notation, presented as a "new useful reduction" and an almost completed proof with a minor "missing lemma". Prompting it to attempt to prove this lemma only resulted in another restatement of the problem in new notation and a demand for an (essentially) equivalent lemma.
To be clear: ChatGPT's work was neither trivial nor useless. I had it create a comprehensive pdf write-up and could then prove the result by actually steering the model and doing some work myself. I'm also like 80% sure that running an actual multi-agent workflow with a higher subscription tier would have succeeded.
multi-agent workflow with a higher subscription tier would have succeeded
you need Mythos/Fable as the research director. GPT 5.6 Sol High for literature search and review, and adversarial result reviewer and critic. Opus 5 (or Sol - it is comparatively cheap) as the agentic mathematician that operates within reviewable rounds or blocks.
Yesterday I sat down with GPT 5.6 Sol High to do some brainstorming. The topic was one of the less appreciated Millenium Problems (the Birch and Swinnerton-Dyer conjecture), and the initial prompt started life as a question on Quora.
Number theory is not a field I know, and I only expected the "discussion" to last for a few exchanges. However, ChatGPT was immediately inspired to try generalizing the BSD conjecture in a specific direction, and this led to an unusually drawn-out and self-sufficient line of "research". Almost every response concluded with a suggestion as to what the next task should be, and my input was just to cheer on what had been accomplished so far, and then endorse the suggested next direction.
What was especially striking to me, was the frequency with which each new response began with a conceptual adjustment regarding the sub-task to be performed. Evidently a vast variety of abstract objects are possible in number theory, and their differences and interrelations can be quite subtle. ChatGPT was regularly adjusting the next sub-task it had set itself, generally in the direction of greater nuance by aiming at a more sophisticated construction than it had first planned.
I was not able to judge what was going on with an expert eye, but it was a kind of interaction I had not quite had before. I have had lengthy brainstorming chat sessions with AI before, but mostly in physics, an area where I know something and could participate as an equal. Here, the AI's line of thought was all but self-sufficient - as I have mentioned, my role was largely just to say "yes, take the next step you just suggested" (I hear that vibe coding can be like this) - and, there seemed to be a lot of creative improvisation along the way, it wasn't just following well-established technical procedures of the field. (But this is where I could be wrong - perhaps it's normal, when problem-solving in advanced number theory, to have to make numerous nuanced judgment calls regarding the specific type of object that you are trying to construct.)
As the end approached today (the end being, arriving at the new generalized form of the conjecture), I felt certain that I was witnessing something that was worth posting about... Then on X, I noticed that today's news included an announcement by Anthropic that one of their employees had last week managed to radically improve a bound related to the Riemann hypothesis, and did so by urging Claude to be ambitious and believe in itself. Oh well. Perhaps I should be satisfied that for $30/month, I got to experience an echo of what the frontier labs get to do, with their billions of dollars and legions of PhDs.
My real message is that I believe most people are underestimating the significance of the recent AI successes in research-level math. This is some of the hardest thinking that humans can do, and it is now being automated. That is not a capability that will remain bottled up in the realm of pure math, affecting only mathematicians. I am much more inclined to think that this is one of the very last signs before AI becomes smarter than humans at absolutely everything.