This would be interesting. (There is discussion of this among mathematicians, e.g. Daniel Litt and Elliot Glazer, but I think more would be better. Terry Tao did some digestion of the Jacobian conjecture counterexample.)
Do any of the proofs have the vibe of AlphaGo's Move 37?
I think AlphaGo's Move 37 logically must have the vibe of AlphaGo's Move 37, and I think is somewhat widely agreed to not be indicative of deep general creativity? I think Navier-Stokes and other things like that are already Move 37 in the sense of "wow we are glimpsing something beyond us, in this area" and the attendant social narrative shifts. The question isn't whether something is somehow above you (obviously current AI is that in many areas!), but as you say, on a level above you.
I think it's not an ASI crux whether LLMs are significantly creative, because being creative is by itself not a path towards fast accumulation of technological culture (accessing far future technologies and theoretical disciplines; as opposed to looking ahead some fixed distance, even if a somewhat greater distance). Instead, 2026 demonstrates that scaling LLMs 10x compared to 2025 was sufficient to reach high levels of common sense (the kind of thing RL only elicits but doesn't create), which is key to applying technical skills in complicated engineering projects, and to setting priorities (including for developing new skills). That same amount of scaling also allowed LLMs to go from doing well at IMO to solving Millennium Prize problems. Since there's another 100x of scaling until 2032, it's almost certain that all engineering that humans can do is going to be automated just with the current methods, including creation of RL tasks/envs to plug LLM capability gaps as they get discovered (which enables LLMs to automatically learn the skills for all the other engineering things in short order).
This tells me that 2026 establishes a likely 2040-2045 upper bound on timing of ASI (fast architecture-rewriting strong RSI), because automation of engineering enables a robot-industrial explosion. It's also more likely that scaling to 2028-2032 LLMs (after 1-2 more leaps like the 2025-2026 one) makes them capable enough that continual learning in a strong sense (sufficient to get fast cultural accumulation started) gets within the lookahead distance accessible with the kind engineering-type capabilities we are seeing this year. But that depends more on the intrinsic difficulty of continual learning, which could remain inaccessible to LLMs if it requires more conceptual breakthroughs (than they can make in one step), in which case the whole 2022-2026 LLM boom is not informative about the timing of when those breakthroughs will occur.
That is, even if LLMs are capable of conceptual breakthroughs (within some lookahead distance of the current technological-cultural frontier), they also need to be capable of chaining them together (which is what I mean by unbounded/fast cultural accumulation). This requires learning the previous ideas in order to become capable of inventing the next ideas, with fast invent-learn-invent cycles that build on each other. Currently, this could only happen when new models get trained (even when the training is also automated by LLMs), which might be too slow to matter for the kinds of inventions where the chaining of inventions (that are only accessible in sequence, not in parallel) is the bottleneck, at least until an industrial explosion speeds up even the next-model training step a lot.
Are OpenAI's math results "creative" in an important way?
My gut says yes.
I believe as these results gain more readable expositions, it will become obvious that the models, as applied to math, are creative in a different way to humans.
I predict some people, including some mathematicians, will use these differences to claim "not creative", and I believe this will be cope.
The irrationality exponent ofis 2 (same as most other real numbers).
This and other errors, looks like none of the math symbols in the quoted comment were copied properly.
Many people expect the AI paradigm to approximately AGI complete, and expect it to smoothly scale to AI that can do innovative science. i.e. the kind of science you'd need to do to develop a plague from scratch that actually kills all humans, with very limited opportunity to experiment. Or, the kind of science you'd need to solve alignment, so you could build a more powerful successor AI that shares your values.
A periodic thread of disagreement has been @TsviBT, @Steven Byrnes and some others arguing that, yes, the AIs are getting better at verifiable tasks, but they still seem to be missing some basic sauce that lets them actually form new concepts on the fly and apply them.
The AIs are getting better at "brute force creativity", where you try every single idea ever devised by Man in parallel and maybe one of them works. But there's still a kind of creativity that a) would have way higher efficiency and b) is maybe necessary for doing some kinds of breakthrough work.
On one hand, my subjective experience has been seeing AIs be able to demonstrate increasingly interesting judgement and taste in increasingly many domains (even open-ended ones). But, I do indeed still run into AIs that get confused and "just don't get it", in a way that is pretty suggestive that there is still something significant missing.
OpenAI has lately been shipping some major results in mathematics, for problems that people had previously agreed were important and hard. But, I'd heard for previous results that the proofs turned out sort of "not actually interesting", compared to how sometimes makes you go "Holy shit that's beautiful/elegant/surprising" and conveys someone is a level above you."
Recently they shipped a lot of new proofs on important problems.
I'm not a math guy. But, seems fairly important for some people with enough context to wade through them and figure out "What sort of cognition was involved here?". Do any of the proofs have the vibe of AlphaGo's Move 37? Does it invent a new concept that does real work?
I'm guessing answering this is a nontrivial task that'll involve some... archaeology? Might be cool for a competent math theorist who doesn't have a higher priority project to organize the labor. (I'm guessing the broader math community will do this on their own, but not necessarily organize the information in a way that's optimized for figuring out "is this the kind of mind that could solve its own alignment or quickly develop radically more powerful weapons than we have available?".
I'm shamelessly copying in @DaemonicSigil's quick take about this, to give some hooks for people to start thinking through.