Gary Marcus:BREAKING, Hysterical News: Half of the Astra problems can be solved [by] Fable. OpenAI didn’t even have a control group. And most of you fell for it!
OpenAI’s PR department suckered you AGAIN
[Emojis deleted, because they format horribly]
What a strange comment. Sure, Astra is a smaller improvement than previously thought (yesterday), but that's because other models are more capable, not because Astra is less. The unreleased-frontier capabilities are exactly what we thought they were based on those ten findings.
Also, the observation of Fable being able to solve half of Astra's 10 problems is consistent with the possibility that there is a different list of 10 problems that Fable can solve, such that Astra can only solve half of them.
What was surprising was how LLMs were so good at language and various other tasks without being able to jumpstart the recursive self-improvement process and automate AI R&D.
Right.
So, two questions pointing to the opposite directions are:
Are there inherent reasons why all those autonomous RSI experiments tend to saturate easily? Is the problem of non-saturating RSI luckily more difficult than we have thought, thus giving us more time?
Or: do we have a lot of "true RSI" overhang right now, and something could unpredictably trigger a transition to a non-saturating mode, basically at any moment?
Our understanding of the "landscape" here is really poor...
Given that the bottleneck now seems to be verification rather than generation for RL training, I'm surprised that we haven't seen any labs going all-in on building verification models.
Math is hard.
Math used to be strangely hard for LLMs. People used to gloat about that. Remember?
Math is getting easier. AI is getting more capable. Life comes at you fast.
Remember this meme?
Why yes. Yes it is.
We don’t know the extent to which Astra is a big jump over Fable and Sol in this realm. We do know that Astra can do math. As in real math.
There are Lean proofs. That does not mean that all ten results prove the things they assert that they prove. So far it is looking good.
At least some of these were possible before Astra, even without a harness, as both Sol and Fable have now proven the existence of nonsofic groups in their chat interfaces.
A lot of the time the barrier for a particular result is as simple as asking the right question and letting the model cook. With Astra, OpenAI asked it to take a crack at a bunch of open math problems, and it solved 10 of them. Once you know what they are, pointing other models at the same problems, even without ‘hints,’ shows there was a mathematical proofs overhang.
We do still have strong statements that Astra is a major step for scientific reasoning.
Some skepticism is always wise, but I believe that the reason they tried this was that there was a substantial jump in scientific reasoning, at least for some types of problems and probably across the board.
Soon we will all have Astra, and other models as good or better than Astra, both at math and at other things, from OpenAI, from Anthropic and soon after that from many other sources.
Approximately zero people have their head around what this means.
Table of Contents
How Impressive Are These Results?
All signs point to pretty damn impressive.
Some instances of Fable find it absurdly impressive.
Here’s another illustration of how this list looks:
Other instances are less impressed. But we should all agree: It’s a big deal.
Certainly that is a fantastic result for $2,000, but you also have to price in some of the costs of developing Astra in the first place, as well as the cost of any failed attempts.
My guess is that in the short term, if you budget more money to mathematicians to go harder, you don’t get that much of a force multiplier if they are not spending the money on compute. They’re still allowed to buy coffee to turn it into theorems, but that has rapidly diminishing returns.
The problem is that the mathematicians can hand off some classes and hire a little help, but they were already pretty motivated to work hard on proofs, and there are not that many good mathematicians, and training more of them takes many years, and the best mathematicians are a lot more talented and productive than the second tier. I am guessing the main thing you could do is lure a bunch of top level math talent to stay in or return to academia.
What you can do is point the mathematicians towards different problems. You still need to find places they have curiosity, but yes if you suspected these particular problems were solvable you could get more shots at these particular goals.
As I understand math, it would likely take years to see those results if those involved are shifting focus into new problem areas. Mathematicians usually need a while to struggle with and understand these kinds of problems before they can make progress.
Indeed, Alexander Gerko points to a different problem. There are not even enough mathematicians to process all the ‘vibe researched’ math results as it is, let alone what we will get with Astra and then models after Astra. If we get, as he predicts, 50 years of math progress in 2 years, who is even going to understand the results? What counts as worthy of a new mathematics PhD when he says over the last month he got AI to do several PhDs worth of math results?
Fable reacted to this in at least one case by calling this ‘the most important day in mathematics.’ The consensus is that Fable was taking things way too far if you are judging purely by the results, as impressive as they are.
The other reason the day is big is that this changes expectations. If we get these ten results now, what about more results soon?
So, yes, very impressive. It’s a big deal.
AI is now superhumanly capable at cyber and coding and superhuman at advanced math, the same way non-AI computers have been superhuman at basic math for a long time. This is distinct from superintelligence, which is a higher bar.
Often the move is to try and rationalize this as ‘oh okay sure it is superhuman at exactly the things it is superhuman at now, but not at other things, and also it is already superhuman so it cannot get substantially better than it already is.’
Please, do not do that.
Could We Have Called Sol or Fable?
Yes, for at least some of these questions, if we already knew where to look.
Levent Alpoge, who had Fable disprove the Jacobian Conjecture, pointed Fable at these ten problems, and in a day had solved five of them.
This is likely a repeat of what happened with Mythos and cyber.
Mythos has what I call ‘The Juice,’ the ability to find and string together exploits without anyone knowing what to be looking for.
Once you know what you are looking for, and point another model at the exact code snippet, usually Sol or Opus and often Kimi or GLM and so on can also find any particular vulnerability. They cannot string them together on their own at the same level, and they cannot go looking for anything at all in the same way. This crosses a threshold where in practice you would point Mythos at code and have it go looking for anything at all.
Astra has a similar version of The Juice with respect to this kind of defined advanced math. You can point it at a variety of major open problems, and it will crack some of them, and this fact motivated OpenAI to actually look for such advancements. Now that we’ve seen this, you can point Fable or Sol at these questions, and sometimes they solve them.
IA’s question is on point. You do this to calibrate advancements and capabilities in math, and because we are following curiosity. Levent is doing a cool public service.
Other people are pointing Fable or Sol at various unsolved problems. They are sometimes getting good results, like the disproof of the Jacobian conjecture. That is worth doing, but success is rarer.
That Fable and Sol can do these problems once asked does not change our answer to ‘how good at math is Astra?’ Rather it changes our answer of how good Fable and Sol are, which does inform the question of whether Astra is a step change a la Mythos. I don’t think we have enough information, from the outside, to answer that question yet.
I do think that it would have been more responsible of OpenAI to have had a control group, at least in terms of asking Sol, and giving it at least a similar budget. Of course I understand why they did not do that. It would be worse marketing, so why do more work to do worse marketing? Because science, because reputation, because all the good things. One would hope.
Still, I get it. OpenAI still did a super cool thing. They were still the first ones to actually Do The Thing and publish a result. That’s what counts. My name in Dnepropetrovsk is cursed, because OpenAI has published first.
Meanwhile, yeah, sometimes it really is that one guy in Dnepropetrovsk named Dan.
It still counts.
It’s Coming
Regardless of whether this extends to unverifiable domains, quite a lot of what constitutes AI R&D very much has verification available. If you speed something up, you know you sped it up. We have a lot of metrics one can maximize, not at the level of math but at a similar level to things like code and cyber.
That is the main reason this result matters. The math progress is cool. Over time I expect it to result in other cool things.
Whoever gets traction on true AI R&D self-improvement loops is going to suddenly find themselves in an overwhelmingly strong position. Knowing where you are relative to the other players becomes crucially important when things start accelerating more, and more information sharing on this would allow everyone to feel less pressure to be reckless. It stops meaningfully being a ‘race’ once the takeoff fully starts, to the extent that was ever the right metaphor.
Yo Shavit speculates that OpenAI might be in an increasingly strong position for AI R&D, due to its focus on RL, TTC and math proving translating well into AI R&D tasks. My guess is that this is not the case, and Anthropic’s specializations matter at least as much and probably moreso. OpenAI has invested a lot in RL, but recent events have shown the need to kind of teardown and rebuild quite a lot of that, and have illustrated some of the reasons why the important investments in deep alignment will be so important in the next phase.
One possible thing that happens if OpenAI proceeds too recklessly is that their AI is misaligned and all is lost and we get a Bad Ending and maybe all die. Another is that their AI is misaligned, and this becomes increasingly obvious and a barrier to using that AI to do the work that matters, and they have to keep pausing or reworking and they fall behind.
They Still Don’t See What Is The It That Is Coming
For example, Daniel Litt can see a future where AIs prove math theorems and no humans digest the proofs, but he thinks this would be because of deskilling. He does not realize this will be because there are no humans around to do the digestion.
If the world goes well otherwise, I am not worried about people not studying math. The right kind of math person loves studying math. There will be plenty of time available for such pursuits, at all levels. I agree with Fernando Borretti on the other types of cope he lists at the link: The AI will have better taste and better everything else than you do, so no you won’t still be in the core math loop. But I think that a lot of math either exists outside of social contexts, or it exists in social contexts that can survive. Math competitions are a lot like chess and a lot of math work was famously useless. See the old joke about someone suggesting a mathematician’s work found an application and they say ‘you take that back.’
Similarly, I do not worry that humans, if they have control over resource allocation, will be unwilling to pay for work on advanced math in these future abundant worlds. Advanced math is cool and proves unexpectedly useful and we will have lots of surplus. We might ‘run out of math to do’ but we won’t leave the math undone.
Whereas I very much do worry that once the AIs are out there doing superintelligent math things, they are soon also doing superintelligent everything else. And then soon the optimization pressures that dictate what happens are not human. We will not be the ones doing resource allocation. And, especially if this happens soon, likely we would not long survive afterwards.
Why indeed would you assume a human is paying the electric bill? I barely even pay my own electric bill now.
This is not to pick on Daniel Litt.
Is This AGI?
The consensus is no, this is not AGI.
To what extent is that goalposts moving, versus realizing that we were wrong about what is intelligence and what would be strong evidence of AGI?
I see a mix of both. I can see the argument that the ‘G’ is about sample efficiency and performance out of distribution, and the frontier has been more jagged than we expected. I also increasingly respect the response that this is basically hogwash, ‘AI is whatever hasn’t been done yet’ and it is becoming increasingly absurd to not admit we have what we were previously thinking about as AGI, the goalposts have moved a lot.
I do not think Astra is AGI, as per the way we currently think about AGI, but I view the other position as totally valid.
The AI Solved His Favorite Problems
This is the perspective of Henry Yuen, who spent a long time working on some of these problems, including finding results that Astra built upon. It hits hard.
That seems like the right things to think about next. What can one work on next, as a new favorite problem, without the AI solving that problem first as well? And most importantly, how are we going to keep this from getting out of control?
Elliot Glazer confirms Henry’s point 4, that the weakest part of Astra (and Sol) is their inability to know which parts of the proof are hard, when explaining it to humans.
Was This Surprising?
Directionally, no.
In terms of how fast this particular part of life came at us? Yes. Notice how Tamay thought he was a little under 50% to win the bet by 2030, got 3:1 odds, and won in 2026.
This result goes well beyond that one, in August 2026 instead of March 2030, at a much lower cost, and as recently as April this was trading around 50%.
A counterpoint is that ‘today’s new AI result’ will always be a particular thing that we did not expect AI to do this soon. There are a bunch of things AI could suddenly become able to do, or we could discover it can do. This is one of them, and whichever one we find today is always a surprise. That does not mean overall progress is surprising.
I do think this is modestly surprising in terms of pace, even adjusting for that. But yes we do have to keep that in mind.
Are People Not Impressed?
The problem with AI being very good at math proofs is that regular people, including those with political power, do not appreciate that the math proofs are impressive.
Even for me, the results here are kind of an Informed Ability. I can read proof summaries, but that tells me almost nothing about how hard was the proof to find, or how impressive or valuable was the result.
You say that now, but my prediction is that if Fable 6 started reliably posting bangers people would say that this did not count as AGI either. The goalposts, they move.
How Much Does This Change Our Predictions?
Somewhat towards faster timelines where it counts most. A lot of this was priced in before this weekend, but not all of it, and if you still have the expectations in this area from 2023 this should be a major surprise and a large update for you.
My update is that I expect things to be slightly faster in general, and for math and coding to be slightly more out in front of other things than I did before, that we are more likely to see more automation of AI R&D sooner, and that we are that much less likely to hit any meaningful walls. And that there might be more overhang than we realize. So yes, speed up those timelines, although not a ton.
Once you can 10x your own training efficiency you can then do that three more times, and then train how to manage a taco truck. Even if you do have to ‘patch each capability one at a time’ all you have to do is speed that up by orders of magnitude. At some point Trinity realizes she has to fly a helicopter and doesn’t know how, and so she has to have that program uploaded individually, but it takes two seconds so who cares. Same idea, except you also create the program and maybe it takes two hours, or two days, which changes very little.
That was always the baseline scenario. What was surprising was how LLMs were so good at language and various other tasks without being able to jumpstart the recursive self-improvement process and automate AI R&D. Perhaps nature is unhealing.
How Narrow Was This?
You can try to explain this away as a relatively narrow set of problems.
Verification is often not easier than generation. Keep an eye on that.
The results here have a number of things in common. They are well-defined, formalized problems, where you can easily do verification on a solution. Often they involve finding a concrete new object, like the Jacobian conjecture counterexample for n=3. There was a lot of surrounding theory around to pick up and use. They are problems that everyone considered important, but got relatively little attention.
They also all got solved at once for less than $2,000. So yes, that is the area that was crossed by this particular advance, with the new model queried at a (relative to problem importance) low budget in straightforward fashion. You start somewhere.
I expect it to not be long before other types of math problems start getting solved.
Coding and cyber are areas where we do not have full ASI (superintelligence) but where AI is clearly more capable than top humans at most central tasks, up to a reasonable high level of abstraction.
Those areas often have verification available, but not on the level of math. No formula can tell you whether the code is good, in various senses, only that it passes its unit tests or that you captured the flag.
If you are counting on your own area to be too illegible for something like that, I would not be confident in that.
Seeing Like an Optimizer
I also would worry a lot about a world in which anything you can formally measure you can not only manage but maximize, but that which you cannot formally measure, or where verification or evaluation requires a human in the loop, is much worse. Perhaps orders of magnitude worse.
Goodhart’s Law on steroids, because of course it would take steroids. Whoops.
A world of benchmaxxers, of KPIs that go ever higher, only to have you figure out why the KPIs are not such KPIs after all. This leads to very paperclip-style or Seeing Like a State scenarios, on every level, everywhere, all at once.
Contrary to the views of some, I believe that ‘can get arbitrarily good at optimizing for fully specified tasks’ is sufficient to shoot straight to superintelligence as part of the effect. In order to maximize one must first understand the universe.
It feels like it should be possible to jury-rig your way out of this issue, if your AI is good enough at verifiable tasks. No lab has offered me that $100 million a year contract but you’d simply [CENSORED] and then you’d…
…in which case, you will know if it succeeded when you see benchminning, as in a large improvement in usefulness and judgment that is not reflected in benchmarks, perhaps even with a regression in some benchmarks where upon investigation the technically right answer is not so useful.