My two cents on the Navier-Stokes drama (as someone who knows mathematicians at both labs):
It seems that Levent and Tristan treat this project as a private personal research project that they've been working on for the past year, with some LLM assistance. They are incensed that a trillion-dollar company heard a rumor about their personal work and decided to devote unprecedented massive amounts of compute at breakneck speed to scoop them in particular, and see this behavior as a gross violation of academic research norms.
It seems that Bubeck et al. at OpenAI are focused on the fact that Levent used Anthropic internal models in this private personal research project. Their story is that they have been extremely careful not to tread on math academia (my understanding is that they are sitting on a mountain of unreleased results and a substantial fraction of mathematical progress in 2026 is known only to OpenAI/Anthropic insiders). However, in their frame any project that uses Anthropic internal models is an Anthropic project (possibly legally so) and it is fair play for one trillion-dollar AI company to race literally as hard as possible against another. They are nonplussed that this race between two AI labs is being reframed as OpenAI bullying the little guy.
My impression is that the math research team at OpenAI are mostly mathematicians who genuinely care about academic norms. Several of my conversations with my friend on that team boil down to me asking him to stop worrying about doing right by normies and start worrying about alignment.
I think you are underestimating how much restraint they are already showing, as researchers who are used to sprinting to the arXiv the moment they have a proof, to sit on the gigantic hoard of results that they surely have stashed away, these results including problems many of their friends and collaborators have studied and pined for for decades.
(Pure speculation here), but I would guess that this is part of why they sprinted so hard on the NS project - finally they found an Anthropic project that they felt it was fair game for them to go all in on and take credit publicly for.
I also rather they'd care about alignment, but I don't think upholding some norm gives you some moral clearance to break it in cases where you really want to - and if that happens you were probably not really caring about the norm so much!
E.g. one alternative would have been to straightforwardly tell mathematicians that they won't hold back publishing AI results, and understand that this may upset the community. This would've still upset the community, but in a more "honest" way.
The first Millennium Prize, Navier-Stokes, has fallen to AI.
A deeply unfortunate situation has arisen involving what should have been some combination of a positive story about new progress in AI-assisted mathematical research and yet another opportunity to freak out about rapid AI progress.
Or, as we call it around here, Tuesday.
The Real Story Is The New Model That Is Better Than Astra
Keep your eyes on the prize. There are three stories here.
The first story is much more important than the second story, which in turn is much more important than the third story.
I cannot emphasize the top story enough.
I am still going to tell all three stories, but again: Eyes on the prize.
Setting the Stage
The story on this particular Tuesday begins in the morning.
Tristan Buckmaster and Levent Alpoge had worked for a year and offer us a series of remarkable results: Finite-time blowup with smooth forcing for incompressible porous media, for Boussinesq, and for 3D incompressible Euler. They also believe they have a blowup for hypo-dissipative Navier-Stokes, but the Lean verification of that result is not finished.
So far this is great. As he notes requires rethinking about how all of advanced math will function going forward. Terence Tao offers commentary on the underlying results.
The part that is not so great is where they felt forced to publish early, without the opportunity to spend the weeks necessary to make the proofs what passes among mathematicians as readable. Buckmaster outright apologizes for the way the reports look, comparing the Euler writeup in particular to AI slop.
What happened?
I Heard a Rumor
On September 1, OpenAI heard a (false) rumor that two Millennium Problems had been resolved. In response to the possibility that someone else might make the world a better place, OpenAI spent millions of dollars in inference to explore all the unsolved Millennium Problems, using an internal model stronger than Astra, and cracked the whole of Navier-Stokes in 88 hours, plus another 17 for Astra to do the Lean formalization and verification.
You can view this as ‘OpenAI got curious to see what their new baby could do’ or you can view this as ‘they spent millions trying to scoop what they thought was an Anthropic project.’ My money is on a little from column A, a little from column B.
All it took was the rumor.
Our Price Cheap
Depending on what costs you count, this would have cost a regular customer on the order of $22 million, or several million for internal marginal costs. A small price if it works, but also it is crazy how many people don’t think two steps ahead.
I mean, no, maybe $150k or if you’re lucky $15k, but the point stands.
A Millennium Prize result is worth vastly more than a million dollars. OpenAI does not intend to claim the prize money. This was never about prize money, for anyone. It was always about the credit.
An Accusation Is Made
After hearing internet rumors, Buckmaster reached out to OpenAI, and according to Buckmaster OpenAI’s Sebastien Bubeck said that an internal OpenAI model had produced a proof of finite time blowup for the forced Navier-Stokes equations over the past few days. Buckmaster claims that over the course of two calls, it became clear he had initially been misled and that an entire team had been working on the problem, using an ‘insane’ amount of compute.
Which we now know was 130 or 300 billion output tokens, depending on what counts.
If this part is true, it would be quite bad.
Or here’s Levent Alpoge telling his side of the story, and Sebastien responding:
Sebastien strongly denies that he did anything wrong, as does Noam Brown, and Sebastien offers screenshots he says demonstrate good faith.
Also a counter-accusation:
Sam Altman also strongly claims his team acted ethically, and tried to collaborate in good faith.
How Did You Get That Idea?
The implication of all this, that many drew, was that there was a soft accusation that OpenAI had stolen part of their result from Codex chat logs, or that the chat logs had been used in training data and that had contributed to the result.
We now have explicit confirmation that the Codex data was not used in any way in training runs:
I agree with Sholto Douglas that it is extremely unlikely that user data had any influence here via either method:
OpenAI Almost Certainly Did Not Misappropriate User Data
It would be an earthquake, perhaps the worst possible story for their business, if we found out OpenAI or Anthropic had raided someone’s user data to gain competitive advantage, and the facts make more sense without this. The funniest hypothesis, from Jerry Tworek, which I also find highly unlikely but which we sadly cannot rule out, is that OpenAI did not intend to use any such data, but the agentic swarm assigned to the problem hacked them and got it anyway, and OpenAI has no idea.
My strong presumption is that the user data was not intentionally accessed, that they would never (or, rather, that they at least recognize they can only do it once and this is not that once), and also that user data would have been of little help.
Thane Ruthenis points out that there was previously a school-shooting incident, where (highly reasonably) the logs were examined and there was internal debate over notifying law enforcement. Even that sent chills into some users. Where will they draw the line? And yes you should understand your data and logs might not stay private, if you give them to a lab. It is reasonable, as Andreas Thorn also is doing, to wonder about whether your proprietary research or math data has stayed private.
I still reiterate that, in practice, I would put a very, very low probability on OpenAI intentionally looking at such data. The risk-reward is just so, so bad. And I believe Roon’s later report that they have now confirmed the logs could not have reached the training data.
Our Top Labs Cannot Get Along Even On A Feel-Good Math Story
No matter who is and is not at fault, it is rather alarming that the labs cannot cooperate on something like assigning credit for a mathematical proof. This is a very bad sign and also a wake-up call.
The continued inability to get along is a big deal, and will only get bigger.
What Next?
If it worked on Navier-Stokes, what else will it work on? Time to find out.
This is indeed reported to be happening, including for fun questions like P=NP. Which, if it happened to be constructively proven true, would break a lot of things. You should also be nonzero worried about what such agent swarms might do as an incremental step, given OpenAI’s history.
The Mathematicians Are Not Happy
Solving math problems is their entire jam. Yet they are very not happy.
This is the general complaint, on top of particular complaints about lab conduct.
One could respond that no, the point of math problems is to solve math problems, and the mathematicians are mad that AI is disrupting their status games. That is certainly part of what is happening, but like Tyler Cowen I am not so cynical, and unlike him I do not think that ‘a bit of patience is needed’ or to wait until the damage is done before once can object. People can extrapolate. Did you know mathematicians are good at drawing straight lines on graphs?
There is a real complaint here. Getting people with talent to do real math for long periods of time, for very little money, and develop real math intuitions and concept mastery, when they could do other things for far more money, is no easy feat. If you ‘take away their status games’ and also far more importantly take away their fun and feelings of accomplishment and elegance and beauty, what then?
Robin Hanson suggests that then we should just fix the incentives and reward mathematicians better.
He doesn’t use the word ‘just’ but this is definitely a just, as in if only people would just [X], and as we all know humans have never justed and are not about to start now. This is not an actual thing humans might do. As even he then realized hours later:
We can and will get some rewards for those who fill in blanks and improve elegance, but it is not going to be the same and will be a Herculean effort, at best.
Scott Armstrong argues that this is a change in order, but not in result. Human mathematicians will take a few weeks to understand the proof, but then they will gain the understanding. In the past, the understanding came before the proof, but either way you get the understanding. That’s what matters. He thinks what people are mad about is that the machines are smarter than them.
As always, AI is the best tool both for learning and not learning. The mathematicians are saying that they are being forced into ‘not learning’ mode, to letting the AI do their homework, and they don’t like it, because actually the homework teaches you.
You can say who cares, that it’s fine to have AI do the frontier math. And yes, I do think a lot of people will still want to do math all day, and look for elegant versions of ugly AI things and such. But it won’t be the same.
OpenAI’s New Model Was A Step Change Above Astra Four Days Into Training
Now for the most important part of the story.
RIP various OpenAI training pauses, June 22, 2026 to August 28, 2026.
On August 18, OpenAI said they had paused frontier RL for two weeks, and their largest planned training run remained on hold. Then on August 28, OpenAI started training another model that quickly became more capable than Astra across the board (by their own reports), now that they’ve solved all those pesky supervision, infrastructure and alignment issues, and it’s good, with a ‘step change’ being observed after only four days:
Terrified? You should be. That model is still training, had been training for less than two weeks, and is not especially aimed at mathematics.
That is all we know about the model. That is enough, when we combine it with the repeated stories of OpenAI and people at OpenAI virtuously freaking out over its alignment and other problems, and over the pace of AI progress, in a preference cascade.
And we also combine it with OpenAI dropping so many pieces of alarming information, that it could have avoided dropping, because they very much need us to understand the situation. The Astra model card counts here, so does An Alien Mind, so does Paul Christiano’s statement, so do the reactions to the HuggingFace attack including Pacing the Frontier.
And so does the OpenAI report on recursive self-improvement (RSI), which is the next section.
OpenAI and Anthropic have entered the RSI era.
OpenAI Research Progress Is Accelerating Due To OpenAI Research Progress
OpenAI has been opening up lately about many things.
The most important issue of all is that of the acceleration and automation of AI R&D. This is the thing that is likely to kill us, the thing that is most scary, the thing all the lab employees are warning about, and also the thing both OpenAI and Anthropic are driving towards as quickly as possible.
They gave us a report about that, too. I agree with Tejal Patwardhan that this is excellent transparency.
Line this up with the timeline revealed in their announcement of the proof of Navier-Stokes. They started training their new model on August 28.
Meanwhile, they believe they have a research intern that the timeline suggests is a different distinct internal model. What will they have next week? Next month? What does Anthropic have?
OpenAI explains that highly capable AI offers great upside, but of course it does. That is not why everyone is so determined to go this fast. They are in a race.
This is one more call, among the many from OpenAI and its employees lately, that are a combination of ‘we need to work together to end this madness’ with a large amount of ‘somebody stop me.’
Once again I object to the view that the prosaic or current problems are the main issue, but yes, they are trying to warn us, over and over that (translating out of corporate speak a bit):
The bulk of the post is a detailed snapshot of how much agentic systems have contributed to RSI progress at OpenAI in recent months, as in since Astra came online internally. The answer is quite a lot.
If anything, the surprise is that use remained this low for so long:
They are not not bragging, and they are not not confessing. They are warning.
The binary measurement here likely makes this look less of an acceleration than it is:
I do wonder how much humans can properly handle 4+ concurrent workflows, says the man with dozens of open tabs and a dozen active post drafts.
It turns out this is not as impressive as it looks, because they’re counting subagents as workflows. That makes it more surprising that, while the median researcher spends $600 a day, there are still the 30% that aren’t doing intense usage.
Code shipments and experiments are accelerating fast starting around the point Claude Code and then Codex started being a big deal.
Until recently, most of what AI got used for was building research and infrastructure code, and the humans mostly did the rest. Now the AI does a lot of work on launching, monitoring and debugging runs, and on technical help and review, as well, and is branching out into analyzing results and other places.
They took our jobs, ‘human troubleshooters’ edition:
Here’s a different variation of the famous METR graph, with production tasks, a bit hard to read as a static screenshot but the progress is rapid. These are not years. These are months:
Quantifying OpenAI’s Pause
They have a graph of RL compute spending over time, I am surprised how little Astra has been used before July 20, but there you go. I notice they do not extend this into the present and presumably the Astra use went back up over time:
This Is the Way the World Ends
Was using the swarm to prove Navier-Stokes, only days into the training of the new model, itself risking a serious loss of control incident? June Jimenez argues yes.
If you unleash a 10,000 agent swarm of a new model you cannot possibly have tested, on a potentially impossible and definitely extremely difficult task, and collectively give it 300 billion tokens, what else might it have done? I definitely considered ‘maybe it hacked into OpenAI in some way to get the logs’ but there are so many possibilities.
Quickly, There’s No Time
When a former lab employee, here Andrew Ho, says ‘I notice that over a >3 month timescale I don’t think my productivity has increased over 100% or perhaps even over 50% because I get distracted’ as a reason to be bearish, that’s not all that reassuring even if true. Imagine only being 50% more productive every three months and thinking that this is bearish.
He is saying ‘even if AGI is eventually achievable, this implies a significantly longer timeline’ on the same day Greg Brockman said ‘welcome to the AGI era.’
On top of that, this level of math progress was not expected a few weeks ago.
Every time one of these problems falls, people try to retcon it into being not so surprising. They are often wrong. No, most people, even most forecasters, did not presume that this would happen. Perhaps you did. After all, you are very intelligent.