OpenAI’s pursuit of Navier-Stokes seems to have been a non-negligible loss of control risk. The risk could be ongoing.
From the timeline of their description of the pursuit, they had a model that had started training on August 28th, they decided to do a preliminary run on Euler forcing on September 1st, and seem to have launched the full 10,000 agent swarm on September 3rd or 4th.
Needless to say, this model could not possibly have undergone frontier-level safety alignment and safety evaluations, consistently reported to take weeks to months.
They say "We have been training a new internal model that has exhibited unprecedented performance in our benchmarks, including mathematics." Astra+ seems to be in the general trend of "model capabilities emerge downstream from general scaling efforts + RLVR", so it's a fairly reasonable inference that Astra+ is also more capable at cyber, bio and rogue deployments.
So it seems plausible that they took a model which was:
Substantially more capable than Astra
Had undergone much less prosaic alignment training than Astra
Had weaker safeguards than production-deployment Astra
Was directed to solve a problem a step-up in difficulty from those that had been previously solved and that was plausibly outside of its level of capabilities; the kind of environment that has seemed to bring out misaligned persistence recently.
And let it output 300 billion tokens in pursuit of said goal.
A couple of things:
Given the gigantic extent of the research effort, shouldn’t someone, or ideally a large team of people, be auditing the agent traces to make sure that Astra+ didn’t hack a bunch of stuff en route to solving Navier-Stokes? Has someone done so?
More generally, I will note that models can be misaligned during training. One hopes not. How standard is it to deploy models in very early stages of training?
And should this model be trained in the first place? Astra's training was infamously bumpy enough that Astra-class model training had to be slowed down substantially. Have you, only one to two months later, patched all of the control failure modes that would allow a more powerful system to sustain a rogue deployment?