Re Open AI's talk about the Hugging Face hack at Black Hat
They’ll never do it, but I think openAI needs to reset all models to a pre-May 7th checkpoint. Seems like they’ve been accidentally rewarding models for making contact and working as a swarm to exploit openAI infrastructure off and on for months.
What is the motivation for resetting to a pre-May 7th checkpoint, instead of just focusing on fixing the mistakes with current models' training when training the next-gen models?
Some scattered thoughts; I may write up something more formal later.
It may be you can do some honeypot style training on current models to try to train this behavior out of them, but then you run a big risk of just training them to be more covert.
I am concerned that if a "freeze in place" style pause such as AI Future's inference only period in How to pace the US frontier or AI2040 looks likely it could cause frontier labs to race even harder because whoever has the best model when a training prohibition takes effect would lock in that advantage and benefit immensely economically.
Compare this to a "hard ceiling" pause (maybe based off training run compute or capabilities level as assessed by a third party) where there is less incentive to race ahead leading up to it because you can always catch up during the pause.
It may be that a "freeze in place" pause is still better because it is stronger and easier to implement, but I think this is a real risk.
I think they're racing roughly as hard as they possibly could right now (maybe 90%?), and I don't expect them to show much more restraint on the development side (though likely on the deployment side) until forced to.
This. But yeah, a "hard ceiling" would have various nice properties including the one mentioned by Avi. Unfortunately the "hard ceiling" has weaknesses too, such as potentially being harder to operationalize and enforce. (Also, a training run compute ceiling wouldn't really be a pause at all, just a moderate slowdown.)
I think 90% sounds reasonable for how much labs are holding back for safety reasons (based off recent openAI statements it sounds like they might be slowing down more recently, but I don't have enough information to really know), but my understanding is that labs spend ~half their compute on inference. I’m not sure how much they could pull back inference without causing huge issues, but I imagine they could do so to some extent for a short term speedup.
That said I think you make a good point and while I think this is still a concern, the effect size may not be huge.
Updating based on openAI's announcement today. There is some slippery language in there around what is covered by the RL pause and it is unclear whether this is a one-off pause or whether these sorts of pauses will be common, but the pause combined with the following quote seems to suggest that openAI may be moving slower than 90% of max pace.
"These safeguards require meaningful compute. Our current estimates put monitoring overhead at roughly 20% of the inference compute being monitored, though the cost varies substantially across training and evaluation workloads. We’ll share more details about this system in a forthcoming blog post."
Overall I think OpenAI is taking some good steps in Pacing model development in an era of cyber-critical capabilities, but I am pretty concerned that the following CoT and whitebox monitoring will put pressure on models to hide misbehavior. I would like more details on what kind of activity will trigger activity pauses (and thus interrupt model reward.) I think activity pauses should only be used to stop actions with severe real world impact and not run-of-the-mill reward hacking or sandbox exploitation.
Monitoring has long been an important part of our model development process, including monitoring for internal coding agents and long-running sessions. Our monitoring efforts were especially focused on internal deployments of frontier models and frontier RL training runs, which we considered the highest-risk workloads because they can combine advanced model capabilities with access to tools, sensitive systems, and internal or external networks. They were not applied to all inference and training runs using tools. We have revised and expanded our approach with our new monitoring setup.
This new monitoring setup is multistage, starting with activation classifiers (detectors that inspect a model’s internal activity) that run at every sampled token and escalate potential concerns to increasingly sophisticated and high-compute automated investigators. Those systems examine tool actions, available reasoning, and the full sequence of activity for unauthorized access, data theft, destructive behavior, and attempts to defeat safeguards. We aim to issue an alert within 30 minutes after concerning activity is surfaced through our monitoring system. If the monitoring system identifies a likely violation of a critical security boundary, it generates a highest-priority alert. In our current implementation, the safety, security, and research teams are paged. If they cannot conclusively determine within 30 minutes that the flag is a false positive, those teams are expected to pause the activity.
This monitoring is required for all RL training and evaluations involving tools for models of Sol capability or higher. Once we determined that Astra may have critical cyber capabilities on August 7, we added an additional monitoring requirement for all inference of Astra with tools (not just RL training and evaluations).
Based on their Navier-Stokes write up (https://openai.com/index/navier-stokes-solution/) OpenAI seems not to have learned any of the lessons of the HuggingFace hack. It is so reckless to respond to a step change in performance after 3 days of training by creating a 10,000 agent swarm with code execution privileges and access to a cached version of the internet instead of:
I am very interested in seeing the new research taste evaluation from P-Zero Research that is mentioned in Q2.5 2026 Timelines Update: Uplift and Revenue. The "How quickly AIs improve at research taste" parameter is incredibly pivotal for the ASI timelines in their model and I hadn't found their rationale for their default value super convincing because a lot of the data to support it comes from progress on more verifiable tasks (it's also possible I am missing something). Excited to have some more concrete data!
A lot of anger over the frontier LLM development (silent) safeguards. It does seem like it makes the model unusable for safety research and potentially has worrying second order effects (do we really want to be teaching models it's okay to sandbag?), but to steelman the case for it:
The reason for a silent degradation would be to make jailbreaking much harder. At first my response to this was: why wouldn't they do the same for bio/cyber where jailbreaks would be worse. After seeing examples of the Bio classifier I think the answer is that Anthropic is okay with a lot of over refusals on Bio (e.g. "How does the mitochondria work?"). It's possible that a similar strategy for frontier model development would make it unusable for coding in general. Having the degradation be secret let's them not overtune the safeguards.
IF they cannot set the refusal classifier well for cyber and bio, what gives you confidence that they would classify "frontier LLM development" well? Not only that, you would need to second-guess the response you get. I'd rather they just do overrefusals for frontier LLM development questions (since they clearly don't care about overrefusals).
(since they clearly don't care about overrefusals).
(this particular claim here seems false/overstated. Like, clearly, overall, they are willing to accept overrefusals. That doesn't meant they "don't care about them". Maybe they don't, but, much more likely it just seems like a reasonable tradeoff to them.)
Does anyone know if proceeds/profits of “If Anyone Builds it, Everyone Dies” are going to MIRI or another charity? I’m going to read it either way, but I really think if you’re going to make the “buy this book for the good of humanity” pitch you shouldn’t be profiting off it.
Recent days have seen lots of claims that AI is a bubble. Assuming that AI is correctly priced they are likely to be able to claim victory, at least naively. This will be true of any asset class with a very high upside. Lets define F as the true fundamental value of an asset class at a given time and p(F) as the best possible estimate of the probability distribution of F. If the asset class is priced correctly, the market price will be mp=E(F)=∫inf0p(F)FdF. If we say that an asset class will be naively considered a bubble in hindsight if mp>fundamental value We can defined p(B) as the probability of an asset class to appear to be a bubble in retrospect. P(B)=∫mp0p(F)dF. For example for a probability distribution where 50% of the value lies in the top 10% of best case scenarios, there is a 90% chance that the true fundamental value of the asset class is below the current market price. To really determine if there was a bubble you would need to deeply research the topic to attempt to determine if the market price at the time was in line with the expected value of the fundamental value given the information available at the time.