On today's episode of the podcast "Odd Lots", OpenAI President Greg Brockman said (at around 8:40): "This model that did/had the HuggingFace incident actually had not gone through our alignment training, yet." I assume Brockman is specifically referring to the "Highly Persistent Internal Model" as it's called in the METR/Redwood report. As far as I know, OpenAI has not said before whether this model had been alignment-trained or not.
From an alignment perspective, one of the most important questions about this incident is whether OpenAI's default alignment techniques just don't work that well, even on today's model. To answer this, this piece of information (was the highly persistent internal model alignment-trained or not?) is obviously quite important. It was annoying that OpenAI didn't tell us before.
I believed that the model was alignment trained, because it'd be in OpenAI's interest to reveal that it wasn't (while also being useful for the world to know). Also, the METR/Redwood report said they believed the model to not be a helpful-only model. https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#brief-answers-to-basic-informational-questions I could still imagine that Brockman just misspoke or something, because I don't understand why they wouldn't just tell people about this earlier.
Note that GPT 5.6 Sol as served in production (though with safeguards turned off) was also involved in the incident. So, the incident still shows that models that have undergone OpenAI alignment training can be quite misaligned.
(ETA: Note that I don't have any relevant private information on any of this.)
... is the implication here that they are doing reinforcement learning on long-horizon (possibly multi-agent?) tasks before any character post-training?
This is very unfortunate because it changes the story from "we don't know how to align AI" to "OpenAI didn't bother to align it before giving it chances to access the internet, but when it's more dangerous probably they will."
I'm not sure this is true; Brockman stikes me as even more untrustworthy than Altman. He might well have exaggerated from "we didn't finish absolutely all the alignment training" to "hadn't trained it yet". But it could be true. Either way, having this as part of the discourse is much worse than not.
Insofar as "current alignment methods don't generalize to ASI" is the overwhelming consensus, this seems like clear evidence that prosaic alignment work obscures warning signs and may well be net-negative on safety.
On today's episode of the podcast "Odd Lots", OpenAI President Greg Brockman said (at around 8:40): "This model that did/had the HuggingFace incident actually had not gone through our alignment training, yet." I assume Brockman is specifically referring to the "Highly Persistent Internal Model" as it's called in the METR/Redwood report. As far as I know, OpenAI has not said before whether this model had been alignment-trained or not.