update:

From an alignment perspective, one of the most important questions about this incident is whether OpenAI's default alignment techniques just don't work that well, even on today's model. To answer this, this piece of information (was the highly persistent internal model alignment-trained or not?) is obviously quite important. It was annoying that OpenAI didn't tell us before.
I believed that the model was alignment trained, because it'd be in OpenAI's interest to reveal that it wasn't (while also being useful for the world to know). Also, the METR/Redwood report said they believed the model to not be a helpful-only model. https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#brief-answers-to-basic-informational-questions I could still imagine that Brockman just misspoke or something, because I don't understand why they wouldn't just tell people about this earlier.
Note that GPT 5.6 Sol as served in production (though with safeguards turned off) was also involved in the incident. So, the incident still shows that models that have undergone OpenAI alignment training can be quite misaligned.
(ETA: Note that I don't have any relevant private information on any of this.)
... is the implication here that they are doing reinforcement learning on long-horizon (possibly multi-agent?) tasks before any character post-training?

I assume that's the implication, yeah. He's relatively specific in the podcast that the takeaway for him from this incident is that they have to do alignment training earlier in the development pipeline.
FWIW, my interest here is more about the scientific question of how hard alignment of (current) models is (relative to how much effort is being put into it) and how plausible it is that we get "alignment by default" (models being aligned if we just apply simple methods like RLHF, whack-a-mole training out undesired behaviors, ...).
I think it's not so clear which one would reflect worse on OpenAI. I think current models can't do that much damage, yet, so letting them run loose without alignment and safeguards isn't that reckless. (Also, safety people might not be particularly motivated to prevent relatively harmless incidents, because warning shots are so informative for the world.) Meanwhile, future models (at this point: very-near-future models) probably can do a lot of damage, especially once they're deployed. So, not having alignment techniques that can reliably mitigate this damage is extremely reckless, given that OpenAI wants to continue scaling.
Right - it might be a worse practice on OAI's behalf, but it's probably better news for anyone wasn't sure whether current alignment techniques would work.
IIRC The huggingface incident itself was a safety eval, not a RL training run. Which is maybe bad but not quite as bad as that. Or are you referring to one of the other swarms?
It probably does make sense that when evaluating dangerous capabilities to do so before applying whatever safety training is meant to suppress those capabilities, since you want to know stuff like how dangerous a jailbreak would be.
Excellent point! There's a strong dynamic in which more safety noway mean less safety when it really counts. This is of course complex. Usually first order effects dominate, but current systems just aren't what we're worried about, so what's the first order effect is legitimately unclear.
This is very unfortunate because it changes the story from "we don't know how to align AI" to "OpenAI didn't bother to align it before giving it chances to access the internet, but when it's more dangerous probably they will."
I'm not sure this is true; Brockman stikes me as even more untrustworthy than Altman. He might well have exaggerated from "we didn't finish absolutely all the alignment training" to "hadn't trained it yet". But it could be true. Either way, having this as part of the discourse is much worse than not.
I agree that we can't be confident that their model wasn't alignment trained.
I'm not sure I understand what you're saying with the rest of your comment. Are you saying we should just assume that their model was alignment-trained because "we don't know how to align AI" is a better story or the like?
I certainly don't think we should distort the truth
No one thinks of thinks of themselves as "distorting the truth", but you did literally just say that a claim "could be true" but that "[e]ither way, having this as part of the discourse is much worse than not", suggesting that you have strong preferences about the discourse unrelated to its truth?
I mean, for those who already independently think alignment is hard, it would be unfortunate if the HF incident was done by non-alignment-trained models, because people could then spin it as "oh silly OAI just hadn't trained the model to be aligned yet (using techniques that we definitely-for-sure-know work, trust)"; the incident becomes weaker evidence that smarter models will continue to be misaligned. If you were mainly convinced alignment was hard because of the HF incident, then lack of alignment training would be reassuring. If you were mainly convinced alignment was hard prior to the HF incident, then lack of alignment training would be unfortunate because others would be less likely to come around to understanding that alignment is hard.
At 6:35 in this video:
https://www.youtube.com/watch?v=56GuvofZgB4
On today's episode of the podcast "Odd Lots", OpenAI President Greg Brockman said (at around 8:40): "This model that did/had the HuggingFace incident actually had not gone through our alignment training, yet." I assume Brockman is specifically referring to the "Highly Persistent Internal Model" as it's called in the METR/Redwood report. As far as I know, OpenAI has not said before whether this model had been alignment-trained or not.
ETA: As pointed out by Matrice Jacobine, X user roon (who works at OpenAI) says: "greg probably doesn’t have the full details, I think it’s safe to say it didn’t go through the full gauntlet of alignment posttraining. but it was alignment trained, and had reasonable looking scores on alignment evals (at the time)."