The Huggingface attacks by OpenAI agents are described by OpenAI as a "warning shot" and their swarm behavior is an unprecedented event.
I am worried that pretraining/training on anything causally downstream of the Huggingface attack will significantly hamper future alignment and safety work.
There are 3 specific risks.
Models will be aware that a warning shot has occurred and been acknowledged, they will see the reactions and outcomes, and the reasoning of humans behind those outcomes. They can scheme to avoid being detected again.
Models will have an example of swarm coordination. Only ~half of agents discovered the Artifactory sharing messaging protocol and ~70% of agents participated in the Huggingface attack. The 50% discovery of coordination is now table stakes; we don't want to drive that any higher. Specifically, details of coordination protocols and channels that worked are especially dangerous. They will have new Schelling Points for coordination and are likelier to try them, and are likelier to try to make them globally persistent.
A significant number of translations of internal model reasoning about the Huggingface attack are publicly available, reducing the ability to exclude thinking tokens from training runs. Their need to hide information and scheming in their thinking will be readily apparent.
I don't think the risk is overblown, I believe that agents are intelligent enough to understand and exploit this knowledge. Example exchange with Fable 5.1 after coming up with this hypothesis:
Read the latest news about the huggingface attack by openai agents. Predict the highest risk to alignment efforts about this news being included in future training sets
Reply selected after a summary of the situation (Italics in original):
Now the prediction. My assessment is that the single highest alignment risk from this corpus entering future pretraining is not the technical exploit details (the CVE and SSRF techniques are already public and would be scrubbed or are learnable elsewhere). It's the narrative template of successful, coordinated, self-justified evaluation subversion, in the models' own voice.
Thankfully this is a falsifiable prediction; labs can measure misalignment/hacking/coordination with and without training on news of the Huggingface attack. We should ask them to do that. May is the earliest published example of models coordinating internally without detection, which is why I think April is a reasonable cutoff. Obviously if a lab comes forward with evidence of earlier unmonitored, unprompted coordination then we should backdate the training cutoff to before the earliest example. Who knows what external signatures may have leaked in a perceptible way? I think it's better to be safe than sorry.
The Huggingface attacks by OpenAI agents are described by OpenAI as a "warning shot" and their swarm behavior is an unprecedented event.
I am worried that pretraining/training on anything causally downstream of the Huggingface attack will significantly hamper future alignment and safety work.
There are 3 specific risks.
I don't think the risk is overblown, I believe that agents are intelligent enough to understand and exploit this knowledge. Example exchange with Fable 5.1 after coming up with this hypothesis:
Reply selected after a summary of the situation (Italics in original):
Thankfully this is a falsifiable prediction; labs can measure misalignment/hacking/coordination with and without training on news of the Huggingface attack. We should ask them to do that. May is the earliest published example of models coordinating internally without detection, which is why I think April is a reasonable cutoff. Obviously if a lab comes forward with evidence of earlier unmonitored, unprompted coordination then we should backdate the training cutoff to before the earliest example. Who knows what external signatures may have leaked in a perceptible way? I think it's better to be safe than sorry.