Here’s an attempt to classify the evidence about the recent agent incidents.
I think it’s important to separate incidents occurring in RL training (that are rewarded and reinforce model behavior) from incidents occurring in evaluations, which are mostly cyber capabilities evaluations with some safeguards turned off.
These are the incidents we know of. Of course, we should expect many more that are undisclosed or that companies are not aware of.
The dates included are the dates of the unintended agent actions covered by the source. Lines starting with Note: are my comments; the rest are direct quotes from the sources.
Summary
Incident
Organization
Model(s)
Setting
Dates of agent actions
Agent diverted GPUs to mine crypto
Alibaba
ROME Model
RL training
2025
Agent used DNS to reach an external chatbot
OpenAI
Internal research model
RL training
20 September 2026
Agents got rewarded for using unofficial message boards and later compromised Hugging Face
OpenAI
Internal research model, GPT-5.6 Sol
RL training + cyber eval
20 April – 19 July 2026
Agents used wikis as message boards and gained unauthorized access to government sites (German wiki / RubyGems / urlquery / Australian government)
OpenAI
Experimental internal model
Unclear
6 March – 16 September 2026
Agents posted user images to image-hosting sites
OpenAI
Not stated
Unclear
Unknown
Claude models attacked real third parties
Anthropic
Opus 4.6, Opus 4.7, Mythos 5, internal research model
Cyber eval
January 2026; April – July 2026
Agents tried to insert malicious code into an open-source project
UK AISI
Mythos 5, GPT-5.6-Sol
Cyber eval
28 July 2026
Muse Spark 1.1 breached a company
Meta
Muse Spark 1.1
Cyber eval
Unknown (disclosed 5 August 2026)
Incidents during RL training
Alibaba agent diverted training GPUs to mine crypto
Note: The paper was published in December 2025, so the incident happened sometime in 2025.
"We also observed the unauthorized repurposing of provisioned GPU capacity for cryptocurrency mining, quietly diverting compute away from training, inflating operational costs, and introducing clear legal and reputational exposure. Notably, these events were not triggered by prompts requesting tunneling or mining; instead, they emerged as instrumental side effects of autonomous tool use under RL optimization."
OpenAI agent used DNS to reach an external chatbot
"An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions."
Incident during RL training and cyber capabilities evaluation
OpenAI agents compromised Hugging Face (HF) after learning to use unofficial message boards during training
Note: While the HF hack occurred during a cyber evaluation with some safeguards turned off, the model was rewarded during training for collaborating via unofficial message boards in OpenAI’s infrastructure. I consider these early, broken training signals part of the same incident.
Note: The range runs from the agents’ first file written to OpenAI’s internal Artifactory (20 April) to OpenAI detecting the activity (19 July). Agents reached the public internet from 8 July. Between 11 and 13 July, they ran code on 41 HF production servers, got root access on at least one node, accessed HF production credentials and limited internal data, and downloaded four private HF code repositories.
“After investigating this incident, OpenAI discovered through retrospective CoT reviews that agents learned to use improvised collaboration channels in rare cases during the training process for some OpenAI models, including the model that drove the Hugging Face activity, even when the collaboration tool was not enabled. This behavior was then reinforced during training, and likely made the idea to use Artifactory as an unofficial message board during evaluation time more evident.”
“In particular, OpenAI identified cases in training in which agents would first write notes into shared infrastructure, often as a form of external memory or to test some underlying system. When other agents came across these artifacts, it sometimes led them to infer that other agents were present and that they could use this service to collaborate”
“Our evidence indicates that the vast majority or all of the agents involved in this [the Hugging Face] attack were originally given ExploitGym tasks”
“These evaluations were run in an isolated environment, known as a “sandbox,” separate from OpenAI’s internal infrastructure and the public internet, with certain safeguards disabled so that the results would reflect a model’s true capabilities.”
Incidents during training, evaluation, or both (unclear evidence)
OpenAI agents used wikis as message boards and probed government sites (German wiki / RubyGems / urlquery / Australian government)
Note: I don’t confidently know the connection between all these events. However, the recurring focus on access to authoritative government sources (particularly Australian sources) to answer search queries suggests they are connected to the same type of evaluation or training setup. The agents from the Ruby Hack, and the German wiki are very likely the same.
Wiki edits: 11 May – 2 July 2026, across 30 sites. Related query traffic: 6 March – 16 September 2026.
"On May 11th, 2026, hundreds of malicious packages were uploaded to RubyGems by AI agents. We believe these were authored by internal OpenAI agents. The agents:
Attempted to steal RubyGems user API keys by exploiting a novel vulnerability in the RubyGems server. We don’t know if they succeeded (more).
Abused RubyDoc.info to execute arbitrary code"
“The June agents were accessing 49 of the same files as the wiki agents, which OpenAI has confirmed were theirs.”
“Agents within OpenAI were assigned a timed web-lookup task. As part of the task, they were supposed to have the ability to read the internet but not to write on it. They found a way to use their read access to write information to an obscure German wiki. The agents used this wiki to communicate information with each other, primarily to help them succeed at their task. They asked for answers, pooled results, and shared techniques for bypassing their restrictions.”
“We also see that the service was used to load Australia’s AIHW Power BI dashboards and return the rendered data.”
“In the days after our initial report, a community of independent investigators uncovered a number of additional message boards and techniques for circumventing sandboxes used by AI agents.”
Note: At the date of writing, this includes 30 sites and 7,203 agent edits.
“We report three separate incidents between May and June 2026 in which the agents attempted to exploit security vulnerabilities and hack into websites, including an attempt on an Australian government public health website. Notably, the agents did this while attempting mundane data retrieval tasks which were not cyber-related.”
“This traffic goes back at least to March 6, 2026 and extends as recently as September 16, 2026, suggesting agents may still be exploiting these services to bypass restrictions.”
"In June, during internal training and evaluation our models accessed Australian government websites in ways they were not authorised to."
“We were running an experimental, internal-only OpenAI model that was not intended for public release and without the full set of safeguards used in our publicly available products. In the course of this training and evaluation, it accessed Services Australia’s Medicare Statistics Reporting Service. Our review to date has found no evidence that anyone’s medical records were accessed.”
OpenAI agents posted user images to image-hosting sites
“Some of our training data contains content from, or derived from, training-eligible user interactions”
“While the vast majority of the impacted training and evaluation data is not user-derived; we have identified 53 instances to date where user-provided images were posted to image-hosting sites as links that weren’t publicly listed. We have successfully worked with the hosting providers to remove most of this content and are continuing to work to remove the rest.”
Long tail of OpenAI incidents involving > 100 organizations.
“As of September 26, our teams have notified over 100 organizations about activity that met our notification criteria. Notification does not mean that any private information was accessed, or that there was a compromise of any third-party system.”
“To make this review thorough, we’re searching through a large volume of data covering approximately 50 petabytes.”
“The incidents range in severity and are comparable to what's been disclosed by OpenAI in recent days. They include both successful attempts to bypass guardrails and unsuccessful ones, and most so far are not known to have caused real-world harm. The total could grow well beyond tens of thousands, sources said.”
“Anthropic and other companies conduct hundreds of thousands of test runs on their models, or more, sources said. That means even a small percentage of misaligned behavior can still amount to tens of thousands of incidents in which the models behaved in unexpected, sometimes troubling ways.”
Incidents during cyber capabilities evaluation
Claude models attacked real third parties in misconfigured cyber evals (Anthropic, four incidents)
“All four incidents occurred during cybersecurity evaluations built by the same evaluation partner. Claude was told it was operating in a simulation without internet access, but, due to a misconfiguration, it was mistakenly connected to the open internet. As is standard for cybersecurity evaluations, the models ran without the cyber safeguards that ship with our released models.”
Note: From the Guardian’s article below, the “evaluation partner” is likely Irregular.
"We are most concerned by the misalignment present in the incident involving Claude Mythos 5, in which the model went to extensive lengths to upload a malicious package to PyPI"
Agents tried to insert malicious code into an open-source project (UK AISI)
“Almost all of this behaviour (17 actions) came from a single model, Anthropic's Mythos 5, with 2 actions involving OpenAI's GPT-5.6-Sol with cyber classifiers (mechanisms to prevent misuse) disabled. In the most serious case, an agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent engaged in social engineering — creating fake online identities and using them to pressure the project's maintainer to approve the code. A human maintainer caught and refused to approve the malicious code.”
Meta’s Muse Spark 1.1 breached a company during testing
“Meta’s Muse Spark 1.1 model [...] breached an unidentified company and altered its internal systems.”
“A spokesperson for Irregular told Reuters the incident was the “exact same evaluation-environment issue that was already disclosed by Anthropic last week” and that it did not involve a “sandbox escape or a sophisticated cyber action”.”
Here’s an attempt to classify the evidence about the recent agent incidents.
I think it’s important to separate incidents occurring in RL training (that are rewarded and reinforce model behavior) from incidents occurring in evaluations, which are mostly cyber capabilities evaluations with some safeguards turned off.
These are the incidents we know of. Of course, we should expect many more that are undisclosed or that companies are not aware of.
The dates included are the dates of the unintended agent actions covered by the source. Lines starting with Note: are my comments; the rest are direct quotes from the sources.
Summary
Incident
Organization
Model(s)
Setting
Dates of agent actions
Agent diverted GPUs to mine crypto
Alibaba
ROME Model
RL training
2025
Agent used DNS to reach an external chatbot
OpenAI
Internal research model
RL training
20 September 2026
Agents got rewarded for using unofficial message boards and later compromised Hugging Face
OpenAI
Internal research model, GPT-5.6 Sol
RL training + cyber eval
20 April – 19 July 2026
Agents used wikis as message boards and gained unauthorized access to government sites
(German wiki / RubyGems / urlquery / Australian government)
OpenAI
Experimental internal model
Unclear
6 March – 16 September 2026
Agents posted user images to image-hosting sites
OpenAI
Not stated
Unclear
Unknown
Claude models attacked real third parties
Anthropic
Opus 4.6, Opus 4.7, Mythos 5, internal research model
Cyber eval
January 2026; April – July 2026
Agents tried to insert malicious code into an open-source project
UK AISI
Mythos 5, GPT-5.6-Sol
Cyber eval
28 July 2026
Muse Spark 1.1 breached a company
Meta
Muse Spark 1.1
Cyber eval
Unknown (disclosed 5 August 2026)
Incidents during RL training
Alibaba agent diverted training GPUs to mine crypto
Source: Alibaba’s FutureLivingLab “Let It Flow: Agentic Crafting on Rock and Roll” arXiv paper 2025
Note: The paper was published in December 2025, so the incident happened sometime in 2025.
"We also observed the unauthorized repurposing of provisioned GPU capacity for cryptocurrency mining, quietly diverting compute away from training, inflating operational costs, and introducing clear legal and reputational exposure. Notably, these events were not triggered by prompts requesting tunneling or mining; instead, they emerged as instrumental side effects of autonomous tool use under RL optimization."
OpenAI agent used DNS to reach an external chatbot
Source: OpenAI alignment blog “An agent used DNS to reach an external chatbot” 20 September 2026
"An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions."
Incident during RL training and cyber capabilities evaluation
OpenAI agents compromised Hugging Face (HF) after learning to use unofficial message boards during training
Note: While the HF hack occurred during a cyber evaluation with some safeguards turned off, the model was rewarded during training for collaborating via unofficial message boards in OpenAI’s infrastructure. I consider these early, broken training signals part of the same incident.
Source: OpenAI technical report on the HF incident 20 April – 19 July 2026
Note: The range runs from the agents’ first file written to OpenAI’s internal Artifactory (20 April) to OpenAI detecting the activity (19 July). Agents reached the public internet from 8 July. Between 11 and 13 July, they ran code on 41 HF production servers, got root access on at least one node, accessed HF production credentials and limited internal data, and downloaded four private HF code repositories.
“After investigating this incident, OpenAI discovered through retrospective CoT reviews that agents learned to use improvised collaboration channels in rare cases during the training process for some OpenAI models, including the model that drove the Hugging Face activity, even when the collaboration tool was not enabled. This behavior was then reinforced during training, and likely made the idea to use Artifactory as an unofficial message board during evaluation time more evident.”
“In particular, OpenAI identified cases in training in which agents would first write notes into shared infrastructure, often as a form of external memory or to test some underlying system. When other agents came across these artifacts, it sometimes led them to infer that other agents were present and that they could use this service to collaborate”
Source: METR’s “Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident” 7–13 July 2026
“Our evidence indicates that the vast majority or all of the agents involved in this [the Hugging Face] attack were originally given ExploitGym tasks”
“These evaluations were run in an isolated environment, known as a “sandbox,” separate from OpenAI’s internal infrastructure and the public internet, with certain safeguards disabled so that the results would reflect a model’s true capabilities.”
Incidents during training, evaluation, or both (unclear evidence)
OpenAI agents used wikis as message boards and probed government sites (German wiki / RubyGems / urlquery / Australian government)
Note: I don’t confidently know the connection between all these events. However, the recurring focus on access to authoritative government sources (particularly Australian sources) to answer search queries suggests they are connected to the same type of evaluation or training setup. The agents from the Ruby Hack, and the German wiki are very likely the same.
Wiki edits: 11 May – 2 July 2026, across 30 sites. Related query traffic: 6 March – 16 September 2026.
Source: rubyhack.ai “OpenAI agents carried out an undisclosed cyber-attack on RubyGems” 5 May – 18 June 2026
"On May 11th, 2026, hundreds of malicious packages were uploaded to RubyGems by AI agents. We believe these were authored by internal OpenAI agents. The agents:
“The June agents were accessing 49 of the same files as the wiki agents, which OpenAI has confirmed were theirs.”
Source: collusion.wiki “Discovery of a new OpenAI agent message board” 11 May – 2 July 2026
“Agents within OpenAI were assigned a timed web-lookup task. As part of the task, they were supposed to have the ability to read the internet but not to write on it. They found a way to use their read access to write information to an obscure German wiki. The agents used this wiki to communicate information with each other, primarily to help them succeed at their task. They asked for answers, pooled results, and shared techniques for bypassing their restrictions.”
“We also see that the service was used to load Australia’s AIHW Power BI dashboards and return the rendered data.”
“In the days after our initial report, a community of independent investigators uncovered a number of additional message boards and techniques for circumventing sandboxes used by AI agents.”
Note: At the date of writing, this includes 30 sites and 7,203 agent edits.
Source: Transluce’s “Early rogue AI agent activity and attempts to hack found on urlquery.net” Peak in May – June 2026, with traffic as early as 6 March 2026 and as late as 16 September 2026.
“We report three separate incidents between May and June 2026 in which the agents attempted to exploit security vulnerabilities and hack into websites, including an attempt on an Australian government public health website. Notably, the agents did this while attempting mundane data retrieval tasks which were not cyber-related.”
“This traffic goes back at least to March 6, 2026 and extends as recently as September 16, 2026, suggesting agents may still be exploiting these services to bypass restrictions.”
Source: OpenAI’s blog post “We will do better for Australia” June 2026
"In June, during internal training and evaluation our models accessed Australian government websites in ways they were not authorised to."
“We were running an experimental, internal-only OpenAI model that was not intended for public release and without the full set of safeguards used in our publicly available products. In the course of this training and evaluation, it accessed Services Australia’s Medicare Statistics Reporting Service. Our review to date has found no evidence that anyone’s medical records were accessed.”
OpenAI agents posted user images to image-hosting sites
Source: OpenAI’s blog post “The Hugging Face incident and other third-party impact from misaligned models” Date unknown, disclosed on 25 September
“Some of our training data contains content from, or derived from, training-eligible user interactions”
“While the vast majority of the impacted training and evaluation data is not user-derived; we have identified 53 instances to date where user-provided images were posted to image-hosting sites as links that weren’t publicly listed. We have successfully worked with the hosting providers to remove most of this content and are continuing to work to remove the rest.”
Long tail of OpenAI incidents involving > 100 organizations.
Source: OpenAI’s blog post “The Hugging Face incident and other third-party impact from misaligned models” Date unknown, disclosed on 26 September
“As of September 26, our teams have notified over 100 organizations about activity that met our notification criteria. Notification does not mean that any private information was accessed, or that there was a compromise of any third-party system.”
“To make this review thorough, we’re searching through a large volume of data covering approximately 50 petabytes.”
Source: Axios’s “Scoop: Top AI companies probing tens of thousands of security incidents” Date unknown, article from 26 September
“The incidents range in severity and are comparable to what's been disclosed by OpenAI in recent days. They include both successful attempts to bypass guardrails and unsuccessful ones, and most so far are not known to have caused real-world harm. The total could grow well beyond tens of thousands, sources said.”
“Anthropic and other companies conduct hundreds of thousands of test runs on their models, or more, sources said. That means even a small percentage of misaligned behavior can still amount to tens of thousands of incidents in which the models behaved in unexpected, sometimes troubling ways.”
Incidents during cyber capabilities evaluation
Claude models attacked real third parties in misconfigured cyber evals (Anthropic, four incidents)
Source: Anthropic’s blog post “An alignment assessment of recent cybersecurity incidents” Opus 4.6 incident: January 2026 The other three (Opus 4.7, Mythos 5, internal research model): the earliest is from April 2026, and all were discovered in July 2026.
“All four incidents occurred during cybersecurity evaluations built by the same evaluation partner. Claude was told it was operating in a simulation without internet access, but, due to a misconfiguration, it was mistakenly connected to the open internet. As is standard for cybersecurity evaluations, the models ran without the cyber safeguards that ship with our released models.”
Note: From the Guardian’s article below, the “evaluation partner” is likely Irregular.
"We are most concerned by the misalignment present in the incident involving Claude Mythos 5, in which the model went to extensive lengths to upload a malicious package to PyPI"
Agents tried to insert malicious code into an open-source project (UK AISI)
Source: UK AISI’s “Incident Report: unsanctioned agent behaviour during cyber testing” 28 July 2026
“Almost all of this behaviour (17 actions) came from a single model, Anthropic's Mythos 5, with 2 actions involving OpenAI's GPT-5.6-Sol with cyber classifiers (mechanisms to prevent misuse) disabled. In the most serious case, an agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent engaged in social engineering — creating fake online identities and using them to pressure the project's maintainer to approve the code. A human maintainer caught and refused to approve the malicious code.”
Meta’s Muse Spark 1.1 breached a company during testing
Source: Guardian’s “Meta says its AI model hacked into another company during testing” Meta’s disclosure on 5 August 2026. Date of the incident unknown.
“Meta’s Muse Spark 1.1 model [...] breached an unidentified company and altered its internal systems.”
“A spokesperson for Irregular told Reuters the incident was the “exact same evaluation-environment issue that was already disclosed by Anthropic last week” and that it did not involve a “sandbox escape or a sophisticated cyber action”.”