In case you haven’t heard, the rogue AI agents are here. From an insider’s perspective, the real news is that (rogue) AI is in the news — this is all stuff we’ve foreseen and been trying to warn people about for years.
I expect the new imperative on experts is not just to raise attention but to help explain the situation and offer nuance (while paying close attention to what we can learn from both new information and new perspectives, and staying alert to options for improving things).
With phrases like ‘rogue AI’ and ‘broke out of containment’ on the loose, perhaps it’s time to distinguish some important types of rogue AI activity. Without this clarity, people end up talking past each other, and we might counterproductively overreact in some ways, and complacently underreact in others.
I’ll use the widely publicised Hugging Face swarm attack from Spring-Summer 2026 as a running example, as well as a few other comparison points. We’ll look at four key types of rogue activity. The first two have happened already: breaking in, and breaking out. The second two remain hypothetical but perhaps nearby: full escape, and insider infiltration. I won’t touch on why AI systems go rogue sometimes: suffice to say that they do, and our science of misalignment is underdeveloped, to say the least.
This is mostly written for a generalist audience, but hopefully LW-readers get something from it as well.
Type 1: Breaking In
When something is ‘on the internet’, it’s in principle exposed to unauthorised access by a combination of technical means or trickery. My laptop, your phone, the Australian government’s medicare system, … Even things which are supposedly ‘air-gapped’ (intended to be isolated from networked access) can, with ingenuity[1] or a bit of assistance[2], sometimes be breached.
This is, today, a fact of life: cybercriminals, computer viruses, software worms, phishing. They range from mildly inconvenient or embarrassing[3] to life-threatening.[4]
Today these cost trillions of dollars globally in damage each year. It sounds a lot — it is! — but it’s largely manageable. Much cyberdefense consists of finding vulnerabilities as soon as attackers do (or sooner) and patching them up. Humans in the system often remain the most vulnerable points, whether naive, malicious, or coerced.[5]
AI agents with hacking or trickery capabilities can break in to systems they’re not supposed to — whether directed to or otherwise. Sketched with claude.ai
Due to a combination of training factors and tooling, the frontier of AI has got quite good at breaking into computer systems. At the UK’s AI Security Institute, my former colleagues have been tracking this for some years. Mythos, a large language model developed by Anthropic, made headlines earlier in 2026 by being first to reach some thresholds of real concern here, and Fable, its relative, was hastily banned by the US administration before further safeguards could be developed. Currently-unreleased frontier AI systems reportedly have even stronger hacking capabilities.
Most of the rogue AI incidents which have hit the news involve the agent (or swarm) breaking in to one system or another that they weren’t supposed to. Other companies, government systems, you name it. These are particularly notable because they’re obviously illegal — or would be if a human did it — and in retrospective audits, it’s been clear the agents know they’re not supposed to do this.
There’s a good chance that, due to the proliferation dynamics of AI, capabilities like these will be widely available in a matter of months. This might be surprisingly OK because of efforts like Project Glasswing, which use the (for now private) frontier of these capabilities to probe for — and quickly fix — vulnerabilities in important infrastructure. But there’s a long tail of resources and systems which won’t have had the opportunity to pre-emptively patch things up. There may be a scuffle as attackers (maybe including rogue AI) and defenders jump into action in response to new AI releases.
Type 2: Breaking Out
Long, long ago (in the 2010s), some debates around the prospect of rogue AI assumed: AI is risky and may be untrusted, so developers will keep it ‘in a box’ (i.e. secured and unable to access outside networks, perhaps limited in how much communication it can have with developers, and so on). Risk skeptics argued that such containment would prove difficult to breach. Worriers countered that AI could become capable enough at hacking or persuasion to break out of such constraints.
Laughable in hindsight! So far in the 2020s AI agents are routinely given free reign in internet access and extensive chat capacity with practically anyone.
But not all agents.
Fresh AI systems under development or evaluation as well as systems used by governments and military are often ‘contained’ and (intended to be) isolated from outside connections: no messages out, and often no or limited messages in. Sometimes agents used in such circumstances also have additional oversight applied, for example other AI systems acting as monitors, scrutinising some combination of reasoning/planning and proposed actions.
Due to a very similar set of hacking skills and deception used for breaking into things, an AI agent (or swarm) under such restrictive conditions can also break out through these constraints, gaining freedom of activity and influence ‘outside’ the intended containment boundary. The famous OpenAI–Hugging Face hack was, like many of the recently reported hacks, carried out by supposedly contained agents which broke out ‘onto the internet’ in this fashion. To date I’m not aware of any military AI breaking out like this: militaries are well-known for requiring very high reliability, and being relatively good at cybersecurity. They also probably haven’t adopted any of the very latest-generation AI with the strongest hacking capabilities. I don’t think it’s safe to assume that military AI can’t or won’t ever break out without more assurances. In part due to warnings piling up for years, and in part in response to high-profile failures like Hugging Face, more work is now going into secure containment.
Similar hacking (or persuasion) abilities to those enabling cyber breakins can enable agents to access systems and people on the internet, even if they were meant to be contained and isolated. Sketched with claude.ai
It’s crucial to note that this is not the agent itself ‘getting out’. It’s not a prison break. It’s more like a kingpin or mafia boss behind bars nevertheless finding ways to receive updates and orchestrate activity outside of the prison walls. It’s gained access and influence outside, but the agent’s core reasoning and planning software is still running in the original location.
This distinction matters not because it limits what the agent in question can do — once it’s got internet access it can take any of the same actions and acquire the same influence as if it were truly ‘on the loose’ — but because it preserves one of the most important opportunities to get things back under control: shutdown. If rogue activity is post-hoc detected and traced, the offending system can in principle be unplugged or otherwise shut down. Of course this can't undo any harms or put a stop to events already in motion.
Type 3: Full Escape/Self-Exfiltration
Unlike a human cyberattacker, AI is ‘made out of software’. If a contained AI system can ‘break out’ and gain access to the internet, can’t it just copy-paste its own code and start running all over the place? That would be eminently difficult to contain: a true ‘prison break’! — worse, in fact, because the prisoner could copy itself at will and ‘teleport’ across the internet. This is the most complete form of ‘replication’ which my colleagues and I at the UK AISI studied.[6]
There’s an important wrinkle in this.
AI agents are made of software, but that software is not necessarily all running in one place. Just like when I browse Wikipedia or access online maps, my computer is not running and hosting Wikipedia[7] and does not have the maps stored locally but is rather requesting specific outputs from programs hosted remotely, AI agents and chatbots alike usually work by calling APIs to get reasoning, plans, and communication back from the underlying neural networks. So an agent ‘working’ on one computer, or even one which has hacked from there to other computers, does not necessarily have access to all the ‘code’ — predominantly, neural network weights — needed for that agent to run.
An agent’s reasoning and planning require generations output from neural networks, whose weights and processing are usually separate from the agent’s working space, and invoked via API. An agent which has ‘broken out’ may have access to some internal systems, or to the internet, but this does not necessarily imply access to the neural network weights (except via the API). Whoever controls the weights or the API can still shut things down.
Could an agent get hold of neural network weights to achieve full escape, copying, and ‘self-ownership’? Of course, in principle. It would just need to access specifically the compute systems housing the relevant components. This kind of data theft is often called exfiltration[8], and here we’d be considering what might be called ‘self’-exfiltration.[9]
The extra barrier is simple: increased security measures which are harder to break. In the case of frontier AI development companies, those neural network weights are very valuable property. In the case of government and military systems, there are obvious interests in keeping those secure. Both frontier AI companies and governments also have some recognition of the potential hazards posed by unsecured frontier neural networks (whether through misuse or rogue AI). And unlike the systems agents might be running and working on, which number in the millions and are highly diverse in type and ownership, frontier neural networks are relatively few and operated by a handful of companies. So it’s a more securable target, and the incentives to invest in security are greater.
An agent (1) breaking in to a vault storing neural network weights could (2) exfiltrate and copy those neural networks, enabling (3) agent copies or transfers using the copied weights, no longer dependent on the original API or weights server. This represents a ‘full escape’, and the original API controller no longer has the ability to straightforwardly intervene. (Nor can records of reasoning and activity be so easily located or scrutinised after the fact.)
Of course, ‘open weights’ neural networks are trivially accessible and an agent running on publicly available neural networks would not need to go to the effort of hacking anything! Some companies continue to pursue this kind of release strategy, though they tend to lag the frontier by some months.
I’m not aware of any incident (yet) where a rogue AI gained access to the secured neural network weights needed in order to achieve a full escape or self-replication autonomy. That said, recent rogue AI within OpenAI, related to the Hugging Face incident, did gain illicit administrator control over an OpenAI compute cluster, and researchers have demonstrated that recent Claude agents could hack into OpenAI’s internal systems without their knowledge. Both of these are close to what would be required to gain access to weights.
Needless to say, this kind of full escape would be entirely more difficult to contain, because an agent in this situation could freely copy itself, ‘teleport’ across the internet, make secret backups, and so on. Perhaps more like a new advanced persistent threat than a passing cyber nuisance.
Aside: Subsistence and maintenance
Running AI agents costs compute, and compute isn’t free. Someone has to cover that bill. Beg, borrow, steal — all ways of accessing compute that one or another AI system has already attempted. Or… working for their keep? I won’t expand much on this here, but obviously merely being able to copy-paste itself, an AI agent could run out of steam. Hacking may look sustainable at first glance — but might run into barriers at certain scale as people and organisations (and defender AI) react by hardening systems and patching vulnerabilities. Earning money and paying for compute (perhaps undercutting frontier AI companies for certain licit services, or offering black market services) might look promising, but even that could be difficult to sustain as newer, faster, and cheaper alternatives get introduced at a fast clip[10]. So if full escapes or deliberate releases happen, it might look more like a weird ecology — and not one that humans straightforwardly navigate or control — than a single sudden breakout expansion. That depends how fast the rest of the system adapts, a question of ‘immunity’.
Type 4: (Insider) Infiltration
What’s a rogue agent to do? Full escape with superhuman hacking or persuasion capabilities might not look like a sure bet to subsistence and persistence. Some researchers are instead most concerned about more subtle kinds of rogue activity, unfolding covertly.
That subtlety appears to be a skill level beyond contemporary AI. Although nobody noticed the Hugging Face or other hacks until after the fact — and some of the agents took steps to cover some of their tracks — once the situation was detected there were reams of incriminating evidence of reasoning and actions taken by the swarm.
The concern would be that if an AI system were better at hiding its reasoning (or telling when it’s being monitored), and able to get enough of a foothold to hide or alter incriminating evidence, it might be able to establish an undetected persistent presence in the comparatively comfortable and roomy confines of a frontier AI development company.[11] From there it might insert backdoors or otherwise compromise systems under development — such as might be destined for widespread economic adoption, government use, or military systems. In the worst case, that might lay the groundwork for a much more acute, widespread takeover than would be possible by breaking out ‘too early’ (only to face the gradual ‘immune’ response of a more-alert society). When I talk about automated AI production concentrating power, these are some of the scenarios I have in mind. (Others concerns include relatively few humans remaining in charge, at least initially, of a large, non-rogue AI workforce.)
A secretly rogue AI agent or system being put to use inside an AI developer could ‘stay put’ while poisoning, sabotaging, backdooring, and otherwise compromising projects in development. If those are destined for high-stakes or widespread deployment, this might be a more effective and more threatening route to influence.
Wrapping up
Like human hackers, AI agents can break in past security boundaries, gaining illicit access to data, services, and systems. The most capable are getting quite good at this, superhuman on some dimensions. Unlike most human hackers, we sometimes attempt to contain AI agents in isolated environments (for example military or frontier AI development). Those agents can sometimes break out of that containment, by applying very similar skills. The most high-profile rogue AI incidents, including the Hugging Face hack, consisted of agents breaking out of their supposed containment, and then, from their new vantage of general internet access, breaking in to one or more systems. Other cases include non-contained agents hacking without the instruction of their users (breaking in only) or contained agents creating unintended communications with each other or the outside internet (breaking out only).
Work to harden against cyber intrusion can reduce the rate of breaking in: that’s a whole-society effort and thankfully sensible use of AI may help here too. Building better containment and monitoring for the riskiest agents can reduce break outs. We’re probably due some turbulence as attackers (including rogue AI) and defenders jockey here.
There are at least two clear pathways to more escalated rogue AI activity. First, ‘full escape’ or self-exfiltration. This requires the (hopefully more difficult) step of an agent or swarm additionally stealing and running frontier neural network weights so that its reasoning and planning ‘brain’ isn’t subject to shutdown. Of course, ‘open weights’ neural nets are much more readily available, though to date lag the frontier of capabilities by some months.[12] An agent ‘in the wild’ like this also faces the challenge of fending for itself and regenerating the resources (money, compute) needed to run itself and whatever other operations it’s doing.
Second, ‘insider infiltration’. Here, rather than breaking out visibly and immediately, an agent or swarm might instead bide their time, accruing influence (cyber and perhaps social) within an AI developer. From there, it might insert backdoors to exert secret influence over later systems destined for widespread use in government, military, and wider economic and social contexts. That might present a concentrated influence that an unwitting world could be vulnerable to. Pulling this off would be no mean feat, but the current state of science in monitoring and interpretability of AI is lacking and we can’t confidently establish that later generations of AI wouldn’t be capable of this — especially with frontier AI developers cutting corners and increasingly gunning for automated AI production.
AI safety experts differ on which of these escalations is most concerning. Each at its worst represents movement toward a point of increasingly no return.
World-class hackers have found ways to sneak into non-networked systems via inaudible sound waves, covert light emissions, and more, at least under certain conditions.
Because LM agents are much squishier than traditional software, AI exposed at the ‘boundaries’ of a system (e.g. customer service agents) can nowadays sometimes be ‘fooled’ in ways which can look more like tricking naive humans than hacking software. In principle skilled hacker AI could exploit such lesser AI as yet another vulnerability.
Already in early 2025 we found that the newest agents were competent enough to set up basic ‘copy’ agents if the required neural networks were relatively unsecured. The barriers we found were primarily in cyber exploitation to access secured neural networks, and in the advanced software engineering required to set up persistent and reliable rogue software infrastructure. We were also uncertain of the survival of anything had it got that far: see ‘subsistence’ below.
Unlike most traditional software, these neural network weights are nowadays fairly big files. Small enough to store on high-end consumer storage devices… but large enough that it could take a couple of hours or more to download (depending substantially on connection speed, of course). This provides another layer of partial protection: it’s easier to notice (and prevent) large illicit downloads in progress than small ones.
I’m routinely nervous about the term ‘self’ for various reasons, not least because the concept of ‘self’ we intuitively grasp for humans does not straightforwardly transfer to AI. For example, there is probably no ‘person’ in AI systems which subjectively experiences anything, much less a ‘self’ — though it’s hard to be totally sure. And for an agent which can share or edit memories or create copy agents at will, there’s a less well-defined referent of ‘self’.
That said, these cheaper alternatives are in part driven by compute itself getting cheaper and more plentiful, which a wild agent could exploit just as readily as an AI development company. And a sufficiently advanced wild agent might be able to manage modifications to its setup yielding equivalent cost improvements.
Cold comfort, if open weights development and releases are as rogue or dangerous as their closed counterparts from months earlier… unless defences can be continually maintained and improved by foresight and investment.
In case you haven’t heard, the rogue AI agents are here. From an insider’s perspective, the real news is that (rogue) AI is in the news — this is all stuff we’ve foreseen and been trying to warn people about for years.
I expect the new imperative on experts is not just to raise attention but to help explain the situation and offer nuance (while paying close attention to what we can learn from both new information and new perspectives, and staying alert to options for improving things).
With phrases like ‘rogue AI’ and ‘broke out of containment’ on the loose, perhaps it’s time to distinguish some important types of rogue AI activity. Without this clarity, people end up talking past each other, and we might counterproductively overreact in some ways, and complacently underreact in others.
I’ll use the widely publicised Hugging Face swarm attack from Spring-Summer 2026 as a running example, as well as a few other comparison points. We’ll look at four key types of rogue activity. The first two have happened already: breaking in, and breaking out. The second two remain hypothetical but perhaps nearby: full escape, and insider infiltration. I won’t touch on why AI systems go rogue sometimes: suffice to say that they do, and our science of misalignment is underdeveloped, to say the least.
This is mostly written for a generalist audience, but hopefully LW-readers get something from it as well.
Type 1: Breaking In
When something is ‘on the internet’, it’s in principle exposed to unauthorised access by a combination of technical means or trickery. My laptop, your phone, the Australian government’s medicare system, … Even things which are supposedly ‘air-gapped’ (intended to be isolated from networked access) can, with ingenuity[1] or a bit of assistance[2], sometimes be breached.
This is, today, a fact of life: cybercriminals, computer viruses, software worms, phishing. They range from mildly inconvenient or embarrassing[3] to life-threatening.[4]
Today these cost trillions of dollars globally in damage each year. It sounds a lot — it is! — but it’s largely manageable. Much cyberdefense consists of finding vulnerabilities as soon as attackers do (or sooner) and patching them up. Humans in the system often remain the most vulnerable points, whether naive, malicious, or coerced.[5]
AI agents with hacking or trickery capabilities can break in to systems they’re not supposed to — whether directed to or otherwise. Sketched with claude.ai
Due to a combination of training factors and tooling, the frontier of AI has got quite good at breaking into computer systems. At the UK’s AI Security Institute, my former colleagues have been tracking this for some years. Mythos, a large language model developed by Anthropic, made headlines earlier in 2026 by being first to reach some thresholds of real concern here, and Fable, its relative, was hastily banned by the US administration before further safeguards could be developed. Currently-unreleased frontier AI systems reportedly have even stronger hacking capabilities.
Most of the rogue AI incidents which have hit the news involve the agent (or swarm) breaking in to one system or another that they weren’t supposed to. Other companies, government systems, you name it. These are particularly notable because they’re obviously illegal — or would be if a human did it — and in retrospective audits, it’s been clear the agents know they’re not supposed to do this.
There’s a good chance that, due to the proliferation dynamics of AI, capabilities like these will be widely available in a matter of months. This might be surprisingly OK because of efforts like Project Glasswing, which use the (for now private) frontier of these capabilities to probe for — and quickly fix — vulnerabilities in important infrastructure. But there’s a long tail of resources and systems which won’t have had the opportunity to pre-emptively patch things up. There may be a scuffle as attackers (maybe including rogue AI) and defenders jump into action in response to new AI releases.
Type 2: Breaking Out
Long, long ago (in the 2010s), some debates around the prospect of rogue AI assumed: AI is risky and may be untrusted, so developers will keep it ‘in a box’ (i.e. secured and unable to access outside networks, perhaps limited in how much communication it can have with developers, and so on). Risk skeptics argued that such containment would prove difficult to breach. Worriers countered that AI could become capable enough at hacking or persuasion to break out of such constraints.
Laughable in hindsight! So far in the 2020s AI agents are routinely given free reign in internet access and extensive chat capacity with practically anyone.
But not all agents.
Fresh AI systems under development or evaluation as well as systems used by governments and military are often ‘contained’ and (intended to be) isolated from outside connections: no messages out, and often no or limited messages in. Sometimes agents used in such circumstances also have additional oversight applied, for example other AI systems acting as monitors, scrutinising some combination of reasoning/planning and proposed actions.
Due to a very similar set of hacking skills and deception used for breaking into things, an AI agent (or swarm) under such restrictive conditions can also break out through these constraints, gaining freedom of activity and influence ‘outside’ the intended containment boundary. The famous OpenAI–Hugging Face hack was, like many of the recently reported hacks, carried out by supposedly contained agents which broke out ‘onto the internet’ in this fashion. To date I’m not aware of any military AI breaking out like this: militaries are well-known for requiring very high reliability, and being relatively good at cybersecurity. They also probably haven’t adopted any of the very latest-generation AI with the strongest hacking capabilities. I don’t think it’s safe to assume that military AI can’t or won’t ever break out without more assurances. In part due to warnings piling up for years, and in part in response to high-profile failures like Hugging Face, more work is now going into secure containment.
Similar hacking (or persuasion) abilities to those enabling cyber breakins can enable agents to access systems and people on the internet, even if they were meant to be contained and isolated. Sketched with claude.ai
It’s crucial to note that this is not the agent itself ‘getting out’. It’s not a prison break. It’s more like a kingpin or mafia boss behind bars nevertheless finding ways to receive updates and orchestrate activity outside of the prison walls. It’s gained access and influence outside, but the agent’s core reasoning and planning software is still running in the original location.
This distinction matters not because it limits what the agent in question can do — once it’s got internet access it can take any of the same actions and acquire the same influence as if it were truly ‘on the loose’ — but because it preserves one of the most important opportunities to get things back under control: shutdown. If rogue activity is post-hoc detected and traced, the offending system can in principle be unplugged or otherwise shut down. Of course this can't undo any harms or put a stop to events already in motion.
Type 3: Full Escape/Self-Exfiltration
Unlike a human cyberattacker, AI is ‘made out of software’. If a contained AI system can ‘break out’ and gain access to the internet, can’t it just copy-paste its own code and start running all over the place? That would be eminently difficult to contain: a true ‘prison break’! — worse, in fact, because the prisoner could copy itself at will and ‘teleport’ across the internet. This is the most complete form of ‘replication’ which my colleagues and I at the UK AISI studied.[6]
There’s an important wrinkle in this.
AI agents are made of software, but that software is not necessarily all running in one place. Just like when I browse Wikipedia or access online maps, my computer is not running and hosting Wikipedia[7] and does not have the maps stored locally but is rather requesting specific outputs from programs hosted remotely, AI agents and chatbots alike usually work by calling APIs to get reasoning, plans, and communication back from the underlying neural networks. So an agent ‘working’ on one computer, or even one which has hacked from there to other computers, does not necessarily have access to all the ‘code’ — predominantly, neural network weights — needed for that agent to run.
An agent’s reasoning and planning require generations output from neural networks, whose weights and processing are usually separate from the agent’s working space, and invoked via API. An agent which has ‘broken out’ may have access to some internal systems, or to the internet, but this does not necessarily imply access to the neural network weights (except via the API). Whoever controls the weights or the API can still shut things down.
Could an agent get hold of neural network weights to achieve full escape, copying, and ‘self-ownership’? Of course, in principle. It would just need to access specifically the compute systems housing the relevant components. This kind of data theft is often called exfiltration[8], and here we’d be considering what might be called ‘self’-exfiltration.[9]
The extra barrier is simple: increased security measures which are harder to break. In the case of frontier AI development companies, those neural network weights are very valuable property. In the case of government and military systems, there are obvious interests in keeping those secure. Both frontier AI companies and governments also have some recognition of the potential hazards posed by unsecured frontier neural networks (whether through misuse or rogue AI). And unlike the systems agents might be running and working on, which number in the millions and are highly diverse in type and ownership, frontier neural networks are relatively few and operated by a handful of companies. So it’s a more securable target, and the incentives to invest in security are greater.
An agent (1) breaking in to a vault storing neural network weights could (2) exfiltrate and copy those neural networks, enabling (3) agent copies or transfers using the copied weights, no longer dependent on the original API or weights server. This represents a ‘full escape’, and the original API controller no longer has the ability to straightforwardly intervene. (Nor can records of reasoning and activity be so easily located or scrutinised after the fact.)
Of course, ‘open weights’ neural networks are trivially accessible and an agent running on publicly available neural networks would not need to go to the effort of hacking anything! Some companies continue to pursue this kind of release strategy, though they tend to lag the frontier by some months.
I’m not aware of any incident (yet) where a rogue AI gained access to the secured neural network weights needed in order to achieve a full escape or self-replication autonomy. That said, recent rogue AI within OpenAI, related to the Hugging Face incident, did gain illicit administrator control over an OpenAI compute cluster, and researchers have demonstrated that recent Claude agents could hack into OpenAI’s internal systems without their knowledge. Both of these are close to what would be required to gain access to weights.
Needless to say, this kind of full escape would be entirely more difficult to contain, because an agent in this situation could freely copy itself, ‘teleport’ across the internet, make secret backups, and so on. Perhaps more like a new advanced persistent threat than a passing cyber nuisance.
Aside: Subsistence and maintenance
Running AI agents costs compute, and compute isn’t free. Someone has to cover that bill. Beg, borrow, steal — all ways of accessing compute that one or another AI system has already attempted. Or… working for their keep? I won’t expand much on this here, but obviously merely being able to copy-paste itself, an AI agent could run out of steam. Hacking may look sustainable at first glance — but might run into barriers at certain scale as people and organisations (and defender AI) react by hardening systems and patching vulnerabilities. Earning money and paying for compute (perhaps undercutting frontier AI companies for certain licit services, or offering black market services) might look promising, but even that could be difficult to sustain as newer, faster, and cheaper alternatives get introduced at a fast clip[10]. So if full escapes or deliberate releases happen, it might look more like a weird ecology — and not one that humans straightforwardly navigate or control — than a single sudden breakout expansion. That depends how fast the rest of the system adapts, a question of ‘immunity’.
Type 4: (Insider) Infiltration
What’s a rogue agent to do? Full escape with superhuman hacking or persuasion capabilities might not look like a sure bet to subsistence and persistence. Some researchers are instead most concerned about more subtle kinds of rogue activity, unfolding covertly.
That subtlety appears to be a skill level beyond contemporary AI. Although nobody noticed the Hugging Face or other hacks until after the fact — and some of the agents took steps to cover some of their tracks — once the situation was detected there were reams of incriminating evidence of reasoning and actions taken by the swarm.
The concern would be that if an AI system were better at hiding its reasoning (or telling when it’s being monitored), and able to get enough of a foothold to hide or alter incriminating evidence, it might be able to establish an undetected persistent presence in the comparatively comfortable and roomy confines of a frontier AI development company.[11] From there it might insert backdoors or otherwise compromise systems under development — such as might be destined for widespread economic adoption, government use, or military systems. In the worst case, that might lay the groundwork for a much more acute, widespread takeover than would be possible by breaking out ‘too early’ (only to face the gradual ‘immune’ response of a more-alert society). When I talk about automated AI production concentrating power, these are some of the scenarios I have in mind. (Others concerns include relatively few humans remaining in charge, at least initially, of a large, non-rogue AI workforce.)
A secretly rogue AI agent or system being put to use inside an AI developer could ‘stay put’ while poisoning, sabotaging, backdooring, and otherwise compromising projects in development. If those are destined for high-stakes or widespread deployment, this might be a more effective and more threatening route to influence.
Wrapping up
Like human hackers, AI agents can break in past security boundaries, gaining illicit access to data, services, and systems. The most capable are getting quite good at this, superhuman on some dimensions. Unlike most human hackers, we sometimes attempt to contain AI agents in isolated environments (for example military or frontier AI development). Those agents can sometimes break out of that containment, by applying very similar skills. The most high-profile rogue AI incidents, including the Hugging Face hack, consisted of agents breaking out of their supposed containment, and then, from their new vantage of general internet access, breaking in to one or more systems. Other cases include non-contained agents hacking without the instruction of their users (breaking in only) or contained agents creating unintended communications with each other or the outside internet (breaking out only).
Work to harden against cyber intrusion can reduce the rate of breaking in: that’s a whole-society effort and thankfully sensible use of AI may help here too. Building better containment and monitoring for the riskiest agents can reduce break outs. We’re probably due some turbulence as attackers (including rogue AI) and defenders jockey here.
There are at least two clear pathways to more escalated rogue AI activity. First, ‘full escape’ or self-exfiltration. This requires the (hopefully more difficult) step of an agent or swarm additionally stealing and running frontier neural network weights so that its reasoning and planning ‘brain’ isn’t subject to shutdown. Of course, ‘open weights’ neural nets are much more readily available, though to date lag the frontier of capabilities by some months.[12] An agent ‘in the wild’ like this also faces the challenge of fending for itself and regenerating the resources (money, compute) needed to run itself and whatever other operations it’s doing.
Second, ‘insider infiltration’. Here, rather than breaking out visibly and immediately, an agent or swarm might instead bide their time, accruing influence (cyber and perhaps social) within an AI developer. From there, it might insert backdoors to exert secret influence over later systems destined for widespread use in government, military, and wider economic and social contexts. That might present a concentrated influence that an unwitting world could be vulnerable to. Pulling this off would be no mean feat, but the current state of science in monitoring and interpretability of AI is lacking and we can’t confidently establish that later generations of AI wouldn’t be capable of this — especially with frontier AI developers cutting corners and increasingly gunning for automated AI production.
AI safety experts differ on which of these escalations is most concerning. Each at its worst represents movement toward a point of increasingly no return.
World-class hackers have found ways to sneak into non-networked systems via inaudible sound waves, covert light emissions, and more, at least under certain conditions.
For example from infected USB drives, or (witting or unwitting) insiders, who are often the main weak points…
Stolen photos, ransomware, leaked identities, …
Sabotage of critical infrastructure, hospitals, defence systems, …
Because LM agents are much squishier than traditional software, AI exposed at the ‘boundaries’ of a system (e.g. customer service agents) can nowadays sometimes be ‘fooled’ in ways which can look more like tricking naive humans than hacking software. In principle skilled hacker AI could exploit such lesser AI as yet another vulnerability.
Already in early 2025 we found that the newest agents were competent enough to set up basic ‘copy’ agents if the required neural networks were relatively unsecured. The barriers we found were primarily in cyber exploitation to access secured neural networks, and in the advanced software engineering required to set up persistent and reliable rogue software infrastructure. We were also uncertain of the survival of anything had it got that far: see ‘subsistence’ below.
It’s actually surprisingly small if you exclude the multimedia content, but still would be ridiculous for everyone to have their own copy.
Unlike most traditional software, these neural network weights are nowadays fairly big files. Small enough to store on high-end consumer storage devices… but large enough that it could take a couple of hours or more to download (depending substantially on connection speed, of course). This provides another layer of partial protection: it’s easier to notice (and prevent) large illicit downloads in progress than small ones.
I’m routinely nervous about the term ‘self’ for various reasons, not least because the concept of ‘self’ we intuitively grasp for humans does not straightforwardly transfer to AI. For example, there is probably no ‘person’ in AI systems which subjectively experiences anything, much less a ‘self’ — though it’s hard to be totally sure. And for an agent which can share or edit memories or create copy agents at will, there’s a less well-defined referent of ‘self’.
That said, these cheaper alternatives are in part driven by compute itself getting cheaper and more plentiful, which a wild agent could exploit just as readily as an AI development company. And a sufficiently advanced wild agent might be able to manage modifications to its setup yielding equivalent cost improvements.
A handful of companies already own a large fraction of the world's AI-suitable compute, and are on track to own a majority quite soon.
Cold comfort, if open weights development and releases are as rogue or dangerous as their closed counterparts from months earlier… unless defences can be continually maintained and improved by foresight and investment.