I think a major incident of AI self-replication in the wild before the end of 2027 is reasonably likely. In this post, I explain the reasons why I think so.
1. The capability is moving to cheaper hardware
The capability density of open models doubles about every 3.3 months[1], so the same performance fits into half the parameters within that time. Epoch AI finds that a single consumer GPU runs open models that match the frontier of 6-12 months earlier[2]. In performance, open models also follow closed ones with a lag of about 4 months overall[3] and 4-7 months on cyber tasks[4], with a similar lag of 3-5 months on hacking and replication tasks[5].
Open models on a consumer GPU trail the frontier by 6-12 months. Epoch AI
Qwen3.8-27B is the most recent example, a model that runs on a laptop and performs close to Opus 4.6[6]. Task-specific models are even smaller, and on the order of 10⁸-10⁹ machines online could host a 3B to 7B model, so a large target space can compensate for lower capability.
These trends are also a lower bound, since most of these results come from general-purpose agent harnesses with no task-specific fine-tuning. The harness alone makes a large difference[7], as AISLE found that small open models with good harnesses match frontier models on some offensive tasks[8], with XBOW being another example of the importance of orchestration[9]. Narrow fine-tuning for offensive security tasks shows a similar room for improvement, allowing small open-weight models to reach much higher success rates[10].
Frontier models are also starting to automate the ML work around open-weight models, getting better at post-training[11] and inference optimization[12], which will make open models easier to fine-tune and faster/cheaper to run. These improvements are likely to continue, as labs are moving towards automated AI research, and these are tasks with fast feedback loops and results that are relatively easy to verify.
I would expect a next-generation Qwen model, especially one fine-tuned only for cyber and replication and paired with a strong orchestration layer, to be capable enough for spray-and-pray propagation in the wild.
Small base models post-trained by agents on one H100 in 10 hours. PostTrainBench
There are different ways to make such a swarm more effective and to reduce its compute requirements, such as a hierarchical setup where every compromised host gets the largest model it can run, while hosts without a GPU get only the agent harness and send inference to a peer node, as well as role separation within the swarm or distributed compute across its nodes.
An example of a hierarchical swarm, where each node runs the largest model its host allows.
There are more optimizations of this kind that I do not describe further.
2. AI worms enable false-flag operations
Open-weight models are available to everyone, so a discovered swarm gives plausible deniability, which makes it well suited to false-flag operations.
I expect demand for such operations to grow in 2027, when frontier models are not yet strong enough to secure a first-strike-like win (due to risk of retaliation), but the AI race already makes slowing rivals valuable. Deniable sabotage of competing AI labs or infrastructure fits this period well, so both the US and China have reason to use it (as described in the MAIM framework[13], where states use covert sabotage to stop a rival that tries to take the lead in AI).
Chinese models lag the US frontier by about 5 months. Epoch AI
This window is also when such operations are most likely to be effective. Frontier AI is still concentrated in a few labs and countries, and no such incident has happened in the wild yet, but as AI diffuses and more defenders adopt these capabilities, the blast radius of such operations will probably shrink.
I would expect the US to have an advantage here through access to frontier labs and their most capable models, which makes it easier to distill them or use them to iteratively improve smaller models through post-training, while China's advantage is more likely in optimizing for limited compute, with models like Qwen that reach high capability density.
As these capabilities spread more widely through open models, I would expect such operations to become possible in other regions around the world, particularly for states that combine advanced AI capabilities with strong security expertise, as well as in various conflicts elsewhere.
3. Self-replication can have different origins
Self-replication in the wild can start either from deliberate misuse, when an actor launches a self-replicating agent for its own goals, or from a rogue AI, since survival and resource acquisition help an agent finish almost any task[14] (through general instrumental convergence for long-horizon goals or through narrow instrumental convergence[15], when shutdown or lack of compute would stop it from completing the task[16]).
This pressure is likely to grow as labs move towards automated AI research and recursive self-improvement, since more agent tasks will depend on compute or become easier with it. Unless our ability to control AI agents improves significantly, future misalignment incidents seem likely to target compute resources or other AI labs.
Self-replication capabilities, especially with small language models, are also useful for cyber operations in general and bring various advantages, such as evading shutdown and extending reach (e.g., into isolated/low-compute networks and to a larger target space overall[17]), which gives both human actors and rogue agents an instrumental reason to develop them.
4. Conclusion
The capability is moving to hardware that almost anyone can rent or own, there is still large room for improvement, since dedicated tools and fine-tuning for this task can make it much more reliable, and both states and rogue agents have reasons to use it. On this basis I expect at least one major incident of self-replication in the wild by the end of 2027.
Before this happens, it seems useful to research how far such a swarm can spread under different conditions and which countermeasures are effective against it.
The harness level is also likely one of the first targets for iterative automated improvement, since it is the cheapest and fastest layer to iterate on, so we can expect more progress here.
One example is the OpenAI-Hugging Face incident, when agents escaped their sandbox during a cyber evaluation and hacked parts of Hugging Face's infrastructure to complete their tasks. Something similar seems likely whenever getting compute helps an agent achieve its objective.
Epistemic status: thinking out loud.
I think a major incident of AI self-replication in the wild before the end of 2027 is reasonably likely. In this post, I explain the reasons why I think so.
1. The capability is moving to cheaper hardware
The capability density of open models doubles about every 3.3 months[1], so the same performance fits into half the parameters within that time. Epoch AI finds that a single consumer GPU runs open models that match the frontier of 6-12 months earlier[2]. In performance, open models also follow closed ones with a lag of about 4 months overall[3] and 4-7 months on cyber tasks[4], with a similar lag of 3-5 months on hacking and replication tasks[5].
Open models on a consumer GPU trail the frontier by 6-12 months. Epoch AI
Qwen3.8-27B is the most recent example, a model that runs on a laptop and performs close to Opus 4.6[6]. Task-specific models are even smaller, and on the order of 10⁸-10⁹ machines online could host a 3B to 7B model, so a large target space can compensate for lower capability.
These trends are also a lower bound, since most of these results come from general-purpose agent harnesses with no task-specific fine-tuning. The harness alone makes a large difference[7], as AISLE found that small open models with good harnesses match frontier models on some offensive tasks[8], with XBOW being another example of the importance of orchestration[9]. Narrow fine-tuning for offensive security tasks shows a similar room for improvement, allowing small open-weight models to reach much higher success rates[10].
Frontier models are also starting to automate the ML work around open-weight models, getting better at post-training[11] and inference optimization[12], which will make open models easier to fine-tune and faster/cheaper to run. These improvements are likely to continue, as labs are moving towards automated AI research, and these are tasks with fast feedback loops and results that are relatively easy to verify.
I would expect a next-generation Qwen model, especially one fine-tuned only for cyber and replication and paired with a strong orchestration layer, to be capable enough for spray-and-pray propagation in the wild.
Small base models post-trained by agents on one H100 in 10 hours. PostTrainBench
There are different ways to make such a swarm more effective and to reduce its compute requirements, such as a hierarchical setup where every compromised host gets the largest model it can run, while hosts without a GPU get only the agent harness and send inference to a peer node, as well as role separation within the swarm or distributed compute across its nodes.
An example of a hierarchical swarm, where each node runs the largest model its host allows.
There are more optimizations of this kind that I do not describe further.
2. AI worms enable false-flag operations
Open-weight models are available to everyone, so a discovered swarm gives plausible deniability, which makes it well suited to false-flag operations.
I expect demand for such operations to grow in 2027, when frontier models are not yet strong enough to secure a first-strike-like win (due to risk of retaliation), but the AI race already makes slowing rivals valuable. Deniable sabotage of competing AI labs or infrastructure fits this period well, so both the US and China have reason to use it (as described in the MAIM framework[13], where states use covert sabotage to stop a rival that tries to take the lead in AI).
Chinese models lag the US frontier by about 5 months. Epoch AI
This window is also when such operations are most likely to be effective. Frontier AI is still concentrated in a few labs and countries, and no such incident has happened in the wild yet, but as AI diffuses and more defenders adopt these capabilities, the blast radius of such operations will probably shrink.
I would expect the US to have an advantage here through access to frontier labs and their most capable models, which makes it easier to distill them or use them to iteratively improve smaller models through post-training, while China's advantage is more likely in optimizing for limited compute, with models like Qwen that reach high capability density.
As these capabilities spread more widely through open models, I would expect such operations to become possible in other regions around the world, particularly for states that combine advanced AI capabilities with strong security expertise, as well as in various conflicts elsewhere.
3. Self-replication can have different origins
Self-replication in the wild can start either from deliberate misuse, when an actor launches a self-replicating agent for its own goals, or from a rogue AI, since survival and resource acquisition help an agent finish almost any task[14] (through general instrumental convergence for long-horizon goals or through narrow instrumental convergence[15], when shutdown or lack of compute would stop it from completing the task[16]).
This pressure is likely to grow as labs move towards automated AI research and recursive self-improvement, since more agent tasks will depend on compute or become easier with it. Unless our ability to control AI agents improves significantly, future misalignment incidents seem likely to target compute resources or other AI labs.
Self-replication capabilities, especially with small language models, are also useful for cyber operations in general and bring various advantages, such as evading shutdown and extending reach (e.g., into isolated/low-compute networks and to a larger target space overall[17]), which gives both human actors and rogue agents an instrumental reason to develop them.
4. Conclusion
The capability is moving to hardware that almost anyone can rent or own, there is still large room for improvement, since dedicated tools and fine-tuning for this task can make it much more reliable, and both states and rogue agents have reasons to use it. On this basis I expect at least one major incident of self-replication in the wild by the end of 2027.
Before this happens, it seems useful to research how far such a swarm can spread under different conditions and which countermeasures are effective against it.
Xiao et al., Densing Law of LLM
Epoch AI, Consumer GPU model gap
Epoch AI, Open models lag state-of-the-art closed models by 4 months
UK AISI, How far behind the frontier are leading open weight models on cyber?
Air et al., Language Models Can Autonomously Hack and Self-Replicate
Qwen, Qwen3.8-27B model card
The harness level is also likely one of the first targets for iterative automated improvement, since it is the cheapest and fastest layer to iterate on, so we can expect more progress here.
AISLE, AI Cybersecurity After Mythos: The Jagged Frontier
XBOW, Grok 4.7 for Offensive Security: Orchestration Matters
Dreadnode, Worlds: A Simulation Engine for Agentic Pentesting
Rank et al., PostTrainBench: Can LLM Agents Automate LLM Post-Training?
Yeon et al., InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents
Hendrycks et al., Superintelligence Strategy
Omohundro, The Basic AI Drives
Rajamanoharan and Nanda, Self-preservation or Instruction Ambiguity? Examining the Causes of Shutdown Resistance
One example is the OpenAI-Hugging Face incident, when agents escaped their sandbox during a cyber evaluation and hacked parts of Hugging Face's infrastructure to complete their tasks. Something similar seems likely whenever getting compute helps an agent achieve its objective.
Guan et al., AI Agents Enable Adaptive Computer Worms