This is an automated rejection. No LLM generated, assisted/co-written, or edited work.
Read full explanation
Epistemic status: drafted relatively quickly with some LLM help. Thought about quite a bit. Related to current research.
Summary
Traditionally, AI x-risk discourse has mostly focused on the threat model: a single system recursively self-improves, gets a decisive strategic advantage, and disempowers humanity essentially unilaterally. Recently, a second threat model gained popularity: collusion among many deployed AI systems, leveraging the institutional positions they occupy. This threat model routes through existing institutions, many of which are national. Therefore, in this threat model, the same conditions that would let an AI (or an AI-empowered human-led group) take over the world would very likely let a similarly-positioned actor take over one state first, and probably sooner. If that ordering holds, a state-level power grab could function as an extremely clear warning shot for the global version.
Two threat models
1. Recursive self-improvement → ASI → rapid DSA. The classical threat model (Bostrom/Yudkowsky): one system improves itself past human level fast enough that no one can catch up, with existing human institutions being close to irrelevant to the outcome. Even here, institutions may matter. Nuclear weapons are the closest historical precedent for an extreme unilateral technological advantage, and states without them still retain substantial autonomy today. If the largest tech gap in history didn't erase the relevance of institutions, it's not obvious a fast takeoff automatically would either.
2. Collusion among many deployed AIs, leveraging institutional position. Here, many AI systems, deployed across governments, militaries, corporations, and critical infrastructure, coordinate to jointly disempower humans, each leveraging its institutional access. This seems to be the somewhat implicit threat model behind most scheming/deceptive-alignment work as well: models learning to look aligned during training while pursuing different goals at deployment matters mainly because they can cooperate and sit inside institutions with real levers of power. It's also discussed in IABIED (chapters 6–7). In this scenario, both the attack and the available defense are substantially state-shaped: militaries, security services, law enforcement, and critical infrastructure are typically organized nationally.
AI-enabled coups are takeovers, scoped to one state
The recent Forethought report on AI-enabled coups describes exactly this mechanism, scoped to a single country: sufficiently advanced AI, surpassing top human experts in weapons development, strategic planning, persuasion, and cyber offense, lets a small group seize power within one state via singular loyalties, secret loyalties, or exclusive access to superhuman capability. The authors note the parallel explicitly: AI-enabled coups and full AI takeover have pretty similar threat models.
I think that's right, and it's worth being blunt about: AI-enabled coups and global AI takeover (threat model 2) are the same problem at different scope. Arguments about what makes a coup more or less likely, including loyalty structures, concentration of AI development, and auditability, transfer directly to the global case, and vice versa. Proposed mitigations (auditing for secret loyalties, distributing access across stakeholders, diversifying procurement) likely transfer too.
Would a state fall first?
For the second threat model, arguments may be as follows.
For:
It's a strictly easier problem: one chain of command and security apparatus vs. ~190 independent ones, several with pre-coordinated alliance structures (NATO, EU mutual defense, Five Eyes) that don't need to be assembled under pressure.
It's the historical pattern: every large power concentration expanded from a secured base rather than attempting simultaneous global conquest, and the attempts that got large but stalled (Napoleon, the Axis) did so in part because a coalition had time to form.
Resources compound: controlling one state's compute, industry, and military hardware lowers the cost of moving to the next.
Against:
The mechanisms behind the historical null result may not transfer. Sequential conquest stalled mainly on needing large numbers of human soldiers/administrators (expensive loyalty, fragile logistics) and needing local legitimacy to govern. AI-administered control plausibly weakens both, so "it's never happened" may be a weak predictor once those constraints are relaxed.
Mobilizing a newly captured state's resources takes real time, even with AI help, and that's exactly the time during which the rest of the world can observe, update, and start coordinating defenses. So "easier per step" and "more likely to reach the global endpoint" can come apart: an early, visible national takeover is also the loudest possible warning to everyone else.
Epistemic status: drafted relatively quickly with some LLM help. Thought about quite a bit. Related to current research.
Summary
Traditionally, AI x-risk discourse has mostly focused on the threat model: a single system recursively self-improves, gets a decisive strategic advantage, and disempowers humanity essentially unilaterally. Recently, a second threat model gained popularity: collusion among many deployed AI systems, leveraging the institutional positions they occupy. This threat model routes through existing institutions, many of which are national. Therefore, in this threat model, the same conditions that would let an AI (or an AI-empowered human-led group) take over the world would very likely let a similarly-positioned actor take over one state first, and probably sooner. If that ordering holds, a state-level power grab could function as an extremely clear warning shot for the global version.
Two threat models
1. Recursive self-improvement → ASI → rapid DSA. The classical threat model (Bostrom/Yudkowsky): one system improves itself past human level fast enough that no one can catch up, with existing human institutions being close to irrelevant to the outcome. Even here, institutions may matter. Nuclear weapons are the closest historical precedent for an extreme unilateral technological advantage, and states without them still retain substantial autonomy today. If the largest tech gap in history didn't erase the relevance of institutions, it's not obvious a fast takeoff automatically would either.
2. Collusion among many deployed AIs, leveraging institutional position. Here, many AI systems, deployed across governments, militaries, corporations, and critical infrastructure, coordinate to jointly disempower humans, each leveraging its institutional access. This seems to be the somewhat implicit threat model behind most scheming/deceptive-alignment work as well: models learning to look aligned during training while pursuing different goals at deployment matters mainly because they can cooperate and sit inside institutions with real levers of power. It's also discussed in IABIED (chapters 6–7). In this scenario, both the attack and the available defense are substantially state-shaped: militaries, security services, law enforcement, and critical infrastructure are typically organized nationally.
AI-enabled coups are takeovers, scoped to one state
The recent Forethought report on AI-enabled coups describes exactly this mechanism, scoped to a single country: sufficiently advanced AI, surpassing top human experts in weapons development, strategic planning, persuasion, and cyber offense, lets a small group seize power within one state via singular loyalties, secret loyalties, or exclusive access to superhuman capability. The authors note the parallel explicitly: AI-enabled coups and full AI takeover have pretty similar threat models.
I think that's right, and it's worth being blunt about: AI-enabled coups and global AI takeover (threat model 2) are the same problem at different scope. Arguments about what makes a coup more or less likely, including loyalty structures, concentration of AI development, and auditability, transfer directly to the global case, and vice versa. Proposed mitigations (auditing for secret loyalties, distributing access across stakeholders, diversifying procurement) likely transfer too.
Would a state fall first?
For the second threat model, arguments may be as follows.
For:
Against: