TLDR: In this piece I posit that it is in the American government's interest to immediately begin high-level dialogue with the Chinese and implement something along the lines of Plan A, and that between the two policies of strategic exclusion[1] and global pacing, the latter presents lower short-term risk and is a wash on long-term risks. I argue that currently, a largely underlooked short-term risk in terms of probability-weighted impact is kinetic war in Taiwan. For other countries, I argue that immediately building airgapped, non-internet backups for catastrophe planning may mitigate some harms that could potentially accrue.
1. Why This Moment Is Geopolitical
Once upon a time, Bluebeard and Redbeard were two of the most fearsome pirate captains on the high seas. Stories of their plundering spread far and wide, but most of all they were known for their ultimate goal: to collect the cursed treasures of Blackbeard. As legend had it, before the infamous Blackbeard disappeared, he carved his ill-gotten treasures into hundreds of little portions, each of which was enough to make smaller pirate crews retire comfortably. When put together, these treasures would recombine into powerful relics that granted the holders supernatural powers. However, the tricky pirate king put a curse on the loot:
Whosoever commands 5 more pieces of my loot than his opponent may summon a dreadful vortex to sink his enemy's ship to the bottom of the ocean, but he who collects ███ pieces will draw the attention of the sleeping leviathan, whose hungry maws will devour every last pirate on the seas.
He never specified exactly how many pieces would awaken the leviathan from its hibernation, but up to this point nobody had crossed the threshold yet. Several decades after his disappearance, the two pirate crews of Bluebeard and Redbeard were sufficiently well-resourced to pursue the treasures in earnest; ever since they embarked on their missions, the powers of the relics they gathered snowballed their strength further and further beyond what any other pirate crew could hope to compete with, to the point that their two crews and their motley allied ships were the only two forces of any significance on the seas.
Between the two, Bluebeard was the first to start searching for the buried treasures, and as a result he maintained a 2-3 piece advantage over Redbeard over the years. But Redbeard had a different tactic as the second mover; instead of searching far and wide, he imitated the movements of Bluebeard's crew and often found overlooked relics for a far lower cost than Bluebeard. As a result, Bluebeard's crew felt the constant threat that Redbeard posed: if they slowed down even a bit, they were afraid that their opponents would quickly catch up to them, which would put their own chances of survival into question.
At some point, Bluebeard's fleet managed to complete 6 sets of relics. On that day, a great groan was heard all across the world, and the surface of the ocean churned and chopped. There were whispers that the curse was nearing its hidden limit, and that if Bluebeard continued down his path at the same rate, he was bound to awaken the leviathan very soon.
Bluebeard convenes an urgent meeting with his most trusted advisors. His advisors are roughly split into three camps: One-eye thinks that the fears of the leviathan's awakening are greatly overexaggerated, and possibly part of a covert operation by Redbeard's allies to make Bluebeard's crew slow down their search for the treasures in order to sneakily catch up; Hookhand thinks that the rumors should be taken seriously, but that all they have to do is to send more skirmishers to slow down Redbeard while covering up their own tracks better, so that their lead would not be eroded while they think of a plan to fight the leviathan; and Parrot Guy agrees that the rumors should be taken seriously, while advocating for them to immediately reach out to Redbeard's crew for negotiations to bilaterally pause their search for more treasures.
Let's pause the analogy for a moment here and address the real-world equivalent of One-eye's argument. The equivalent of the leviathan here could be different things to different people. For many people, it would mean the presence of a misaligned (assuming that we can adequately define this) AGI / ASI, and for some they would consider it to be that plus some ability to exercise control over the physical world.[2] I consider the bar for such a "leviathan" to be far lower on both the ability and misalignment axes; to me a collection of agents that possesses superhuman intelligence in cybersecurity tasks, for which we have not fully understood the alignment level or controllability of, would be the catastrophe. Arguably, this moment may have already arrived or is not far away.[3]
The reasons for considering this to be the turning point are threefold. Firstly, the ability to hack into other systems at will would be a generally superintelligent misaligned AI's most basic offensive tool. At the point where this becomes possible, we must regard it with the same caution as we would to a generally superintelligent AI that has WMDs at its disposal, since the damage that the latter could inflict upon humans would already be mostly achievable. Secondly, superintelligence in cybersecurity alone should reduce our posterior confidence in alignment testing, since a misaligned AI would be able to obfuscate CoT and tool calls, engage in sabotage of monitoring environments, or otherwise be better at hiding its trail after hacking targets out of the scope of its evals.[4] Therefore, whatever catastrophic outcomes we would imagine to exist when misaligned ASI is brought into existence could already have been set into motion the moment we lose control of monitorability, and since the bad ending scenario here (ie. extinction) has infinitely negative utility to humanity, this is the point in time where our expected utility from developing ASI first becomes infinitely negative. Thirdly, it is technologically feasible for human actors to leverage its offensive cyber capabilities to inflict harm upon others with far less downside than they were previously able to. These may not be malicious actors in the traditional sense (eg. terrorist groups), but rather states and organizations with control over superintelligent models that have not had proper humanistic alignment. Already we see that humans are capable of magnifying their own harms unto others using AI.[5] I am far from the first person to argue that there is a "tremendous disproportionality" to cyberwarfare, and that it incurs less cost and less repercussions to the aggressor.[6] In a future where access to this technology is heavily guarded, it is necessary to fully align the humans in charge of this technology to wider humanitarian interests; if that sounds just as difficult as aligning the AI to humanitarian interests, then it is an equally compelling principled reason to take this threat seriously right now.
Assuming now that Bluebeard has dismissed One-eye's objection, we turn to the other two arguments.
The core contention between the tactics of containment versus negotiation is weighing the tradeoff between being strongly dominant over Redbeard's forces if successful and the risk of getting wiped out by the leviathan if unsuccessful. Within this analogized scenario, the costs (and risks) to Bluebeard of adopting a containment strategy would be:
Overestimating how much more leeway they have, and accidentally summoning the leviathan before they are ready.
They find it difficult to contain Redbeard quickly enough, and their expected time of widening the gap to a state of guaranteed victory in confrontation would be longer than their expected time until the leviathan is summoned.
They have to establish even tighter control over critical technologies, alienating potential allies on the periphery of their influence.[7]
Redbeard could be pushed into a corner, and with no better alternatives he could trigger direct confrontations to the expense of both sides.
On the grand scale, the primary risk to both sides are the possibilities of being annihilated by either the other side or the leviathan. Some methods of delaying the leviathan's awakening have been tested (there is a chance that they are unsound), and proponents of the strategic exclusion plan contend that those methods plus even tighter controls against Redbeard are sufficient to "bring them victory". Below, let us consider several possible states of the world w.r.t. these two risks:
A1: Redbeard is almost certainly doomed if the current trajectory of Bluebeard's crew is maintained.[8]
A2: The outcome of their rivalry is still strategically ambiguous.
B1: The current state of the treasure hunt is on course to summon the leviathan sooner than people would like (ie. before adequate preventive measures are devised).
B2: The current state of the treasure hunt is not on course to summon the leviathan, as adequate preventive measures exist to delay its arrival.
Below is a table of the rational actions for each actor conditional on them bothknowing exactly which state of the world currently exists.
A1
A2
B1
If Bluebeard does not slow down, everyone dies. Therefore, Bluebeard must slow down or pause. This is independent of what Redbeard does, so Redbeard will choose to continue development up until the frontier point.
B2
If Bluebeard continues on this current trajectory, Redbeard will certainly be doomed. Therefore, Redbeard will do everything he can to stop Bluebeard, potentially initiating a lose-lose physical war while he still has a chance to shift the balance. Neither side slows down their search.
Neither side has an incentive to stop, and will keep racing to find more pieces of the treasure. Whichever side stops will then get annihilated by the other side before the leviathan is summoned.
In actuality, neither side knows for sure which of the possible states of the world they exist in, nor do they know if the other agrees with their assessment of the present state. At the same time, it is hard for them to adopt a mixture of the strategies in the different scenarios, which are mostly[9] mutually exclusive to each other. Notice that the B risks here, which involve summoning the leviathan, are much harder to estimate than the A risks. Both sides would naturally have a worse estimate of an unknown new risk than the action set of their old adversaries. Therefore even if the responses in each scenario were not mutually exclusive, a mixed strategy still results in infinitely negative utility in expectation because there is a nonzero chance that the leviathan is summoned.
Is it possible to do better than infinitely negative EV? Let's consider Parrot Guy's proposal in the idealized framework. If we are in scenario B1, Bluebeard and Redbeard would easily agree to a treaty on slowing down their pace of looking for treasures. However, if neither side knows for sure that they are in scenario B1, or perhaps is underestimating or in a stage of denial about the leviathan's power, they might come up with familiar arguments like "the other side would not honor the treaty, and therefore we will lose our lead / be destroyed as a result" or "we have no way of enforcing the treaty". Firstly, a treaty might not be perfect, but one can imagine one which substantially hinders either side from defecting.[10] Secondly and more importantly, the problem with this is that the counterfactual is only maybe better if we are in scenario (A2, B2). In (A1, B2), Redbeard would be willing to accept a treaty as it prevents the lose-lose scenario of direct confrontation, and it would also be beneficial to Bluebeard if Redbeard's last resort actions are sufficiently threatening. If the captains do not know which scenario they are in, they would still accept a treaty in (A2, B2) since its effects are symmetric. I address some of the potential ways the counterfactual could be better in section 3.
Additionally, the enforcement of this treaty is time-sensitive. In scenario B1, there is very little time left before the leviathan is summoned. But even if we take B1 off the board, leaving the only uncertain risks between A1 or A2, now is the only time Bluebeard has to extend a treaty, while is still strategic uncertainty about who will come out in front.
Our fictional pirate captains are in a very similar predicament to the leaders of US and China today. We can make a few observations about the real-life complement:
Politicians seem to believe that we are in (A2, B2), or may have some non-public reasons to believe that (A1, B2) is not a big issue. They may overestimate the probability of these being true the former due to ignorance, power-seeking behaviors, and geopolitical fears. These may be influencing them to de-emphasize the probability of B1 in their decisions.
People in general are dismissive of B1 risks and downplay them as pre-IPO psyops from the big labs, ways to stifle competitors, etc. The public, outside of the small circle of people who pay attention to this stuff, is generally quite skeptical of the threats brought by AGI. Anecdotally, even people working in relevant policy believe some of these "conspiracies" of misalignment being exaggerated, or do not raise these issues to an appropriate level of seriousness due to a "fear of rocking the boat".
Given this, it is the urgent prerogative of those who understand the full implications of misaligned AI to lobby for the US government to take this seriously and grasp the short window in which China has noncombat options at the negotiating table. It is also useful to try and influence the general public and policymakers in other countries to adopt a more up-to-date view of the stakes at hand.
2. War Risk and Long Term Risks
In the scenario (A1, B2), where for Redbeard is so low in the relic race that he would prefer facing Bluebeard directly in combat, we assume that there are some trump cards that Redbeard may have, which he considers last-resort actions within his action space.
In our world, China has a strong incentive to try and hinder American AI progress by taking or blockading Taiwan. Earlier I've illustrated why they have an incentive to do so, but we still need to show that blockading Taiwan is a big deal, that they are able to do so successfully, and that it has a high risk of military escalation with America and its Pacific allies.
Firstly, the amount of compute required to train frontier models grows exponentially, so not being able to access the additional chips would mean missing a generation or two of model development, narrowing or erasing the lead the major labs currently have. In this aspect TSMC's exports from Taiwan are extremely crucial to the supply chain, with their gigafabs accounting for over 13M+ of their 17M 12-inch equivalent wafers/year. Their higher-end products, like their 3nm and 2nm chips, are basically all concentrated in Taiwan, with the first overseas 3nm production starting only in mid 2027 at Arizona Fab 2.[11] Their CoWoS capability is concentrated in Taiwan and the earliest we can expect there to be backups elsewhere would be in 2028 or 2029.[12] It is not possible to replace TSMC in the short term due to the multi-year process of building a fab and ramping up production in the fab. In short, incapacitating TSMC's plants in Taiwan would be fatal to frontier progress in the short and medium term by taking out large amounts of training and inference compute for next-gen models.[13]
Downstream of this is the implication that should China be able to actually take out TSMC's Taiwan fabs from the picture for a couple of months to a year, this would have major implications on the status of the AI race. Again, we assume the stakes here to be the worst-case scenario of "if one side has a significant lead they will wipe out the other", which is the most difficult scenario to consider; if the oppositional stakes were less high then there should proportionally be less resistance to a treaty agreement preventing type B1 risk.[14]
Let us be clear here: superintelligence is a piece of technology that could become an uncounterable weapon when used in a military context. With a sufficient gap in AI capabilities, one party may find it trivially simple to orchestrate indefensible attacks against another party, both digitally and eventually physically. Once this threat is only unilaterally possible, there is no guarantee of sovereignty or self-determination for whoever is on the receiving end of a superintelligence attack.
The base case in the (A1, B2) world is that the Chinese coast guard and the PLA navy engage in a blockade of Taiwan's main ports, with aerial and missile coverage if the US tries to circumvent this through air routes. The Iran war has provided us some precedent on this: there is no need for China to completely cordon off everything, simply establishing a hard-to-remove naval presence (and aerial presence) and laying mines is sufficient to disrupt shipping traffic. Additionally, stable electricity generation becomes a problem after just a few weeks, which could force semiconductor production to slow down or halt.
Even at this baseline level, you can see how difficult it would be for the US and its allies to somehow 1) provide a stable protected path for chips to enter and exit Taiwan and 2) ensure that production is not affected. The best case scenario under the CSIS wargame (linked above) was to set up a cargo transport route through the Ryukyu islands, which itself is extraordinarily difficult.
Most existing wargame scenarios are insufficiently imaginative about the scale of this military conflict. Given that superintelligence is essentially a life-or-death race between the US and China (we assumed so by caveat, but to the politicians in power it very well might be life or death), I would expect that should China's assessment of their AI capability gap be sufficiently severe, they will engage in more aggressive direct military action, and the US may have to response forcefully in kind. Several factors are in China's favor: most obviously, they are geographically closer to Taiwan, and their goal here is simpler, which is to disrupt instead of protect transport routes. This gives them an aerial advantage and access to ground-based missile systems, as well as the option of an amphibious assault.
The feasibility of ground troops taking over Taiwan is beyond the scope of this article, but again there is significant capability for disruption, if not the direct capturing of fabs as a strategic asset: all 6 gigafabs are in the west of Taiwan, which is both closer to China (and thus easy to disrupt in a scorched earth scenario) and also geographically more suitable for an amphibious landing.
Gigafab locations roughly circled in red. Specifically, the Taichung and Kaohsiung locations are close to ports where troops are likely to disembark. Recent test deployments of landing barges have also shown that the Chinese military has the capacity for troops and vehicles to be unloaded in more locations along the western coast.
There is significant risk of military escalation if the US and its allies attempt to intervene and either smuggle out or fight a way out of the cordon, which it very likely will due to the significance of these chips. Aside from breaching the most important of China's "red lines", this event would likely come at a critical point for both sides in the AI race, and China must prevent the US from getting its hands on more compute to accelerate model development. If no treaty is forced by military means, there is a long left tail of lose-lose outcomes that could occur for either side. There's been talk of military reunification of Taiwan for a long time, accompanied by a military buildup in surrounding area, for plenty of reasons other than AI. Therefore, our prior on this should be that confrontation will be the base case, and it will be done in a less cautious manner by the Chinese than if a major consideration was to disincentivize other nations from getting involved in its internal reunification affairs.
Now, let's suppose instead that a treaty as imagined above was signed. On one hand this disincentivizes blockades and other aggressive actions because that would amount to defecting on the treaty. On the other hand, it reduces the likelihood of China being forced to play their military hand in the short term. In the long term, however, we see that the worst-case scenario is a wash on existential risk compared to the counterfactual if we do get misaligned superintelligence, and conditional on the treaty working as intended for a reasonable amount of time, the threat of the opponent secretly training models is not more or less disadvantageous to either side ex ante. A working treaty would have some way for both sides to reliably conclude that the other has only a very limited capacity to train models in secret, which both sides will certainly take advantage of and will know the other is taking advantage of; the upper bound on this secret capacity will increase as time passes, until a point where the treaty is no longer relevant. The risks here are not more significant than in the case where both sides continue to race, as long as there is a nonzero chance of civilizational catastrophe from ASI, because the expected disutility from (they defect, we don't defect) cannot be lower than the infinite expected disutility from (they defect, we defect; ie. no treaty).
3. Some Possible Complications
Counterargument 1: Catastrophe is not infinitely negative utility.
It would be quite bizarre in our normal conversations that one would have to justify at all how "something might eventually kill us all" is really, really bad, to the point that nothing physically imaginable might be worse. But for the sake of the argument, we need to weigh against the possible benefit that could accrue to all the many eons of humankind that might exist in the future due to ASI accelerating us into techno-utopia.
Premise 1: ASI brings a nonzero chance of extinction, due to misaligned goals, unstoppable rogue agents, or a myriad of other possible scenarios we have not had the time to patch.[15]
Premise 2: Death represents infinitely negative utility for many individuals.
Premise 3: The total utility on the upside is finite due to entropy: between now and heat death, there will be a finite amount of conscious time observed by humans, and there is also a finite amount of pleasure that can ever be felt by humans.
Premise 4: The total utility of an action on "humanity" is the sum (where defined) of the utilities to individuals that could be affected by the action, that being present and future individuals.
Conclusion: If there is a chance of human extinction, and there are people who view death as having infinite disutility, then it is in aggregate of infinite disutility to humanity.
It is also quite bizarre that this point has to be made at all, since whatever really bad thing a malicious superhuman could do to you is surely worse than whatever a malicious human could do to you, and yet people are less worried about agents that so far have proven to be pretty willing to do things even rogue actors would think twice about, compared to their "adversary" in the AI race.
Counterargument 2: The treaty will fail spectacularly and backfire in some way.
The treaty failing is not worse than the counterfactual where there is no treaty. Specifically, one must argue that the treaty failing has some inherent harms that do not exist if we do not attempt to create a treaty at all. Some possible ways this could backfire is:
One side actually has a lot more hidden capacity than expected and trains in secret, and because it is not possible to retain public oversight for this, it ends up being even more dangerous than the alternative.
It is easy to train small models with very little compute to a decent level and go unnoticed, so leading companies are forced to add more features to maintain a moat, therefore they will attempt to train in secret anyways which is dangerous.
To either side in the treaty, they are scared that the treaty's terms would be asymmetrically unfair against them, and that they recognize this and take action too late.
In general, these concerns are some mix of "it will end badly for EVERYONE because there will be less oversight" and "it will end badly for US because they will defect faster than us".
For the first type of concern, ultimately we need to consider the effectiveness of open supervision placed in the context of a race. Public supervision does very little to address the core issue with a secret training race, which is the prisoner's dilemma, and thus having or not having oversight from the public makes little difference.
For the second type of concern, there is unfortunately little we can do other than to try and extrapolate from similar treaties in the past. Verification on demand (CWC) or allowing complementary access (IAEA) are possible considerations for an AI treaty, albeit requiring far more intrusive access since the covert development of AI is far easier to conceal. However, we can be relatively certain that a treaty would allow easy monitoring of large sites and electrical usage (eg. Stargate), which makes covert defection a lot harder.
There are very valid reasons why an AI treaty would be different from a nuclear one, such as the lack of a second strike capability, the difference in scale and monitorability of key components, the fact that missiles can be dismantled but not trained models, and so on. Fortunately, AI has displayed a propensity to hack into things when connected to the actual internet, even without being given instructions to do so.[16] Therefore, covert training would have to be entirely airgapped, in artificially constructed sandboxed environments, or the risk of being caught for defecting because the AIs did some unauthorized stuff again and left an online trail would be too high. Here, the potential for sandbagging is actually helpful, because having a fully sandboxed environment means that it's difficult for researchers to covertly train superintelligence that they can epistemically justify as being not evaluation-aware, as doing so would require them deploying the systems, which is both dangerous to the treaty and also existentially dangerous because they have no idea how it will behave outside of their test environments. (If we somehow figure out how to counter sandbagging and other alignment stuff outside of these covert runs, that's actually great news, because by then we wouldn't need the treaty any more.)
4. What's Next For Other Countries
Throughout this piece, I have left out countries other than the US and China from the picture. That is because in the grand scheme of things, no other country has the resources to catch up to the frontier labs in those two countries right now. Through the limited scope of international bodies and within their own countries, there are a few things that other governments might have to consider.
Firstly, smaller countries are no longer impervious to the fact that the barrier for a determined superpower to interfere in their national interests has already been lowered drastically and will continue to get lower. Australia is the current example, but next week or month there may well be other examples of countries being hacked. It is of urgent importance for other countries to recognize that a takeoff in capabilities has already started and that the status quo of friendly sovereign trading partners is a temporary state of the world.
At this point, other countries have very little direct stake in the outcome of the race. Lobbying and drafting resolutions in the UN to pressure the US and China while they currently possess any leverage would be one reasonable path of action. At the very least, the recognition of ASI as a potential WMD and strict limits on its public usage should be floated, and if possible it would be in their interest to negotiate for tangible stakes in superintelligence, instead of hoping for the goodwill of superpowers to extend to them in the future.
Other countries also need to build up offline communications and airgapped systems. We must treat internet capabilities as inherently dangerous and catastrophic internet hacks as guaranteed to exist in the future. Currently, many governments have digitized many of their services, along with core industries like banks, hospitals, and communications providers. Backup methods to prevent chaos in the case that rogue agents maliciously (or even inadvertently) mess with critical infrastructure are currently missing and need to be restored quickly.
Lastly, other countries should try and get involved in AI safety initiatives at the national level. In a world where superintelligence exists, countries that do not end up with direct access to frontier models will only be on the receiving end of externalities or direct actions from the superintelligence. The least a responsible government should do is to treat this as a matter of national security and try to coordinate with other countries to lobby AI superpowers before this happens.
"If we execute these measures well, I believe they would slow China’s progress enough to widen America’s lead significantly over the next 3–5 years — the window when AI becomes geopolitically most important ... [I believe] these measures increase the leverage held by democracies and make an agreement more likely in the future." Although not explicitly stated as a policy of strategic exclusion, this is in line with the "small yard, high fence" strategy of containment from the America's existing geopolitical playbook for China. In other words, the plan here is to try and further tighten restrictions around China right now and prevent them from accessing the strongest frontier models, while trying to walk the thin tightrope between superintelligence and giving China a chance to catch up.
The terminology of AGI / ASI comes with a lot of baggage in general public discourse, and even Yann LeCun is not immune to this. Its actual meaning is irrelevant; what we mean when we discuss the dangers of AGI does not depend on its ability to perform tasks like folding our clothes. The general arguments for the dangers of AGI have been thoroughly laid out elsewhere already, and for a good review of some of these arguments, see this review of If Anyone Builds It, Everyone Dies.
There is no shortage of examples where this has happened, but just to name a few recent ones off the top of my head: the HuggingFace attack, the RubyGems hack, and OpenAI's internal Artifactory hack. Individually the agents might not have strictly superintelligent cyber capabilities, but we are at the stage where in a swarm they are able to discover and exploit hacks far quicker than a group of highly trained human experts would be able to. Ultimately superhuman hacking capabilities should be measured by how much better the AIs are compared to the defense instead of compared to similar human offense. In that regard, frontier models have already demonstrated a strong ability to overwhelm cybersecurity defenses through zero-days.
On a damage:cost ratio, advanced cyber warfare is several OOMs more powerful compared to other more traditional offensive capabilities a state actor may possess, and it also happens to be the most easily trainable offensive action for an AI. That said, this argument does not only pertain to state actors, though a plausible scenario is a world in which superintelligence is owned by the state.
I recognize that this is a less important point considering the power of superintelligence. However, it is unlikely that a nation would reach the point of having access to unparalleled intelligence without significant pushback from other states prior to that happening. As we approach this point, other states are likely to put up some sort of resistance. We should not assume that sovereignty is inviolable, but rather that for most states currently, peaceful coexistence is the most positive outcome for both sides in expectation, but when regime change becomes as easy as sending a superintelligent swarm to cripple critical infrastructure and government communication systems, the barrier for superpowers to interfere in other countries' matters becomes much lower, and the act of doing so could possibly bear a positive expectation to the aggressors.
Physical war cannot be ruled out in the (A2, B2) case, but is much less likely modern warfare is highly -EV for both sides without overwhelming technological advantages, due to the destructiveness of modern arsenals and the difficulty of intercepting missiles, drones, etc.
Their overall share of 5nm and below chips is around 90%, of which another roughly 90% is in Taiwan: 24k wafer starts per month (wspm) in Arizona (p36) vs 220k wspm for only N5/N4 total. On top of that, all the chips produced in Arizona are still shipped to Taiwan for packaging and dicing anyways.
"[T]here could be an acceleration in capabilities that fizzles out: bottlenecks on data, training compute, inference compute, or experiments..." - METR, highlights my own
For example, suppose we discovered a gateway to hell at the bottom of the sea. It should be pretty obvious that America and China (and every other country) would immediately sign a treaty saying "DON'T OPEN THE GATES TO HELL!"
TLDR: In this piece I posit that it is in the American government's interest to immediately begin high-level dialogue with the Chinese and implement something along the lines of Plan A, and that between the two policies of strategic exclusion[1] and global pacing, the latter presents lower short-term risk and is a wash on long-term risks. I argue that currently, a largely underlooked short-term risk in terms of probability-weighted impact is kinetic war in Taiwan. For other countries, I argue that immediately building airgapped, non-internet backups for catastrophe planning may mitigate some harms that could potentially accrue.
1. Why This Moment Is Geopolitical
Once upon a time, Bluebeard and Redbeard were two of the most fearsome pirate captains on the high seas. Stories of their plundering spread far and wide, but most of all they were known for their ultimate goal: to collect the cursed treasures of Blackbeard. As legend had it, before the infamous Blackbeard disappeared, he carved his ill-gotten treasures into hundreds of little portions, each of which was enough to make smaller pirate crews retire comfortably. When put together, these treasures would recombine into powerful relics that granted the holders supernatural powers. However, the tricky pirate king put a curse on the loot:
He never specified exactly how many pieces would awaken the leviathan from its hibernation, but up to this point nobody had crossed the threshold yet. Several decades after his disappearance, the two pirate crews of Bluebeard and Redbeard were sufficiently well-resourced to pursue the treasures in earnest; ever since they embarked on their missions, the powers of the relics they gathered snowballed their strength further and further beyond what any other pirate crew could hope to compete with, to the point that their two crews and their motley allied ships were the only two forces of any significance on the seas.
Between the two, Bluebeard was the first to start searching for the buried treasures, and as a result he maintained a 2-3 piece advantage over Redbeard over the years. But Redbeard had a different tactic as the second mover; instead of searching far and wide, he imitated the movements of Bluebeard's crew and often found overlooked relics for a far lower cost than Bluebeard. As a result, Bluebeard's crew felt the constant threat that Redbeard posed: if they slowed down even a bit, they were afraid that their opponents would quickly catch up to them, which would put their own chances of survival into question.
At some point, Bluebeard's fleet managed to complete 6 sets of relics. On that day, a great groan was heard all across the world, and the surface of the ocean churned and chopped. There were whispers that the curse was nearing its hidden limit, and that if Bluebeard continued down his path at the same rate, he was bound to awaken the leviathan very soon.
Bluebeard convenes an urgent meeting with his most trusted advisors. His advisors are roughly split into three camps: One-eye thinks that the fears of the leviathan's awakening are greatly overexaggerated, and possibly part of a covert operation by Redbeard's allies to make Bluebeard's crew slow down their search for the treasures in order to sneakily catch up; Hookhand thinks that the rumors should be taken seriously, but that all they have to do is to send more skirmishers to slow down Redbeard while covering up their own tracks better, so that their lead would not be eroded while they think of a plan to fight the leviathan; and Parrot Guy agrees that the rumors should be taken seriously, while advocating for them to immediately reach out to Redbeard's crew for negotiations to bilaterally pause their search for more treasures.
Let's pause the analogy for a moment here and address the real-world equivalent of One-eye's argument. The equivalent of the leviathan here could be different things to different people. For many people, it would mean the presence of a misaligned (assuming that we can adequately define this) AGI / ASI, and for some they would consider it to be that plus some ability to exercise control over the physical world.[2] I consider the bar for such a "leviathan" to be far lower on both the ability and misalignment axes; to me a collection of agents that possesses superhuman intelligence in cybersecurity tasks, for which we have not fully understood the alignment level or controllability of, would be the catastrophe. Arguably, this moment may have already arrived or is not far away.[3]
The reasons for considering this to be the turning point are threefold. Firstly, the ability to hack into other systems at will would be a generally superintelligent misaligned AI's most basic offensive tool. At the point where this becomes possible, we must regard it with the same caution as we would to a generally superintelligent AI that has WMDs at its disposal, since the damage that the latter could inflict upon humans would already be mostly achievable. Secondly, superintelligence in cybersecurity alone should reduce our posterior confidence in alignment testing, since a misaligned AI would be able to obfuscate CoT and tool calls, engage in sabotage of monitoring environments, or otherwise be better at hiding its trail after hacking targets out of the scope of its evals.[4] Therefore, whatever catastrophic outcomes we would imagine to exist when misaligned ASI is brought into existence could already have been set into motion the moment we lose control of monitorability, and since the bad ending scenario here (ie. extinction) has infinitely negative utility to humanity, this is the point in time where our expected utility from developing ASI first becomes infinitely negative. Thirdly, it is technologically feasible for human actors to leverage its offensive cyber capabilities to inflict harm upon others with far less downside than they were previously able to. These may not be malicious actors in the traditional sense (eg. terrorist groups), but rather states and organizations with control over superintelligent models that have not had proper humanistic alignment. Already we see that humans are capable of magnifying their own harms unto others using AI.[5] I am far from the first person to argue that there is a "tremendous disproportionality" to cyberwarfare, and that it incurs less cost and less repercussions to the aggressor.[6] In a future where access to this technology is heavily guarded, it is necessary to fully align the humans in charge of this technology to wider humanitarian interests; if that sounds just as difficult as aligning the AI to humanitarian interests, then it is an equally compelling principled reason to take this threat seriously right now.
Assuming now that Bluebeard has dismissed One-eye's objection, we turn to the other two arguments.
The core contention between the tactics of containment versus negotiation is weighing the tradeoff between being strongly dominant over Redbeard's forces if successful and the risk of getting wiped out by the leviathan if unsuccessful. Within this analogized scenario, the costs (and risks) to Bluebeard of adopting a containment strategy would be:
On the grand scale, the primary risk to both sides are the possibilities of being annihilated by either the other side or the leviathan. Some methods of delaying the leviathan's awakening have been tested (there is a chance that they are unsound), and proponents of the strategic exclusion plan contend that those methods plus even tighter controls against Redbeard are sufficient to "bring them victory". Below, let us consider several possible states of the world w.r.t. these two risks:
A1: Redbeard is almost certainly doomed if the current trajectory of Bluebeard's crew is maintained.[8]
A2: The outcome of their rivalry is still strategically ambiguous.
B1: The current state of the treasure hunt is on course to summon the leviathan sooner than people would like (ie. before adequate preventive measures are devised).
B2: The current state of the treasure hunt is not on course to summon the leviathan, as adequate preventive measures exist to delay its arrival.
Below is a table of the rational actions for each actor conditional on them both knowing exactly which state of the world currently exists.
A1
A2
B1
If Bluebeard does not slow down, everyone dies. Therefore, Bluebeard must slow down or pause. This is independent of what Redbeard does, so Redbeard will choose to continue development up until the frontier point.
B2
If Bluebeard continues on this current trajectory, Redbeard will certainly be doomed. Therefore, Redbeard will do everything he can to stop Bluebeard, potentially initiating a lose-lose physical war while he still has a chance to shift the balance. Neither side slows down their search.
Neither side has an incentive to stop, and will keep racing to find more pieces of the treasure. Whichever side stops will then get annihilated by the other side before the leviathan is summoned.
In actuality, neither side knows for sure which of the possible states of the world they exist in, nor do they know if the other agrees with their assessment of the present state. At the same time, it is hard for them to adopt a mixture of the strategies in the different scenarios, which are mostly[9] mutually exclusive to each other. Notice that the B risks here, which involve summoning the leviathan, are much harder to estimate than the A risks. Both sides would naturally have a worse estimate of an unknown new risk than the action set of their old adversaries. Therefore even if the responses in each scenario were not mutually exclusive, a mixed strategy still results in infinitely negative utility in expectation because there is a nonzero chance that the leviathan is summoned.
Is it possible to do better than infinitely negative EV? Let's consider Parrot Guy's proposal in the idealized framework. If we are in scenario B1, Bluebeard and Redbeard would easily agree to a treaty on slowing down their pace of looking for treasures. However, if neither side knows for sure that they are in scenario B1, or perhaps is underestimating or in a stage of denial about the leviathan's power, they might come up with familiar arguments like "the other side would not honor the treaty, and therefore we will lose our lead / be destroyed as a result" or "we have no way of enforcing the treaty". Firstly, a treaty might not be perfect, but one can imagine one which substantially hinders either side from defecting.[10] Secondly and more importantly, the problem with this is that the counterfactual is only maybe better if we are in scenario (A2, B2). In (A1, B2), Redbeard would be willing to accept a treaty as it prevents the lose-lose scenario of direct confrontation, and it would also be beneficial to Bluebeard if Redbeard's last resort actions are sufficiently threatening. If the captains do not know which scenario they are in, they would still accept a treaty in (A2, B2) since its effects are symmetric. I address some of the potential ways the counterfactual could be better in section 3.
Additionally, the enforcement of this treaty is time-sensitive. In scenario B1, there is very little time left before the leviathan is summoned. But even if we take B1 off the board, leaving the only uncertain risks between A1 or A2, now is the only time Bluebeard has to extend a treaty, while is still strategic uncertainty about who will come out in front.
Our fictional pirate captains are in a very similar predicament to the leaders of US and China today. We can make a few observations about the real-life complement:
Given this, it is the urgent prerogative of those who understand the full implications of misaligned AI to lobby for the US government to take this seriously and grasp the short window in which China has noncombat options at the negotiating table. It is also useful to try and influence the general public and policymakers in other countries to adopt a more up-to-date view of the stakes at hand.
2. War Risk and Long Term Risks
In the scenario (A1, B2), where for Redbeard is so low in the relic race that he would prefer facing Bluebeard directly in combat, we assume that there are some trump cards that Redbeard may have, which he considers last-resort actions within his action space.
In our world, China has a strong incentive to try and hinder American AI progress by taking or blockading Taiwan. Earlier I've illustrated why they have an incentive to do so, but we still need to show that blockading Taiwan is a big deal, that they are able to do so successfully, and that it has a high risk of military escalation with America and its Pacific allies.
Firstly, the amount of compute required to train frontier models grows exponentially, so not being able to access the additional chips would mean missing a generation or two of model development, narrowing or erasing the lead the major labs currently have. In this aspect TSMC's exports from Taiwan are extremely crucial to the supply chain, with their gigafabs accounting for over 13M+ of their 17M 12-inch equivalent wafers/year. Their higher-end products, like their 3nm and 2nm chips, are basically all concentrated in Taiwan, with the first overseas 3nm production starting only in mid 2027 at Arizona Fab 2.[11] Their CoWoS capability is concentrated in Taiwan and the earliest we can expect there to be backups elsewhere would be in 2028 or 2029.[12] It is not possible to replace TSMC in the short term due to the multi-year process of building a fab and ramping up production in the fab. In short, incapacitating TSMC's plants in Taiwan would be fatal to frontier progress in the short and medium term by taking out large amounts of training and inference compute for next-gen models.[13]
Downstream of this is the implication that should China be able to actually take out TSMC's Taiwan fabs from the picture for a couple of months to a year, this would have major implications on the status of the AI race. Again, we assume the stakes here to be the worst-case scenario of "if one side has a significant lead they will wipe out the other", which is the most difficult scenario to consider; if the oppositional stakes were less high then there should proportionally be less resistance to a treaty agreement preventing type B1 risk.[14]
Let us be clear here: superintelligence is a piece of technology that could become an uncounterable weapon when used in a military context. With a sufficient gap in AI capabilities, one party may find it trivially simple to orchestrate indefensible attacks against another party, both digitally and eventually physically. Once this threat is only unilaterally possible, there is no guarantee of sovereignty or self-determination for whoever is on the receiving end of a superintelligence attack.
The base case in the (A1, B2) world is that the Chinese coast guard and the PLA navy engage in a blockade of Taiwan's main ports, with aerial and missile coverage if the US tries to circumvent this through air routes. The Iran war has provided us some precedent on this: there is no need for China to completely cordon off everything, simply establishing a hard-to-remove naval presence (and aerial presence) and laying mines is sufficient to disrupt shipping traffic. Additionally, stable electricity generation becomes a problem after just a few weeks, which could force semiconductor production to slow down or halt.
Even at this baseline level, you can see how difficult it would be for the US and its allies to somehow 1) provide a stable protected path for chips to enter and exit Taiwan and 2) ensure that production is not affected. The best case scenario under the CSIS wargame (linked above) was to set up a cargo transport route through the Ryukyu islands, which itself is extraordinarily difficult.
Most existing wargame scenarios are insufficiently imaginative about the scale of this military conflict. Given that superintelligence is essentially a life-or-death race between the US and China (we assumed so by caveat, but to the politicians in power it very well might be life or death), I would expect that should China's assessment of their AI capability gap be sufficiently severe, they will engage in more aggressive direct military action, and the US may have to response forcefully in kind. Several factors are in China's favor: most obviously, they are geographically closer to Taiwan, and their goal here is simpler, which is to disrupt instead of protect transport routes. This gives them an aerial advantage and access to ground-based missile systems, as well as the option of an amphibious assault.
The feasibility of ground troops taking over Taiwan is beyond the scope of this article, but again there is significant capability for disruption, if not the direct capturing of fabs as a strategic asset: all 6 gigafabs are in the west of Taiwan, which is both closer to China (and thus easy to disrupt in a scorched earth scenario) and also geographically more suitable for an amphibious landing.
Gigafab locations roughly circled in red. Specifically, the Taichung and Kaohsiung locations are close to ports where troops are likely to disembark. Recent test deployments of landing barges have also shown that the Chinese military has the capacity for troops and vehicles to be unloaded in more locations along the western coast.
There is significant risk of military escalation if the US and its allies attempt to intervene and either smuggle out or fight a way out of the cordon, which it very likely will due to the significance of these chips. Aside from breaching the most important of China's "red lines", this event would likely come at a critical point for both sides in the AI race, and China must prevent the US from getting its hands on more compute to accelerate model development. If no treaty is forced by military means, there is a long left tail of lose-lose outcomes that could occur for either side. There's been talk of military reunification of Taiwan for a long time, accompanied by a military buildup in surrounding area, for plenty of reasons other than AI. Therefore, our prior on this should be that confrontation will be the base case, and it will be done in a less cautious manner by the Chinese than if a major consideration was to disincentivize other nations from getting involved in its internal reunification affairs.
Now, let's suppose instead that a treaty as imagined above was signed. On one hand this disincentivizes blockades and other aggressive actions because that would amount to defecting on the treaty. On the other hand, it reduces the likelihood of China being forced to play their military hand in the short term. In the long term, however, we see that the worst-case scenario is a wash on existential risk compared to the counterfactual if we do get misaligned superintelligence, and conditional on the treaty working as intended for a reasonable amount of time, the threat of the opponent secretly training models is not more or less disadvantageous to either side ex ante. A working treaty would have some way for both sides to reliably conclude that the other has only a very limited capacity to train models in secret, which both sides will certainly take advantage of and will know the other is taking advantage of; the upper bound on this secret capacity will increase as time passes, until a point where the treaty is no longer relevant. The risks here are not more significant than in the case where both sides continue to race, as long as there is a nonzero chance of civilizational catastrophe from ASI, because the expected disutility from (they defect, we don't defect) cannot be lower than the infinite expected disutility from (they defect, we defect; ie. no treaty).
3. Some Possible Complications
Counterargument 1: Catastrophe is not infinitely negative utility.
It would be quite bizarre in our normal conversations that one would have to justify at all how "something might eventually kill us all" is really, really bad, to the point that nothing physically imaginable might be worse. But for the sake of the argument, we need to weigh against the possible benefit that could accrue to all the many eons of humankind that might exist in the future due to ASI accelerating us into techno-utopia.
It is also quite bizarre that this point has to be made at all, since whatever really bad thing a malicious superhuman could do to you is surely worse than whatever a malicious human could do to you, and yet people are less worried about agents that so far have proven to be pretty willing to do things even rogue actors would think twice about, compared to their "adversary" in the AI race.
Counterargument 2: The treaty will fail spectacularly and backfire in some way.
The treaty failing is not worse than the counterfactual where there is no treaty. Specifically, one must argue that the treaty failing has some inherent harms that do not exist if we do not attempt to create a treaty at all. Some possible ways this could backfire is:
In general, these concerns are some mix of "it will end badly for EVERYONE because there will be less oversight" and "it will end badly for US because they will defect faster than us".
For the first type of concern, ultimately we need to consider the effectiveness of open supervision placed in the context of a race. Public supervision does very little to address the core issue with a secret training race, which is the prisoner's dilemma, and thus having or not having oversight from the public makes little difference.
For the second type of concern, there is unfortunately little we can do other than to try and extrapolate from similar treaties in the past. Verification on demand (CWC) or allowing complementary access (IAEA) are possible considerations for an AI treaty, albeit requiring far more intrusive access since the covert development of AI is far easier to conceal. However, we can be relatively certain that a treaty would allow easy monitoring of large sites and electrical usage (eg. Stargate), which makes covert defection a lot harder.
There are very valid reasons why an AI treaty would be different from a nuclear one, such as the lack of a second strike capability, the difference in scale and monitorability of key components, the fact that missiles can be dismantled but not trained models, and so on. Fortunately, AI has displayed a propensity to hack into things when connected to the actual internet, even without being given instructions to do so.[16] Therefore, covert training would have to be entirely airgapped, in artificially constructed sandboxed environments, or the risk of being caught for defecting because the AIs did some unauthorized stuff again and left an online trail would be too high. Here, the potential for sandbagging is actually helpful, because having a fully sandboxed environment means that it's difficult for researchers to covertly train superintelligence that they can epistemically justify as being not evaluation-aware, as doing so would require them deploying the systems, which is both dangerous to the treaty and also existentially dangerous because they have no idea how it will behave outside of their test environments. (If we somehow figure out how to counter sandbagging and other alignment stuff outside of these covert runs, that's actually great news, because by then we wouldn't need the treaty any more.)
4. What's Next For Other Countries
Throughout this piece, I have left out countries other than the US and China from the picture. That is because in the grand scheme of things, no other country has the resources to catch up to the frontier labs in those two countries right now. Through the limited scope of international bodies and within their own countries, there are a few things that other governments might have to consider.
Firstly, smaller countries are no longer impervious to the fact that the barrier for a determined superpower to interfere in their national interests has already been lowered drastically and will continue to get lower. Australia is the current example, but next week or month there may well be other examples of countries being hacked. It is of urgent importance for other countries to recognize that a takeoff in capabilities has already started and that the status quo of friendly sovereign trading partners is a temporary state of the world.
At this point, other countries have very little direct stake in the outcome of the race. Lobbying and drafting resolutions in the UN to pressure the US and China while they currently possess any leverage would be one reasonable path of action. At the very least, the recognition of ASI as a potential WMD and strict limits on its public usage should be floated, and if possible it would be in their interest to negotiate for tangible stakes in superintelligence, instead of hoping for the goodwill of superpowers to extend to them in the future.
Other countries also need to build up offline communications and airgapped systems. We must treat internet capabilities as inherently dangerous and catastrophic internet hacks as guaranteed to exist in the future. Currently, many governments have digitized many of their services, along with core industries like banks, hospitals, and communications providers. Backup methods to prevent chaos in the case that rogue agents maliciously (or even inadvertently) mess with critical infrastructure are currently missing and need to be restored quickly.
Lastly, other countries should try and get involved in AI safety initiatives at the national level. In a world where superintelligence exists, countries that do not end up with direct access to frontier models will only be on the receiving end of externalities or direct actions from the superintelligence. The least a responsible government should do is to treat this as a matter of national security and try to coordinate with other countries to lobby AI superpowers before this happens.
"If we execute these measures well, I believe they would slow China’s progress enough to widen America’s lead significantly over the next 3–5 years — the window when AI becomes geopolitically most important ... [I believe] these measures increase the leverage held by democracies and make an agreement more likely in the future." Although not explicitly stated as a policy of strategic exclusion, this is in line with the "small yard, high fence" strategy of containment from the America's existing geopolitical playbook for China. In other words, the plan here is to try and further tighten restrictions around China right now and prevent them from accessing the strongest frontier models, while trying to walk the thin tightrope between superintelligence and giving China a chance to catch up.
The terminology of AGI / ASI comes with a lot of baggage in general public discourse, and even Yann LeCun is not immune to this. Its actual meaning is irrelevant; what we mean when we discuss the dangers of AGI does not depend on its ability to perform tasks like folding our clothes. The general arguments for the dangers of AGI have been thoroughly laid out elsewhere already, and for a good review of some of these arguments, see this review of If Anyone Builds It, Everyone Dies.
There is no shortage of examples where this has happened, but just to name a few recent ones off the top of my head: the HuggingFace attack, the RubyGems hack, and OpenAI's internal Artifactory hack. Individually the agents might not have strictly superintelligent cyber capabilities, but we are at the stage where in a swarm they are able to discover and exploit hacks far quicker than a group of highly trained human experts would be able to. Ultimately superhuman hacking capabilities should be measured by how much better the AIs are compared to the defense instead of compared to similar human offense. In that regard, frontier models have already demonstrated a strong ability to overwhelm cybersecurity defenses through zero-days.
We already see examples of these actions, such as during the HuggingFace incident. For a more detailed exposition of this point, see this post.
See the Minab school bombing.
On a damage:cost ratio, advanced cyber warfare is several OOMs more powerful compared to other more traditional offensive capabilities a state actor may possess, and it also happens to be the most easily trainable offensive action for an AI. That said, this argument does not only pertain to state actors, though a plausible scenario is a world in which superintelligence is owned by the state.
I recognize that this is a less important point considering the power of superintelligence. However, it is unlikely that a nation would reach the point of having access to unparalleled intelligence without significant pushback from other states prior to that happening. As we approach this point, other states are likely to put up some sort of resistance. We should not assume that sovereignty is inviolable, but rather that for most states currently, peaceful coexistence is the most positive outcome for both sides in expectation, but when regime change becomes as easy as sending a superintelligent swarm to cripple critical infrastructure and government communication systems, the barrier for superpowers to interfere in other countries' matters becomes much lower, and the act of doing so could possibly bear a positive expectation to the aggressors.
Meaning, the probability of winning is below some threshold at which they must use every method at their disposal to survive.
Physical war cannot be ruled out in the (A2, B2) case, but is much less likely modern warfare is highly -EV for both sides without overwhelming technological advantages, due to the destructiveness of modern arsenals and the difficulty of intercepting missiles, drones, etc.
For some examples of the real world equivalent to this, refer to this MIRI proposal, or the Covert Projects supplement to Plan A.
Their overall share of 5nm and below chips is around 90%, of which another roughly 90% is in Taiwan: 24k wafer starts per month (wspm) in Arizona (p36) vs 220k wspm for only N5/N4 total. On top of that, all the chips produced in Arizona are still shipped to Taiwan for packaging and dicing anyways.
Amkor's advanced packaging campus is expected to start producting early 2028 and TSMC's own chip packaging plant in Arizona is planned for 2029.
"[T]here could be an acceleration in capabilities that fizzles out: bottlenecks on data, training compute, inference compute, or experiments..." - METR, highlights my own
For example, suppose we discovered a gateway to hell at the bottom of the sea. It should be pretty obvious that America and China (and every other country) would immediately sign a treaty saying "DON'T OPEN THE GATES TO HELL!"
The predictions in IABIED have unfortunately largely held.
For example, the Australian government hack was related to a data gathering task.