Statement on AI use: AI models (mostly Fable 5 and Opus 5) were extremely helpful in 1. iterating through lots of variations on the game theory models presented, 2. helping to confirm my understanding of the math, 3. fact checking, finding sources, and catching errors, and 4. providing editorial feedback. The game theory visualizations are entirely vibecoded; I have checked the math, but not the code. All prose is my own except where marked and I take full responsibility for everything in the post.
I. Introduction
When I teach my philosophy students about existential risk from artificial superintelligence, many students' first reaction after fully grasping the argument is "Why are we still building this?" And the general public arguably shares the sentiment: an Opus 5 agent found that
63% of Americans told Pew in February 2026 that AI is advancing too quickly (against 2% who said too slowly), and a Yahoo News/YouGov poll the previous October found that 53% think it at least somewhat likely that AI will 'destroy humanity' someday.
Why not just shut it all down until we know it's safe? The Discourse has an answer to this question: it's because AGI is an arms race. Here is Leopold Aschenbrenner in 2024:
Some hope for some sort of international treaty on safety. This seems fanciful to me. The world where both the CCP and USG are AGI-pilled enough to take safety risk seriously is also the world in which both realize that international economic and military predominance is at stake, that being months behind on AGI could mean being permanently left behind. If the race is tight, any arms control equilibrium, at least in the early phase around superintelligence, seems extremely unstable. In short, "breakout" is too easy: the incentive (and the fear that others will act on this incentive) to race ahead with an intelligence explosion, to reach superintelligence and the decisive advantage, too great. At the very least, the odds we get something good-enough here seem slim. (How have those climate treaties gone? That seems like a dramatically easier problem compared to this.) (Source)
And here is Hunter Ash earlier this year:
AGI is not analogous to nukes
I should start by saying I'm about as far from an AI safetyist as it's possible to be. But I'm going to engage with that frame, and show why the situation is in no way analogous to that of nuclear weapons, which is their only real example of the kind of international coordination that would be required to stop AI progress.
The basic issue is this: you can build nuclear weapons but not use them. And this is what mutually assured destruction hangs on. On the safetyist view, you cannot do this with AGI. Making it and "firing" it are the same action.
The game with nukes is this: if we shoot nukes at another nuclear-armed country, there is an extremely high chance that every major population center in our country gets wiped out by their nukes. And the same is true for them. So no one fires. No one defects.
With AGI, the only way to not defect is to not be able to return fire if the other side defects. Unless your P(doom) is well north of 50%, then the deterrence math doesn't work. If the other team builds it in secret then turns it on, and they did achieve decent alignment, then you're absolutely screwed. There's no mutually assured destruction. They win and you lose. You can't suddenly build your own in retaliation when the missiles are already in the air, so to speak.
I also think compliance with any treaty would be near-impossible to enforce, but that's a subject for another post.
We cannot get out of this arms race even if we wanted to. Winner takes all, and there is no pause button. (Source)
So the argument is basically that if we don't build it, then they will, where "they" could be DeepMind or Anthropic or OpenAI or China—especially China. And if they build it, that would be worse. We'd like to stop, you see, but we can't. That is to say, it has the structure of a Prisoner's Dilemma. A Prisoner's Dilemma is defined by the fact that Defect is the dominant strategy, even though both players prefer (Cooperate, Cooperate) to (Defect, Defect). The argument is that, even though both the US and China (or both Anthropic and OpenAI) would prefer that they both not build, they also both prefer that they themselves build, regardless of what the other does.
The problem is that the AI race is not a Prisoner's Dilemma. I'm not the first to point this out—Katja Grace has talkedaboutit inseveralposts over the years, and there are a coupleofpapers making a similar point.[1](And Zvi actually made the point in one of his newsletters while I was in the process of writing this post.) But since the Discourse at large still seems to mostly take the "AI is a prisoner's dilemma" meme for granted, I want to go through in detail why it is not.
II. Nuclear Arms Races
A classic "arms race" (e.g. nuclear weapons) is a Prisoner's Dilemma. A Prisoner's Dilemma is defined by four payoffs, which we can call R, S, T, and P:
The first entry in each cell is the Row player's payoff, the second is the Column player's payoff. R is the payoff if both cooperate (the "reward payoff"). S is the payoff for cooperating against a defector (the "sucker's payoff"). T is the payoff for defecting against a cooperator (the "temptation payoff"). And P is the payoff if both defect (the "punishment payoff"). Prisoner's Dilemma is defined by the fact that .
The nuclear arms race might look like this. Building nukes costs 1, but if you beat your opponent to nukes you get to take 2 from them. If both players build, there's a standoff, and both pay the cost of building without getting to take anything from the other. We can represent this with the following payoff matrix:
This is how the game structure looks for building nuclear weapons. But Hunter Ash, in his tweet, is talking about the game structure for using them—the classic "Mutually Assured Destruction" scenario. Given a few assumptions, this turns out to be a Stag Hunt game. A Stag Hunt is defined by the fact that : Both prefer cooperating (hunting Stag together) if the other side will cooperate, but both prefer defecting (hunting Rabbit alone) if the other side will defect, and both prefer mutual cooperation to mutual defection.
This means that, unlike the Prisoner's Dilemma which has only one equilibrium, namely (Defect, Defect),[2]Stag Hunt has two equilibria, (Stag, Stag) and (Rabbit, Rabbit). And both players prefer (Stag, Stag), so as long as each can be assured the other will hunt Stag, that's what they'll choose.
In MAD, we assume that both players have a "safe second strike," meaning that if the other player attacks them they can safely retaliate, but not a "safe first strike" (meaning they can't strike first without being retaliated against). We'll also assume that, by striking first, you can diminish your opponent's second strike capacity, even if you can't fully eliminate it.[3]Under these assumptions, the payoffs look something like this:
In this case, both sides prefer to refrain if the other side will (because the -2 they'd get from being retaliated against is less than the 0 they'd get by refraining), but they also prefer to strike if the other will (because they have a chance of diminishing the other's second-strike capacity by striking first, meaning in expectation they get -4 instead of -5 by striking if the opponent does).
III. The AI Race
It turns out that, contra Hunter Ash, this structure is closely analogous to the AI race, under some plausible assumptions.
We can define R, S, T, and P for the AI race with a few parameters. Let's say that:
is the payoff from winning the race and creating an aligned AI
is the payoff from losing the race and getting dominated by the opponent's aligned AI (which will be negative).
Each side will use their aligned ASI to take some amount from the other, so Winning is just as good as Losing is bad; i.e., [4]
Each side has an equal chance of getting to ASI first, so ignoring the risk from loss of control, the expected payoff if both race is (which we can call ).[5]
, the chance that the winning side loses control of their AI.[6](I'll generally continue to call this P(Doom) in the prose parts of this essay)
is the payoff if the winning side loses control of their AI, the Doom payoff.
If neither side builds, nothing happens.
Given this:
: if neither side builds, nothing happens
The other three payoffs are lotteries between Doom and some other outcome , of the form , where is either , , or depending on which payoff we're talking about:
Basically: if no one builds, nothing happens. If one side builds, that side has probability of aligning their AI; if they succeed, they get the Winner's payoff and the other side gets the Loser's payoff, if they fail, they get the Doom payoff. If both sides build, they each have a 50-50 shot of winning, and the winner has probability of aligning their AI and getting the Winner's payoff vs failing to align their AI and getting the Doom payoff.
Now we can just build a model and see that the game type depends on the payoffs and on :
So we can see that for low values of P(Doom), the game is indeed a Prisoner's Dilemma, but for higher values of P(Doom) it becomes a Stag Hunt.[7]There's a crossover point where the game type switches, and the value of only depends on the ratio between and . Specifically, the crossover point happens when . So we can write , and some algebra gets us . Since is negative, the meaning of this is more transparent if we write it as .[8]
So the worse you consider Doom to be (relative to the spoils of victory), the lower P(Doom) has to be for the game to be a Stag Hunt.
Because a Stag Hunt has an equilibrium at (Pause, Pause), an agreement not to build ASI is feasible as long as the game is a Stag Hunt. (Note: I mean game-theoretically feasible. Obviously there are other factors that impact the feasibility of an agreement—messy politics, personalities and relationship between the specific parties involved, etc. I'm not making any claims here about those.)[9]In a Stag Hunt, once both sides are sure the other will comply, a treaty doesn't need enforcement, strictly speaking: neither side has an incentive to defect, and both sides prefer the cooperative equilibrium. The hard part is verification and assurance.
This means that (contra Hunter Ash) you don't need P(Doom) "well north of 50%" for an agreement on not building ASI to be feasible (you only need P(Doom) to be this high if equals ). In fact, if you think Doom is at least 9x worse than Winning the race is good, then P(Doom) only needs to be 10% for an agreement to be feasible. This is a pretty conservative assumption I'd say, since Doom implies complete loss of human control of the future, and likely extinction. I'd go so far as to say it's plausible that Doom is at least 99x worse than Winning is good, in which case P(Doom) only needs to be 1% for an agreement to be feasible.
And most of those with decisionmaking authority here take Doom seriously enough that this math applies to them. Most Western frontier lab CEOs are on the record saying that AI takeover is a real concern; Dario Amodei and Elon Musk havegiven explicit probabilities of 10%+. Sam Altman and Demis Hassabis have both criticized the "P(Doom)" framing, but Sam has said it's "not zero" and Demis has said it's "definitely non-zero and it's probably non-negligible." And all except Elon (and Mark Zuckerberg) signed the 2023 CAIS statement saying that "Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war." Even Xi Jinping mentioned "loss of control" risks from AI in a recent speech, though, as ever when looking at China as a layperson, it's hard to know what to make of this. But in any case, by lab CEOs' own lights, cooperation with China should be (game-theoretically) possible.
In conclusion: STOP SAYING THE AI RACE IS A PRISONER'S DILEMMA![10]
Appendix: Expanded Model
I haven't had a chance to read all of these yet, but I have had Fable check how they relate to what I'm doing here, and although they are making a similar point, they use somewhat different modeling assumptions. ↩︎
With perfectly secure second strikes, the game is actually a "Harmony" game: both sides prefer to Refrain unilaterally. But second strikes in the real world are never perfectly secure. ↩︎
Nothing much interesting happens if we relax this assumption. If Losing is worse than Winning is good, the game type doesn't change. If Losing is less bad than Winning is good, all that happens is that, for low enough values of P(Doom), the game becomes a "Deadlock", where both players actually prefer (Build, Build) to (Pause, Pause). You can see this in the expanded model in a later footnote. ↩︎
Again, nothing much interesting happens if we relax this assumption. If one side is favored to get to ASI first, then for low enough values of P(Doom), the game can be a Deadlock for them even if it isn't for the other side. You can see this in the expanded model in a later footnote. ↩︎
We'll assume that whichever ASI comes online first can defeat another ASI that comes online later, or prevent it from becoming ASI in the first place. So whether ASI is aligned or misaligned depends only on whether the winner's ASI is aligned or misaligned—once a misaligned ASI exists, there's no chance to defend against it with an aligned ASI of your own, and conversely, once an aligned ASI is created, all future ASIs will be aligned as well. ↩︎
For it actually becomes a Harmony game, but is not 1 so this is irrelevant for real-world analysis. ↩︎
Note that, in particular, it does not depend on L, or on the probability of winning the race given that both sides Build. So our two assumptions from earlier don't affect , and Ash's worry from the introduction about how bad it would be to lose the race never takes us from a Stag Hunt to a Prisoner's Dilemma. Here is an expanded model with these parameters allowed to vary that you can play with in the appendix at the end of the post. ↩︎
As an aside, this is basically isomorphic to another well-known variant of the nuclear arms race problem. There was a serious fear, at one point during the process of building the atomic bomb, that setting it off could ignite the atmosphere and kill everyone on earth. In actuality, before testing the bomb they did the calculations and found that the probability of this was small, but the same logic applies: if the probability of killing everyone were high enough, then coordination not to build the bomb would be a Stag Hunt. ↩︎
Unless your own P(Doom), W, and D satisfy (e.g. P(Doom) below 10% if you think Doom is only 9x as bad as Winning is good). ↩︎
Statement on AI use: AI models (mostly Fable 5 and Opus 5) were extremely helpful in 1. iterating through lots of variations on the game theory models presented, 2. helping to confirm my understanding of the math, 3. fact checking, finding sources, and catching errors, and 4. providing editorial feedback. The game theory visualizations are entirely vibecoded; I have checked the math, but not the code. All prose is my own except where marked and I take full responsibility for everything in the post.
I. Introduction
When I teach my philosophy students about existential risk from artificial superintelligence, many students' first reaction after fully grasping the argument is "Why are we still building this?" And the general public arguably shares the sentiment: an Opus 5 agent found that
63% of Americans told Pew in February 2026 that AI is advancing too quickly (against 2% who said too slowly), and a Yahoo News/YouGov poll the previous October found that 53% think it at least somewhat likely that AI will 'destroy humanity' someday.
Why not just shut it all down until we know it's safe? The Discourse has an answer to this question: it's because AGI is an arms race. Here is Leopold Aschenbrenner in 2024:
And here is Hunter Ash earlier this year:
So the argument is basically that if we don't build it, then they will, where "they" could be DeepMind or Anthropic or OpenAI or China—especially China. And if they build it, that would be worse. We'd like to stop, you see, but we can't. That is to say, it has the structure of a Prisoner's Dilemma. A Prisoner's Dilemma is defined by the fact that Defect is the dominant strategy, even though both players prefer (Cooperate, Cooperate) to (Defect, Defect). The argument is that, even though both the US and China (or both Anthropic and OpenAI) would prefer that they both not build, they also both prefer that they themselves build, regardless of what the other does.
The problem is that the AI race is not a Prisoner's Dilemma. I'm not the first to point this out—Katja Grace has talked about it in several posts over the years, and there are a couple of papers making a similar point.[1](And Zvi actually made the point in one of his newsletters while I was in the process of writing this post.) But since the Discourse at large still seems to mostly take the "AI is a prisoner's dilemma" meme for granted, I want to go through in detail why it is not.
II. Nuclear Arms Races
A classic "arms race" (e.g. nuclear weapons) is a Prisoner's Dilemma. A Prisoner's Dilemma is defined by four payoffs, which we can call R, S, T, and P:
The first entry in each cell is the Row player's payoff, the second is the Column player's payoff. R is the payoff if both cooperate (the "reward payoff"). S is the payoff for cooperating against a defector (the "sucker's payoff"). T is the payoff for defecting against a cooperator (the "temptation payoff"). And P is the payoff if both defect (the "punishment payoff"). Prisoner's Dilemma is defined by the fact that .
The nuclear arms race might look like this. Building nukes costs 1, but if you beat your opponent to nukes you get to take 2 from them. If both players build, there's a standoff, and both pay the cost of building without getting to take anything from the other. We can represent this with the following payoff matrix:
This is how the game structure looks for building nuclear weapons. But Hunter Ash, in his tweet, is talking about the game structure for using them—the classic "Mutually Assured Destruction" scenario. Given a few assumptions, this turns out to be a Stag Hunt game. A Stag Hunt is defined by the fact that : Both prefer cooperating (hunting Stag together) if the other side will cooperate, but both prefer defecting (hunting Rabbit alone) if the other side will defect, and both prefer mutual cooperation to mutual defection.
This means that, unlike the Prisoner's Dilemma which has only one equilibrium, namely (Defect, Defect),[2]Stag Hunt has two equilibria, (Stag, Stag) and (Rabbit, Rabbit). And both players prefer (Stag, Stag), so as long as each can be assured the other will hunt Stag, that's what they'll choose.
In MAD, we assume that both players have a "safe second strike," meaning that if the other player attacks them they can safely retaliate, but not a "safe first strike" (meaning they can't strike first without being retaliated against). We'll also assume that, by striking first, you can diminish your opponent's second strike capacity, even if you can't fully eliminate it.[3]Under these assumptions, the payoffs look something like this:
In this case, both sides prefer to refrain if the other side will (because the -2 they'd get from being retaliated against is less than the 0 they'd get by refraining), but they also prefer to strike if the other will (because they have a chance of diminishing the other's second-strike capacity by striking first, meaning in expectation they get -4 instead of -5 by striking if the opponent does).
III. The AI Race
It turns out that, contra Hunter Ash, this structure is closely analogous to the AI race, under some plausible assumptions.
We can define R, S, T, and P for the AI race with a few parameters. Let's say that:
Given this:
Basically: if no one builds, nothing happens. If one side builds, that side has probability of aligning their AI; if they succeed, they get the Winner's payoff and the other side gets the Loser's payoff, if they fail, they get the Doom payoff. If both sides build, they each have a 50-50 shot of winning, and the winner has probability of aligning their AI and getting the Winner's payoff vs failing to align their AI and getting the Doom payoff.
Now we can just build a model and see that the game type depends on the payoffs and on :
So we can see that for low values of P(Doom), the game is indeed a Prisoner's Dilemma, but for higher values of P(Doom) it becomes a Stag Hunt.[7]There's a crossover point where the game type switches, and the value of only depends on the ratio between and . Specifically, the crossover point happens when . So we can write , and some algebra gets us . Since is negative, the meaning of this is more transparent if we write it as .[8]
So the worse you consider Doom to be (relative to the spoils of victory), the lower P(Doom) has to be for the game to be a Stag Hunt.
Because a Stag Hunt has an equilibrium at (Pause, Pause), an agreement not to build ASI is feasible as long as the game is a Stag Hunt. (Note: I mean game-theoretically feasible. Obviously there are other factors that impact the feasibility of an agreement—messy politics, personalities and relationship between the specific parties involved, etc. I'm not making any claims here about those.)[9]In a Stag Hunt, once both sides are sure the other will comply, a treaty doesn't need enforcement, strictly speaking: neither side has an incentive to defect, and both sides prefer the cooperative equilibrium. The hard part is verification and assurance.
This means that (contra Hunter Ash) you don't need P(Doom) "well north of 50%" for an agreement on not building ASI to be feasible (you only need P(Doom) to be this high if equals ). In fact, if you think Doom is at least 9x worse than Winning the race is good, then P(Doom) only needs to be 10% for an agreement to be feasible. This is a pretty conservative assumption I'd say, since Doom implies complete loss of human control of the future, and likely extinction. I'd go so far as to say it's plausible that Doom is at least 99x worse than Winning is good, in which case P(Doom) only needs to be 1% for an agreement to be feasible.
And most of those with decisionmaking authority here take Doom seriously enough that this math applies to them. Most Western frontier lab CEOs are on the record saying that AI takeover is a real concern; Dario Amodei and Elon Musk have given explicit probabilities of 10%+. Sam Altman and Demis Hassabis have both criticized the "P(Doom)" framing, but Sam has said it's "not zero" and Demis has said it's "definitely non-zero and it's probably non-negligible." And all except Elon (and Mark Zuckerberg) signed the 2023 CAIS statement saying that "Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war." Even Xi Jinping mentioned "loss of control" risks from AI in a recent speech, though, as ever when looking at China as a layperson, it's hard to know what to make of this. But in any case, by lab CEOs' own lights, cooperation with China should be (game-theoretically) possible.
In conclusion: STOP SAYING THE AI RACE IS A PRISONER'S DILEMMA![10]
Appendix: Expanded Model
I haven't had a chance to read all of these yet, but I have had Fable check how they relate to what I'm doing here, and although they are making a similar point, they use somewhat different modeling assumptions. ↩︎
In one-shot settings, and ignoring Functional Decision Theory. ↩︎
With perfectly secure second strikes, the game is actually a "Harmony" game: both sides prefer to Refrain unilaterally. But second strikes in the real world are never perfectly secure. ↩︎
Nothing much interesting happens if we relax this assumption. If Losing is worse than Winning is good, the game type doesn't change. If Losing is less bad than Winning is good, all that happens is that, for low enough values of P(Doom), the game becomes a "Deadlock", where both players actually prefer (Build, Build) to (Pause, Pause). You can see this in the expanded model in a later footnote. ↩︎
Again, nothing much interesting happens if we relax this assumption. If one side is favored to get to ASI first, then for low enough values of P(Doom), the game can be a Deadlock for them even if it isn't for the other side. You can see this in the expanded model in a later footnote. ↩︎
We'll assume that whichever ASI comes online first can defeat another ASI that comes online later, or prevent it from becoming ASI in the first place. So whether ASI is aligned or misaligned depends only on whether the winner's ASI is aligned or misaligned—once a misaligned ASI exists, there's no chance to defend against it with an aligned ASI of your own, and conversely, once an aligned ASI is created, all future ASIs will be aligned as well. ↩︎
For it actually becomes a Harmony game, but is not 1 so this is irrelevant for real-world analysis. ↩︎
Note that, in particular, it does not depend on L, or on the probability of winning the race given that both sides Build. So our two assumptions from earlier don't affect , and Ash's worry from the introduction about how bad it would be to lose the race never takes us from a Stag Hunt to a Prisoner's Dilemma. Here is an expanded model with these parameters allowed to vary that you can play with in the appendix at the end of the post. ↩︎
As an aside, this is basically isomorphic to another well-known variant of the nuclear arms race problem. There was a serious fear, at one point during the process of building the atomic bomb, that setting it off could ignite the atmosphere and kill everyone on earth. In actuality, before testing the bomb they did the calculations and found that the probability of this was small, but the same logic applies: if the probability of killing everyone were high enough, then coordination not to build the bomb would be a Stag Hunt. ↩︎
Unless your own P(Doom), W, and D satisfy (e.g. P(Doom) below 10% if you think Doom is only 9x as bad as Winning is good). ↩︎