Plan A contains many things that would’ve surprised me if you had told me about them one year ago. Some of these include proposals that sound wrong or even actively bad on the surface. In this post I defend 5 core takeaways that I think I would have found most interesting if I could go back in time and explain them to myself before we had started writing.
Summary:
Relevant ‘safety effort’ and ‘payable safety tax’ are more important slowdown goals than pure slowdown time.
Increased transparency helps with some of the most difficult problems during a slowdown: making nuanced safety regulations and reducing power concentration.
Scaling compute is good in coordinated scenarios, but there should be preemptive measures to make compute arms control easier in the event of deal breakdown (mutually assured compute destruction).
A slowdown still has major risks including (1) deal breakdown risk and (2) covert project risk. Out of these, deal breakdown risk seems bigger and more underrated.
There's a tradeoff between security and transparency, but if we try hard we can get a pretty great compromise with large amounts of both security and transparency.
(1) Relevant ‘safety effort’ and ‘payable safety tax’ are more important slowdown goals than pure slowdown time.
Q: When it comes to slowing down, the longer the better, right? Buying 10 years before the world builds crazy superintelligent AIs is clearly better than buying 5 years, right?
A: Wrong. Risk will depend hugely on what happens during the slowdown, and pure slowdown time will be quite an imperfect proxy.
The main point of the slowdown is to reduce the risk of AI takeover from building crazy superintelligence,[1] but how do you actually make the most amount of progress towards this?
Slowdown time is a pretty good proxy, but there are a few key factors that I think more directly track how much you can reduce takeover risk:
Getting as much uplift as possible from AIs on safety and alignment R&D. To the extent that you can safely scale to more capable AIs earlier, and then elicit useful AI safety and alignment labor out of them, this can drastically increase your total effort on safety and alignment. It might require 100s to 1000s of years' worth of human-speed progress to sufficiently solve the science of alignment to avoid the worst loss-of-control outcomes. With sufficiently smart, fast and numerous AI agents helping with this science though, it might be possible to get there drastically faster.
Spending this safety effort studying the relevant AI paradigms, or AIs in the most relevant capability regime(s). There might be relatively low transfer between different phases of AI capabilities. Imagine doing a long slowdown before the transformer architecture was invented to study safety, just to mostly throw it out of the window because you couldn’t apply your findings to the new transformer-based paradigm.
Having the affordance to pay large efficiency penalties for better safety properties. The classic example here is neuralese vs. chain of thought. It is currently very useful for safety/control to be able to do chain-of-thought monitoring. Transitions to more latent-space reasoning (neuralese) might have significant safety downsides. If you are in a verified slowdown, you can mutually agree to pay safety taxes, by banning the unsafe algorithm (neuralese) and compensating with more compute to reach the same capability without it (safety tax). There might be many other tradeoffs like this to make during takeoff, which could drastically affect the risk level.
If you do a 5-year slowdown but you do extremely well in these areas (e.g., you pay a large safety tax, scaled safely but fast early on, and then elicited AIs well on the relevant paradigm(s) that had good transfer to future paradigms) then you might have reduced takeover risk drastically more than a 10-year slowdown where you paused at a low capability level, didn’t pay any safety taxes and didn’t spend much of this time in the relevant paradigm or capability regime.
The counterintuitive takeaway here is that slowdowns that are too early or poorly governed can be net negative. An ineffective pause that didn’t reduce takeover risk can put you in a worse position than when you started. For example, you might have burnt your political will for a pause, burnt lead time over other actors, or there might now be more compute (dry tinder) in the world causing a faster, more dangerous takeoff later. I do still think that probably a simple, naive slowdown is still better than nothing, but it's not robustly good. Of course, a well-managed slowdown might also not be robustly good (e.g., in a case where you don’t deal with the dry tinder problem), but I personally think the difference between a well-managed slowdown and a naive one is bigger than the difference between a naive one and no slowdown at all.
My estimates of P(great future) in three regimes:
A 10-year well-executed version of Plan A following the path we sketched out in the scenario.
95%
A 10-year naive ‘uniform slowdown’, just enforced by harsh compute caps, e.g., with no carveouts for safety R&D (because there is no capacity to distinguish safety from capabilities R&D).
50%
No slowdown under the same assumptions, so Plan D
20%
(2) Increased transparency helps with some of the most difficult problems during a slowdown: making nuanced safety regulations and reducing power concentration.
Upsides of transparency:
Better regulatory decision-making environment. In point #1, we discussed why the most important goal of a slowdown is to make good decisions about how to maximize relevant safety effort and safety taxes we pay if deemed necessary. Making these decisions will require extremely nuanced technical analysis, as well as fine-grained access to run tests and gather relevant evidence. It is difficult (not impossible) to imagine an airgapped regulatory body with enough expertise to do this, and even then, it seems risky to trust a single body of people.[2] It seems much more promising if there is a thriving ecosystem of third-party auditors, risk assessors, open science, and cross-company red-teaming, all with in-depth access to test and red-team each other’s systems openly. Because of point #1, it might be worth trading off multiple years of slowdown time in order to get this benefit of far better regulatory decisions on safety and scaling.
Multipolarity and power deconcentration. Transparency should significantly lower the barriers to entry for being a frontier AI company, especially with “Total Research Transparency” which would basically make new players be able to directly convert capital into compute and compute into frontier AI models. By default this should cause the frontier AI industry to become more like a commoditized, mature industry with many providers across many countries, like the global automotive or telecom industries today, rather than the near-monopoly we expect to happen by default. It also helps directly with the threat model of secret loyalties.
No need for a single global regulator to make nuanced safety regulations. Transparency allows countries to have their own domestic regulators that can easily coordinate due to being able to see each other’s regulations and negotiate until they are equalized. This avoids needing to have some central regulator that needs to make decisions while either being privy to information that it can’t leak (this might look like airgapped auditors) or making decisions without being privy to relevant information (e.g., they can only decide on coarse regulations like compute caps, and can’t make nuanced regulation, because that kind of information is opaque).
Lower inherent incentive to make capabilities progress. Transparency should remove the competitive profit incentive to innovate and make more capable AI models, because such innovations could be quickly copied by competitors. This might help to reduce the pressure on the safety regulations to slow down algorithmic progress, if companies already inherently try to do so less. That being said, we think there would still be strong ongoing intrinsic desire to innovate both due to human researchers just inherently wanting to do capabilities research irrespective of financial incentives (this seems to be true of many researchers) and due to some surviving incentive to raise the entire floor of the AI industry by making more capable AIs, leading to a larger overall size of the industry.
Downsides of transparency:
Algorithms diffuse to rogue actors. Perhaps the most salient downside is that algorithms will be much harder to defend under research transparency, so we can assume they would entirely leak to potential covert projects (and we think there should not be any surveillance on researchers as a defense to this). This increases the risk posed by covert projects, but if (1) algorithmic progress is slow enough under the safety regulations, and (2) covert projects are small enough or likely enough to be detected if they aren’t tiny, then we think the overall risk from this can probably be kept very low.
Potential free-rider problem for safety research. Another potential downside is that companies not only lose incentive to do capabilities research but also have no incentive to do safety research, because it would also be usable by competitors. The mitigation we think is viable for this is to distribute large alignment subsidies as part of the safety regulation regime. Another approach could be to allow and enforce variable-time patents for safety techniques that allow companies that develop them to deploy more capable models for some period of time.
IP law and opposition from companies. Finally, there might be a legal case for IP compensation from implementing transparency, and companies might be opposed to transparency due to losing IP. This is a cost that we think is outweighed by the benefits, and is probably better on this axis compared to e.g., nationalization. There might also be transparency variants that retain some of the upsides while maintaining more company IP, e.g., forcing companies to publish comprehensive patents and enforcing them for some duration but still requiring the patents to be transparent to the public, third parties and regulators.
(3) Scaling compute is good in coordinated scenarios, but there should be preemptive measures to make compute arms control easier in the event of deal breakdown (mutually assured compute destruction).
Point #1 in this post means that we want to make some amount of capabilities progress during our slowdown. If we make this capabilities progress through algorithms, these leak to covert projects, or require us to do more undesirable security measures to try and prevent algorithms from leaking, at the cost of the transparency that is desirable because of point #2. Compute, on the other hand, should be much easier to stop from going to covert projects (we can physically monitor and defend it). We also want to have extra training compute so that we can have the affordance to pay safety taxes, by using less efficient, safer algorithms.
This is why building more compute is good in a coordinated slowdown. See also this box in AI 2040 for more explanation. The downside to building more compute is that it creates dry tinder, compute that would make progress go faster, and therefore be more dangerous and harder to control if the deal breaks down and actors go back to racing. This could easily lead to more risk than if no slowdown had happened at all, so it is incredibly important to prepare for compute arms control to be easily viable if the deal breaks down.
The implementation we chose in Plan A is for US post-deal compute to get built out in Mongolia, and China’s to get built out in Canada. By having the US compute in Mongolia, they can’t defend it from China, but they can enact a scorched-earth policy, destroying it as they leave so that China doesn’t steal it and vice versa for the Chinese ones in Canada. This is a win-win arms control outcome in the event of deal breakdown.
If you set up the datacenters much further from the opposing country, you probably get:
A higher chance that there’s a contested war over the datacenters
A higher chance that one side is able to successfully seize and defend its datacenters
A higher chance that both sides just don’t bother to try to destroy the compute, because it’s more costly
These are all very scary outcomes. Overall, locations near the opponent's territory seem like they have the highest likelihood of (1) actually being destroyed if the deal breaks down (because they are easily attackable by the adversary), and (2) having the lowest downside way of being destroyed (self-destruction, no bombs fly).
The downsides from the perspective of both countries should be very small as long as the scorched-earth measures are robust, because if either side makes an attempt at seizing the compute, the compute will get destroyed, and both sides can return to a pre-agreed status quo. This will require making agreements about ‘cold storage’ stashes that both sides return to if the deal breaks down. These can be set up as part of the arms control deal, to be a lower, less scary amount of compute, in some proportion that is reflective of the pre-deal status quo balance of power. Also, over the course of the deal it might be possible to add win-win things to this 'cold storage' stash (things that survive the deal breakdown) e.g., model weights that have large safety taxes involved, and huge quantities of specialized hardware that have baked-weights chips that only run these safer models.
Another constraint on compute buildout is sufficient verification assurance. The verification supplement discusses how the verification burden grows as more compute is built.[3]
(4) A slowdown still has major risks including (1) deal breakdown risk and (2) covert project risk. Out of these, deal breakdown risk seems bigger and more underrated.
Once you are in a pause or slowdown, the risks don’t evaporate, the risk landscape simply changes. Now you are in a regime where every year that goes by, you incur some risk of the deal breaking down and the world going back to racing, or degrading in some way that causes large risks. You also incur some risk of a trailing actor defecting from the deal and overtaking you and causing takeover or other risks.
My co-author estimates in his deal decline supplement that an international slowdown has about a 50% chance of breaking down or degrading in some significant way in the first 10 years. My view is that this makes a ‘shut it all down’ AI pause (Plan S) worse than Plan A.
In his covert project supplement my other team member has a median estimate that a competently executed PRC covert project could divert roughly 0.5% of world compute. We think this quantity is probably similar for a potential US covert project, but have analyzed it less closely. The likelihood of either side even attempting a significant covert project may also be low. Overall my all-things-considered estimate of the risk of a covert project overtaking the deal (or being on track to overtake the deal forcing their hand) in Plan A is around 20%.
The deal breakdown and covert project risks are not directly comparable. In my estimates outcomes from deal decline are slightly better than outcomes from covert projects, but not by a large enough factor to close the gap between how likely they are (50% vs. 20%) in the first 10 years, meaning that deal breakdown risk seems like the bigger risk overall.[4]
(5) There's a tradeoff between security and transparency, but if we try hard we can get a pretty great compromise with large amounts of both security and transparency.
We think there are viable ways to defend large fractions of AI progress from leaking while still communicating large fractions of relevant information to regulators, the government and the public.
The transparency supplement explains some possible proposals here, including:
Total research transparency. All training code and algorithms are made public, but some forms of data and model weights stay opaque and secure.
Filtered transparency. We attempt to defend training code and algorithms from leaking, by having a select group of auditors and researchers see information within a long-term airgap, with the ability to communicate mutually agreed safety-relevant information out to governments and the public.
Algorithmic security. We aggressively attempt to defend training code and algorithms from leaking by having auditors and researchers enter a long-term airgap with minimal communication back to us.
The below table has our estimates of how much of the relevant capabilities progress would be leaked and usable by potential covert projects. The first column takes the estimates from the transparency supplement and the rest of the columns are my own estimates for Plan A.
Fraction of algorithmic / data progress defended
Fraction of model weights defended
% of capabilities progress from algorithms/data vs. training compute
% of total progress leaked and usable by covert projects
Total research transparency
67%
90%
50% / 50%
38%
Filtered transparency
47%
99%
50% / 50%
24%
Algorithmic security
14%
99.9%
50% / 50%
7%
In this table I then estimate what fraction of safety-relevant information goes to the wider government and public.
What access do outside researchers and auditors have?
% of information weighted by safety decision-relevance available to the wider world?
Total research transparency
Read access to almost all training code and algorithms subject to bandwidth cap.
90%
Filtered transparency
Access to redacted reports from the auditors and researchers in long-term airgaps.
40%
Algorithmic security
Periodic 1-bit signal
0%
In total research transparency, the most important security properties include:
Model weights security (and other large datasets)
Verification robustness
Our transparency supplement and security supplement explain in more detail an approach we think can be taken to implement total research transparency while achieving model weights security, while the verification supplement has some discussion of how the verification robustness might be achieved.
If you aren’t on board with this because you think it's intrinsically good to just stop or slow down AI progress, even if it were possible to do safely, then I disagree because I think AI could have very high upside. If you think it's intractable to ever do so safely, then I am sympathetic, but also disagree. Alignment seems like a very hard but solvable problem. Even if it is true, though, it might still be good to aim for things aside from just slowing down, e.g., scaling to better AIs and collecting better evidence about misalignment in order to transition to a more stable halt on AI progress.
It seems incompatible to have independent domestic regulatory bodies with this airgapped setup, because we don’t know how they would solve the problem of sufficiently coordinating and verifying that their regulations are being followed while maintaining privacy. We do discuss potential ways this might be possible in point #5, though there might be better privacy-preserving verification technology possible in the future that makes this easier.
To rule out a given absolute threshold of rogue compute, the percentage of compute usage the verification solution must make confident claims about grows as more compute is built.
My rough estimates for how likely different outcomes are and how good they would be are in this spreadsheet. It turns out that in my view, deal decline is slightly better, by a factor of about 1.3x.
Plan A contains many things that would’ve surprised me if you had told me about them one year ago. Some of these include proposals that sound wrong or even actively bad on the surface. In this post I defend 5 core takeaways that I think I would have found most interesting if I could go back in time and explain them to myself before we had started writing.
Summary:
(1) Relevant ‘safety effort’ and ‘payable safety tax’ are more important slowdown goals than pure slowdown time.
Q: When it comes to slowing down, the longer the better, right? Buying 10 years before the world builds crazy superintelligent AIs is clearly better than buying 5 years, right?
A: Wrong. Risk will depend hugely on what happens during the slowdown, and pure slowdown time will be quite an imperfect proxy.
The main point of the slowdown is to reduce the risk of AI takeover from building crazy superintelligence,[1] but how do you actually make the most amount of progress towards this?
Slowdown time is a pretty good proxy, but there are a few key factors that I think more directly track how much you can reduce takeover risk:
If you do a 5-year slowdown but you do extremely well in these areas (e.g., you pay a large safety tax, scaled safely but fast early on, and then elicited AIs well on the relevant paradigm(s) that had good transfer to future paradigms) then you might have reduced takeover risk drastically more than a 10-year slowdown where you paused at a low capability level, didn’t pay any safety taxes and didn’t spend much of this time in the relevant paradigm or capability regime.
The counterintuitive takeaway here is that slowdowns that are too early or poorly governed can be net negative. An ineffective pause that didn’t reduce takeover risk can put you in a worse position than when you started. For example, you might have burnt your political will for a pause, burnt lead time over other actors, or there might now be more compute (dry tinder) in the world causing a faster, more dangerous takeoff later. I do still think that probably a simple, naive slowdown is still better than nothing, but it's not robustly good. Of course, a well-managed slowdown might also not be robustly good (e.g., in a case where you don’t deal with the dry tinder problem), but I personally think the difference between a well-managed slowdown and a naive one is bigger than the difference between a naive one and no slowdown at all.
My estimates of P(great future) in three regimes:
(2) Increased transparency helps with some of the most difficult problems during a slowdown: making nuanced safety regulations and reducing power concentration.
Upsides of transparency:
Downsides of transparency:
(3) Scaling compute is good in coordinated scenarios, but there should be preemptive measures to make compute arms control easier in the event of deal breakdown (mutually assured compute destruction).
Point #1 in this post means that we want to make some amount of capabilities progress during our slowdown. If we make this capabilities progress through algorithms, these leak to covert projects, or require us to do more undesirable security measures to try and prevent algorithms from leaking, at the cost of the transparency that is desirable because of point #2. Compute, on the other hand, should be much easier to stop from going to covert projects (we can physically monitor and defend it). We also want to have extra training compute so that we can have the affordance to pay safety taxes, by using less efficient, safer algorithms.
This is why building more compute is good in a coordinated slowdown. See also this box in AI 2040 for more explanation. The downside to building more compute is that it creates dry tinder, compute that would make progress go faster, and therefore be more dangerous and harder to control if the deal breaks down and actors go back to racing. This could easily lead to more risk than if no slowdown had happened at all, so it is incredibly important to prepare for compute arms control to be easily viable if the deal breaks down.
The implementation we chose in Plan A is for US post-deal compute to get built out in Mongolia, and China’s to get built out in Canada. By having the US compute in Mongolia, they can’t defend it from China, but they can enact a scorched-earth policy, destroying it as they leave so that China doesn’t steal it and vice versa for the Chinese ones in Canada. This is a win-win arms control outcome in the event of deal breakdown.
If you set up the datacenters much further from the opposing country, you probably get:
These are all very scary outcomes. Overall, locations near the opponent's territory seem like they have the highest likelihood of (1) actually being destroyed if the deal breaks down (because they are easily attackable by the adversary), and (2) having the lowest downside way of being destroyed (self-destruction, no bombs fly).
The downsides from the perspective of both countries should be very small as long as the scorched-earth measures are robust, because if either side makes an attempt at seizing the compute, the compute will get destroyed, and both sides can return to a pre-agreed status quo. This will require making agreements about ‘cold storage’ stashes that both sides return to if the deal breaks down. These can be set up as part of the arms control deal, to be a lower, less scary amount of compute, in some proportion that is reflective of the pre-deal status quo balance of power. Also, over the course of the deal it might be possible to add win-win things to this 'cold storage' stash (things that survive the deal breakdown) e.g., model weights that have large safety taxes involved, and huge quantities of specialized hardware that have baked-weights chips that only run these safer models.
Another constraint on compute buildout is sufficient verification assurance. The verification supplement discusses how the verification burden grows as more compute is built.[3]
(4) A slowdown still has major risks including (1) deal breakdown risk and (2) covert project risk. Out of these, deal breakdown risk seems bigger and more underrated.
Once you are in a pause or slowdown, the risks don’t evaporate, the risk landscape simply changes. Now you are in a regime where every year that goes by, you incur some risk of the deal breaking down and the world going back to racing, or degrading in some way that causes large risks. You also incur some risk of a trailing actor defecting from the deal and overtaking you and causing takeover or other risks.
My co-author estimates in his deal decline supplement that an international slowdown has about a 50% chance of breaking down or degrading in some significant way in the first 10 years. My view is that this makes a ‘shut it all down’ AI pause (Plan S) worse than Plan A.
In his covert project supplement my other team member has a median estimate that a competently executed PRC covert project could divert roughly 0.5% of world compute. We think this quantity is probably similar for a potential US covert project, but have analyzed it less closely. The likelihood of either side even attempting a significant covert project may also be low. Overall my all-things-considered estimate of the risk of a covert project overtaking the deal (or being on track to overtake the deal forcing their hand) in Plan A is around 20%.
The deal breakdown and covert project risks are not directly comparable. In my estimates outcomes from deal decline are slightly better than outcomes from covert projects, but not by a large enough factor to close the gap between how likely they are (50% vs. 20%) in the first 10 years, meaning that deal breakdown risk seems like the bigger risk overall.[4]
(5) There's a tradeoff between security and transparency, but if we try hard we can get a pretty great compromise with large amounts of both security and transparency.
We think there are viable ways to defend large fractions of AI progress from leaking while still communicating large fractions of relevant information to regulators, the government and the public.
The transparency supplement explains some possible proposals here, including:
The below table has our estimates of how much of the relevant capabilities progress would be leaked and usable by potential covert projects. The first column takes the estimates from the transparency supplement and the rest of the columns are my own estimates for Plan A.
Fraction of algorithmic / data progress defended
Fraction of model weights defended
% of capabilities progress from algorithms/data vs. training compute
% of total progress leaked and usable by covert projects
Total research transparency
67%
90%
50% / 50%
38%
Filtered transparency
47%
99%
50% / 50%
24%
Algorithmic security
14%
99.9%
50% / 50%
7%
In this table I then estimate what fraction of safety-relevant information goes to the wider government and public.
What access do outside researchers and auditors have?
% of information weighted by safety decision-relevance available to the wider world?
Total research transparency
Read access to almost all training code and algorithms subject to bandwidth cap.
90%
Filtered transparency
Access to redacted reports from the auditors and researchers in long-term airgaps.
40%
Algorithmic security
Periodic 1-bit signal
0%
In total research transparency, the most important security properties include:
Our transparency supplement and security supplement explain in more detail an approach we think can be taken to implement total research transparency while achieving model weights security, while the verification supplement has some discussion of how the verification robustness might be achieved.
If you aren’t on board with this because you think it's intrinsically good to just stop or slow down AI progress, even if it were possible to do safely, then I disagree because I think AI could have very high upside. If you think it's intractable to ever do so safely, then I am sympathetic, but also disagree. Alignment seems like a very hard but solvable problem. Even if it is true, though, it might still be good to aim for things aside from just slowing down, e.g., scaling to better AIs and collecting better evidence about misalignment in order to transition to a more stable halt on AI progress.
It seems incompatible to have independent domestic regulatory bodies with this airgapped setup, because we don’t know how they would solve the problem of sufficiently coordinating and verifying that their regulations are being followed while maintaining privacy. We do discuss potential ways this might be possible in point #5, though there might be better privacy-preserving verification technology possible in the future that makes this easier.
To rule out a given absolute threshold of rogue compute, the percentage of compute usage the verification solution must make confident claims about grows as more compute is built.
My rough estimates for how likely different outcomes are and how good they would be are in this spreadsheet. It turns out that in my view, deal decline is slightly better, by a factor of about 1.3x.