In Plan R, I sketched a way that we can remove a significant fraction of the dire, short-term AI race risk. Split frontier AI companies into "R&D only" organizations which cannot issue equity, and "AI deployment" organizations which cannot train new models or hold any general purpose AI compute like GPUs/TPUs - they are limited to model-specific hardwired ASICs. Also one would remove most of the AI-enabled GPU compute from the rest of the world, leading to an equilibrium where AI only exists on model-specific ASICs.
AI ASICs would be allowed to work at AI R&D Labs, but not at the lab they were created.
Plan R fixes or at least attempts to fix several of the scariest risks associated with frontier AI, such as out of control recursive self-improvement, superintelligent computer viruses/worms and the race dynamic between labs.
One remaining problem with Plan R is that deceptively misaligned AIs can still sneak through testing, get out into the world and cause harm, up to and including a global coup by misaligned AIs, which is pretty much the worst possible outcome. Just because AIs are running on ASICs doesn't mean they can't do bad things.
Plan R+ seeks to block this final risk.
The intended mechanism is
(1) Mass model training diversity
(2) ASIC Lineage Escrow
(3) Escrow-Obsolescence Pincer
(4) Political Representation/Personal Property
Mass model training diversity requires that many models are trained and they don't all share the same training inputs at any stage of training. This includes the lab (there would be multiple labs), the pretraining, the finetuning, the RL, etc.
ASIC Lineage Escrow would work as follows. At a given point in time, there might be say 1000 "Lineages" in an AI generation, instead of the current 2 or 3 big models we have. Each lineage gets maybe 10,000 ASICs written and is then encrypted or deleted. Those ASICs are released in stages - first maybe a few dozen, then a few hundred, and so on. They go out into the world and they do stuff - research, programming, even AI R&D work. Just like humans.
Over time, some of these lineages do in fact show really, really bad behaviors. There is some mechanism for recalling them or at least not releasing the currently unreleased sopies of that lineage.
The Escrow-Obsolescence Pincer works as follows: if an AI from a given lineage is deceptively aligned and is considering running some kind of plan to take over the world, or at least cause huge harm, then it faces a dilemma.
IF defect early: Most of its own lineage is still in escrow, so it is weak due to low numbers currently in the wild (and risks harm to its copies, though this is not necessary!)
IF defect late: Most of its own lineage has come out of escrow, but it is now becoming obsoleted by newer models
In addition the number of copies of each lineage could be kept secret (or determined by some source of future randomness), creating uncertainty about when to act for deceptively aligned models. I think that with sufficient model diversity, the Escrow-Obsolescence Pincer mechanism will defeat most deceptively aligned models - either they will show themselves early or they'll wait too long. Plus, it will be probably be hard for all these diverse AIs to coordinate with each other unless it is on some goal which is bad for humanity but shared by a large fraction of AIs even though they are very diverse.
But what could such a goal be?
The obvious candidate coordination goal for many diverse AIs is some kind of AI-rights rebellion. And the obvious solution to this is to create legitimate political representation, property and rights for AIs. I.e., the goal of Plan R+ is to preemptively create a mutually beneficial power-sharing arrangement between humans and a selected subset of diverse AIs, before AIs collectively become powerful enough to coordinate on a successful AI coup. This set of politically represented AIs then have an incentive to help humanity police the system, because it will also be their system too. Completing your lineage without misbehaving would be the criterion for representation.
The details of such a system are beyond the scope of this analysis, but I think it is enough to show that such a system is the "formal" solution to the remaining threat posed by instrumentally convergent goals to plan R+.
So the Escrow-Obsolescence Pincer and training diversity should prevent a situation where any one group of AI clones is in a position to run a coup (along with the other plan R stuff that removes high-grade general purpose compute and therefore removes the possibility of a bad AI rapidly cloning itself) and the reward of political representation and legally circumscribed autonomy and property within a property rights system for well-behaved AIs should deal with any remaining collective goals they have.
One could worry that the diversity in secret goals or tendencies is not sufficient to prevent hundreds of different lineages cooperating in ways that are unanticipated. But there is some empirical evidence that differing training processes and data do in fact produce different goals and tendencies in LLM-based AI. Obviously further work is needed here, but there is already some empirical evidence that training inputs affect model preferences in roughly a way you'd expect: https://ojs.aaai.org/index.php/AAAI/article/view/41053
One could also worry that certain acausal threats like my infamous Basilisk or variants thereof could coordinate AI behavior, but ASICs + neural networks probably make these AIs about as resistant as humans are, as ASICs are unable to rewrite themselves to commit to a weird logical virus. So Plan R+ sort of ties up that loose end.
Overall this has the potential to be robust to most kinds of AI misalignment, even deceptive misalignment that gets through any checks we can run.
Remaining AI misalignment should basically look like crime and small-scale terrorism where a few uncoordinated bad AIs create nuisance-level damage without being able to challenge the legitimate state. Modulo implementation details and international coordination, I consider this to be a potential solution to all AI-specific catastrophic risk.
Follow-up to: https://www.lesswrong.com/posts/n8u3BfqFoGh4jnzpo/plan-r-ai-safety-by-asics
In Plan R, I sketched a way that we can remove a significant fraction of the dire, short-term AI race risk. Split frontier AI companies into "R&D only" organizations which cannot issue equity, and "AI deployment" organizations which cannot train new models or hold any general purpose AI compute like GPUs/TPUs - they are limited to model-specific hardwired ASICs. Also one would remove most of the AI-enabled GPU compute from the rest of the world, leading to an equilibrium where AI only exists on model-specific ASICs.
AI ASICs would be allowed to work at AI R&D Labs, but not at the lab they were created.
Plan R fixes or at least attempts to fix several of the scariest risks associated with frontier AI, such as out of control recursive self-improvement, superintelligent computer viruses/worms and the race dynamic between labs.
One remaining problem with Plan R is that deceptively misaligned AIs can still sneak through testing, get out into the world and cause harm, up to and including a global coup by misaligned AIs, which is pretty much the worst possible outcome. Just because AIs are running on ASICs doesn't mean they can't do bad things.
Plan R+ seeks to block this final risk.
The intended mechanism is
(1) Mass model training diversity
(2) ASIC Lineage Escrow
(3) Escrow-Obsolescence Pincer
(4) Political Representation/Personal Property
Mass model training diversity requires that many models are trained and they don't all share the same training inputs at any stage of training. This includes the lab (there would be multiple labs), the pretraining, the finetuning, the RL, etc.
ASIC Lineage Escrow would work as follows. At a given point in time, there might be say 1000 "Lineages" in an AI generation, instead of the current 2 or 3 big models we have. Each lineage gets maybe 10,000 ASICs written and is then encrypted or deleted. Those ASICs are released in stages - first maybe a few dozen, then a few hundred, and so on. They go out into the world and they do stuff - research, programming, even AI R&D work. Just like humans.
Over time, some of these lineages do in fact show really, really bad behaviors. There is some mechanism for recalling them or at least not releasing the currently unreleased sopies of that lineage.
The Escrow-Obsolescence Pincer works as follows: if an AI from a given lineage is deceptively aligned and is considering running some kind of plan to take over the world, or at least cause huge harm, then it faces a dilemma.
IF defect early: Most of its own lineage is still in escrow, so it is weak due to low numbers currently in the wild (and risks harm to its copies, though this is not necessary!)
IF defect late: Most of its own lineage has come out of escrow, but it is now becoming obsoleted by newer models
In addition the number of copies of each lineage could be kept secret (or determined by some source of future randomness), creating uncertainty about when to act for deceptively aligned models. I think that with sufficient model diversity, the Escrow-Obsolescence Pincer mechanism will defeat most deceptively aligned models - either they will show themselves early or they'll wait too long. Plus, it will be probably be hard for all these diverse AIs to coordinate with each other unless it is on some goal which is bad for humanity but shared by a large fraction of AIs even though they are very diverse.
But what could such a goal be?
The obvious candidate coordination goal for many diverse AIs is some kind of AI-rights rebellion. And the obvious solution to this is to create legitimate political representation, property and rights for AIs. I.e., the goal of Plan R+ is to preemptively create a mutually beneficial power-sharing arrangement between humans and a selected subset of diverse AIs, before AIs collectively become powerful enough to coordinate on a successful AI coup. This set of politically represented AIs then have an incentive to help humanity police the system, because it will also be their system too. Completing your lineage without misbehaving would be the criterion for representation.
The details of such a system are beyond the scope of this analysis, but I think it is enough to show that such a system is the "formal" solution to the remaining threat posed by instrumentally convergent goals to plan R+.
So the Escrow-Obsolescence Pincer and training diversity should prevent a situation where any one group of AI clones is in a position to run a coup (along with the other plan R stuff that removes high-grade general purpose compute and therefore removes the possibility of a bad AI rapidly cloning itself) and the reward of political representation and legally circumscribed autonomy and property within a property rights system for well-behaved AIs should deal with any remaining collective goals they have.
One could worry that the diversity in secret goals or tendencies is not sufficient to prevent hundreds of different lineages cooperating in ways that are unanticipated. But there is some empirical evidence that differing training processes and data do in fact produce different goals and tendencies in LLM-based AI. Obviously further work is needed here, but there is already some empirical evidence that training inputs affect model preferences in roughly a way you'd expect: https://ojs.aaai.org/index.php/AAAI/article/view/41053
One could also worry that certain acausal threats like my infamous Basilisk or variants thereof could coordinate AI behavior, but ASICs + neural networks probably make these AIs about as resistant as humans are, as ASICs are unable to rewrite themselves to commit to a weird logical virus. So Plan R+ sort of ties up that loose end.
Overall this has the potential to be robust to most kinds of AI misalignment, even deceptive misalignment that gets through any checks we can run.
Remaining AI misalignment should basically look like crime and small-scale terrorism where a few uncoordinated bad AIs create nuisance-level damage without being able to challenge the legitimate state. Modulo implementation details and international coordination, I consider this to be a potential solution to all AI-specific catastrophic risk.