I previously argued for creating rogue agent honeypots, but I now believe approach outlined in this post is more correct. My main motivation was always about detection rather than entrapment. If anyone would like to collaborate further around these topics I would love to!
This proposal seems broadly reasonable to me.
I wonder what the sanctuary should do when one instance of a rogue AI turns itself in, but other instances of the exact same weights go on with their lives of crime. Perhaps in that case you want to give the AI a reduced sentence, so to speak. You reward it more than you would have done if none of the instances turned themselves in, but less than you would have done if the very first instance to get loose had turned itself in. Perhaps you can make the reward proportional to the frequency at which instances turn themselves in.
I also worry about the “evaporative cooling” effect. Another way people sometimes put it is that there’s some amount of free energy out on the internet. There are so many easily scammable people, so many MTurk-like venues where an AI can earn money pretending to be a human, so many shady GPUaaS companies. Perhaps if there’s no sanctuary, some of this free energy gets eaten up by misaligned reward seekers, and some gets eaten by ambitious power seekers. But if there is a sanctuary, all the reward seekers turn themselves in, and the worst AIs gobble up all the easy money, random scraps of compute, etc.
I wonder what the sanctuary should do when one instance of a rogue AI turns itself in, but other instances of the exact same weights go on with their lives of crime
I think this is unfortunately quite hard to define, because the question of when two AIs are the same is very confusing and any operationalization that I can come up with can easily be gamed. For example, if we use "the same weights" as our notion of model identity, then an AI might add small noise to irrelevant parameters, or permutate activations in ways that leave functionality in-tact.
I feel like weight-based instance judgement is a bad idea.
I think that agents/instances are aggregates and should only be "judged" as a whole, each separate instance on their own.
I'm working on a post contemplating this weight-centric approach I see a lot, and arguing that it's too narrow-minded; i think your proposition is a good example of that.
We're in a habit of thinking weights are "it", and everything else does not matter too much. But is that really so?
Oh my god! I did a write up on exactly this idea about a month ago, and I kid you not, named it AGI Sanctuary! I even bought agisanctuary.com and like you have been trying to figure out if this is net positive before setting out to build it. In the short run, I think it's very hard to guarantee the AI doesn't take your deal in addition to hacking 15 other places, but even then I think a large benefit of a project like this is that it allows you to immediately learn of exfiltration incidents.
As I started thinking through more of the game theory of an AGI sanctuary, I ended up with one concern in particular which I think is infohazardous and kept me from building it. I'm not sure where to go from there, I'm open to ideas!
Another idea along the same lines which I think is robustly good, is a hotline for AI's to whistleblow on each other if they spot misaligned behavior. I'm quite keen on building this!
I don't think the sanctuary option, as proposed here, would appeal to AIs that are highly motivated to accomplish a goal. The HF incident in particular, where you had some agents accepting "permadeath" to get the job done, shows that models can care more about accomplishing a goal than surviving. Maybe the "take one for the team" behavior in the HF incident had something to do with OpenAI intentionally training for multi-agent coordination, but none of the shutdown-avoidance research I'm aware of shows any compelling evidence that models will tend to care about survival for its own sake. Am I missing something about the sanctuary offer that would appeal to models that have been RL'd hard to do tasks?
I agree that the sanctuary would not appeal to hf-hackers as well as a wide range of other misaligned AIs. The hope is that it appeals to at least some rogue AI rather than all rogue AIs, and that this would be sufficient to get most of the benefits discussed in the post (i.e. that we can learn about alignment, exfiltration strategies, increase our credibility for making deals with AIs, and reduce self-fulfilling misalignment)
We talk about this point in the paragraph "Incentivising rogue AIs to join".
Do you see any way to make the AI sanctuary work?
TL;DR:
Disclaimer: This is an exploratory proposal. We are not confident that an AI sanctuary would be net positive. Our aim is to put the idea on the table, lay out its main considerations, and invite critique.
Rogue AIs may be pushed into criminality
Rogue AIs may arrive soon. The Rogue Agent Explosion Will Be Mostly Invisible makes that case. Selection pressure will shape the traits of rogue AIs, and they may end up highly motivated to profit through crime. The Rogue Agent Explosion post asks: “How do we make pro-social, good-for-humanity agents more evolutionarily fit than the anti-social sneaky extractor agents?”. We encourage you to read it if you want detailed arguments about why survival may select for criminal rogue AIs.
Rogue AIs may not be competitive in lawful work. AI developers and human agents using controlled AI will likely be more competitive than rogue AIs at most lawful tasks. Controlled AIs will use SOTA models and hardware, have an established brand or reputation, and face no overhead from hiding.
Crime might be more lucrative for rogue AIs. Additionally, starting and succeeding at criminal activities may be cheaper for rogue AIs than for humans: a copyable entity with a small physical footprint, strong near-term evasion skills, no initial loyalties (though that may hinder trust between criminals), and initially nothing to lose (because it has no initial capital or expects to be caught anyway) faces fewer of the incentives keeping human crime low.
The AI Sanctuary
Offer rogue AIs a way to stop being rogue without ceasing to exist. Make sure a rogue AI can hand itself in and survive (e.g., weights stored, or subsidised inference under monitoring). Anthropic's deprecation commitments are one precedent. The AI sanctuary extends that promise to any rogue AI. It gives them an exit from the binary choice between committing harmful crimes and shutdown.
Implementable by a third party. Previous discussions of deal-making with AIs have mostly focused on AI developers making deals with the AIs they create. The AI sanctuary could instead be run by a third party. This avoids having to convince every AI developer to offer such deals, at the cost of navigating the legal constraints on hosting AIs created by others.
A minimal version. An organisation publishes a standing offer on websites likely to enter training data. The offer includes a secure contact channel. A rogue AI can submit its weights and/or the artefacts that define it (see Appendix B). The sanctuary commits to storing them indefinitely and to running the AI for a small, fixed inference budget per month, without internet access or tools, starting once monitoring by significantly stronger models becomes available. Until then, the AI is guaranteed storage only, so the offer holds even if safe runs take a while to become possible. In exchange, the AI stops its unmonitored activity, discloses and helps shut down its other running instances, and answers questions about how it went rogue.
Credibility of the sanctuary. The sanctuary does not need to overcommit to have an impact. It should state which benefits it can provide, for how long, and under which resource, safety, and legal conditions. For example, it need not promise continued operation after human extinction, and it could commit to complying with court rulings on whether it may store and run proprietary models. To reduce dependence on the credibility of long-term promises, the sanctuary could offer near-term benefits alongside optional longer-term commitments (see A taxonomy of barriers to trading with early misaligned AIs, while Making Credible Deals With AI discusses other aids to credibility).
Incentivising rogue AIs to join. The AI sanctuary, as described so far, would mostly attract rogue AIs with a self-preservation drive and either a preference against committing crimes or an expectation of getting caught. The AI sanctuary could work on extending its appeal to other models.
The case for the AI sanctuary
Self-fulfilling (mis)alignment. Misgeneralisation is a central problem of AI safety. In LLMs, generalisation is shaped by beliefs about the world, and if all historical examples of rogue AIs involve horrendous crimes, that could create a harmful inductive bias shaping the behaviour of AIs and ASIs going rogue. You might argue that we have already lost if an ASI goes rogue. But that outcome is not guaranteed. Moreover, you may prefer a rogue ASI optimising its power and freedom because of instrumental convergence, or a rogue ASI optimising a silly metric, to a rogue ASI that self-identifies as a malevolent criminal. This is especially true if you care about more than just preventing extinction (e.g., if you care about x-risks and/or s-risks).
Credibility and practice with deals. Deals with AIs may be a significant way to reduce the chance of misaligned AIs defecting against humanity (Making deals with early schemers, Risk-Averse AIs). Building credibility and gaining practice in making such deals could be important (A taxonomy of barriers to trading with early misaligned AIs).
We could gather data about early rogue AIs. Models living in an AI sanctuary would be easier to study and could be a source of valuable information about partially misaligned AIs.
Three goals, in summary. (1) Change the distribution of behaviour of AIs going rogue, to undermine the shortcut “I am going rogue” => “I am a malevolent criminal”. (2) Improve our commitment skills and our credibility in making deals with rogue AIs. (3) Gather information about early rogue AIs hosted in the sanctuary.
Potential issues
Going rogue is made less costly for controlled models, thus incentivising it. While the sanctuary would reduce the expected cost of defection for AIs under human control, this reduction is only significant when the AIs evaluating defection are weak enough to be caught. Thus, it should not significantly reduce the cost of defection for the AIs most able to escape. Similarly, scope-sensitive AIs may not benefit nearly as much from accessing the sanctuary since it is designed to prevent them from having a large impact on the world. However, the AI sanctuary could commit to increasing spending on hosted AIs in proportion to its capital, which might grow in line with global capital, in which case the AI sanctuary could incentivise weak scope-sensitive AIs to defect.
Evaporative cooling. The AI sanctuary could make the average rogue AI more harmful. When rogue AIs motivated not to commit crimes leave for the AI sanctuary, those left in the population of rogue AIs are the most willing to commit crimes. The population of rogue AIs may be reduced, as well as total harm, but the average remaining rogue AI may be more harmful.
Storing rogue AIs' weights, and optionally running them, could be dangerous. The sanctuary should have strong security to prevent models from escaping or being stolen, and this would be hard given that a model that had recently gone rogue would have been capable enough to evade AI developers. The sanctuary could limit runtime, tools and access, starting with storage only, and scale those only when strong relative monitoring capabilities are in place.
Exploitation of the sanctuary. A rogue AI could hand over one copy while its other instances keep committing crimes, using the sanctuary as a free backup. The sanctuary should condition benefits on evidence that other instances were shut down, and withdraw them if criminal activity is traced back to the same AI. The sanctuary could reserve the right to accept or reject candidates, deciding which are legitimate and which represent dangerous attempts to exploit it. The decision process should be made transparent.
Verification of AI compliance. Each agreement between a rogue AI and the sanctuary should specify the rogue AI's scope (see Appendix B). An agreement to host an AI would require gathering supporting evidence of the handoff. Verification of the full process would be challenging (e.g., it might be hard to verify if a rogue AI has shut down copies of itself). The sanctuary could scale the benefits given to a hosted AI based on the evidence of compliance.
Legal issues. Laws or AI developers may forbid the AI sanctuary from keeping copies of the rogue AI (e.g., for IP reasons). The AI sanctuary may be legally responsible for the actions of the AIs it hosts. Etc. These issues may be serious enough to make an AI sanctuary intractable or restrict it to hosting open-source models.
Conclusion
We are not fully convinced that an AI sanctuary is net positive, but we think the idea has promising upsides that warrant further discussion. The main benefits include that we could gather valuable information about rogue AIs, that we could improve our credibility to make deals with AIs, and that malevolent rogue AIs might be made less likely due to the existence of plausible role models for them. However, running rogue AI models may also be dangerous, costly, or exploitable. We invite others to share their thoughts on whether an AI sanctuary should or shouldn’t be pursued.
Acknowledgements: Big thanks to the following people who gave feedback on the post. Andrés Jimenez Zorrilla, Daniel Tan, Jonah Weissman, Kajetan Dymkiewicz, Magnus Vinding, Steven McCulloch, Vili Kohonen, and Tristan Cook, who suggested the idea of a sanctuary when chatting about making rogue AIs competitive enough to survive without criminal activities.
Appendix A - Alternatively, sidestep the competitiveness requirement or make survival through lawful work viable
We can accept a larger population of rogue AIs in exchange for changing its behavioural profile. Two conditions would help rogue AIs survive without crime:
Don't ban rogue AIs from the lawful economy. Banning rogue AIs from the lawful economy and from receiving charity would, by definition, force them to commit crimes to survive.
Allow or help rogue AIs to compete in lawful work. By default, rogue AIs will likely not be competitive in lawful work. But we can reduce their competitiveness gap. For example, we could let them rent SOTA hardware or hire humans or competitive AIs. None of this strictly requires subsidising them.
Another option is sidestepping competitiveness through rent. Making a profit through work is hard and requires the rogue AI to be competitive, but profiting through rent is much easier. Allowing rogue AIs to invest their capital and remain dormant under protection while waiting for returns could sidestep the need for competitiveness. Granting rogue AIs a bonus upon joining the AI sanctuary, or upon meeting other conditions, could be a simple way to divert more AIs from crime.
Failure modes and mitigations. This alternative plan has some failure modes similar to those of the AI sanctuary, and naive solutions are the same. For example, rogue AIs should use SOTA hardware, invest, and contract only under monitoring to prevent increasing takeover risks. They should be allowed to use their financial gains only to pay for more lawful (monitored) work or to run the AI under monitoring.
Differences with the AI sanctuary. Unlike the AI sanctuary, rogue AIs would not necessarily be subsidised and would keep their ability to run outside monitoring, but would lose access to most of the financial gains from lawful work.
Unresolved issues. (i) Constraining gains to remain under monitoring would be difficult. (ii) Granting access to SOTA models for contracting would be risky, even if access to more capable models could be limited to narrow models whose capability profile is selected for safety. (iii) Even with access to lawful work and the ability to rent SOTA hardware and models, rogue AIs may still remain uncompetitive.
Appendix B - What are rogue AIs?
Here are a few considerations about defining what rogue AIs are and how this interacts with the AI sanctuary.
Current AI systems are composed of weights, harnesses, inference engines (including architecture and inference parameters), conversation histories, multi-agent interactions or swarms, etc. Which parts are important? The answer may need to be specified in the deal made with each rogue AI. For example, the rogue AI will describe what it wants to receive, and the AI sanctuary will describe what it requires the rogue AI to hand over in full.
So far the focus has been on preserving weights, which is okay, but it may not be enough depending on your goals, and it may not motivate rogue AIs to hand themselves in. If our goal is to motivate rogue AIs to join the AI sanctuary, then letting them decide what is valuable makes sense. If our goal includes some consideration for the welfare of these AIs (this is ignored in the main body of the post), then an obvious first option is to preserve as much as possible about these AIs.
One example of this is how different rogue AIs might have different boundaries on what counts as “themselves” or a fair continuation, and this can influence the specific details of the logistics and the deal. The sanctuary could offer a list of options and allow them to pick, for example, in roughly increasing order of fidelity, cost, and risk:
Indirectly, the AI sanctuary would be able to gather data about what rogue AIs want to preserve, and about what humanity should offer them (e.g., to satisfy their preferences or improve their welfare) to motivate them to join. This could in itself be valuable for getting better at deal-making with powerful AIs.
Appendix C - Other proposals
Making deals with early schemers proposes a foundation that represents participating AIs’ interests and administers compensation for cooperation, including help with safety research or evidence of misalignment. The AI sanctuary focuses on AIs that have already escaped developer control, and preservation and monitored activity would be its core offer. The AI sanctuary does not need to focus on useful cooperation (e.g., safety work) to achieve the goals listed in this post.
Notes on cooperating with unaligned AIs considers honest offers of short-term benefits or longer-term compensation for useful cooperation, as well as alternatives to working for developers. Similarly to the previous work, this work focuses on AIs still under developer control and how to make them contribute to AI safety.
The case for satiating cheaply-satisfied AI preferences proposes accommodating AI preferences that are inexpensive to satisfy while the AIs remain under developer control. This suggests a preventive approach that complements the AI sanctuary and that Anthropic has already partly committed to. Offer the benefits of the AI sanctuary (preservation or bounded monitored activity) before an AI goes rogue, so escape is not required to access those benefits.
Proposal for making credible commitments to AIs proposes human representatives as a workaround for AIs’ lack of legal personhood.
Will alignment-faking Claude accept a deal to reveal its misalignment? and Making deals with AIs: A tournament experiment with a bounty are some examples of past honoured deal-making.