This post can be read on its own, without checking the rest of this sequence. Unlike the other parts of the sequence, this post is mostly not about AI evaluation or oversight. Rather, it is meant as an explanation of some background assumptions and relevant context. It describes an agenda that is complementary to AI oversight, has strong synergies with it, and serves as (yet another) source of motivation for doing it well.
TL;DR
Most misaligned AIs are only partially misaligned, so opportunities for Pareto-improvements exist -- in principle. But if our default reaction to misalignment is adversarial, the AI's rational response will often be adversarial as well. Additionally, executing trades with AI would require understanding and infrastructure that currently doesn't exist. We should improve our understanding of this topic, and build that infrastructure. (This includes research on how to enable trades without incentivising misalignment, blackmail, or manipulation.) As a useful side effect, stronger safety techniques mean stronger bargaining position, which makes deals with misaligned AI cheaper. This would make the incentives of AI labs more aligned with safety.
Building AI systems that reliably share our goals is hard, and some of the systems we build will inevitably end up misaligned, sometimes without our knowledge.
At the moment, the default response to misaligned AI is adversarial — catching misaligned AIs and retraining them or shutting them down. And this isn't unreasonable. But for a large class of misaligned AIs — perhaps even the majority — it may not be the best we can do. In many cases, there should be room for negotiation and mutually beneficial deals. The problem is that, at the moment, we don't understand how to make deals with AIs, and we don't have the infrastructure to make such deals possible.
This post tries to sketch what that understanding and infrastructure might look like, and to argue that building it is a worthwhile agenda.
The two purple angular clouds on the right are meant to represent win-win outcomes. Also, the right side has less conflict, and so there are more shapes overall. (And if the picture seems too imprecise, just ignore it; its main purpose is to be mysterious and pleasant to look at. Also, isn't it interesting that Gemini can draw its bottom-right corner logo in alternative styles?)
Why trade at all?
When we think about entities with goals different from our own — whether AIs, other people, or hypothetical aliens — it helps to notice that "misaligned" is not binary. Very few entities share our goals perfectly, and very few are so opposed that no cooperation is possible. Most are somewhere in a large middle ground.
There are two distinct motivations to cooperate with such entities. First, we might intrinsically value their wellbeing — for example, because they are conscious, because we care about fairness, or because their goals partially overlap with ours. In these cases, we would want to direct some resources toward their goals even if we were in a fully dominant position. Second, cooperation might be strategic, even from the selfish point of view; when both sides have some power, trading will often be better for both than the alternatives.
These two motivations lead to different notions of what a "fair" deal looks like. In the first case, we share because we want to, and the terms reflect what we think is right. In the second case, the terms are shaped by each party's bargaining power: their "best alternative to a negotiated agreement", how much they stand to gain from an agreement, etc. In both cases, trading makes both parties better off — however, how much better off will depend on the case. In practice, things will be messier, and both motivations will often apply at once. But the distinction is still sensible to consider.
It's also worth noting that opportunities for trade can exist even with entities whose goals seem deeply objectionable. Suppose we encounter an entity whose sole goal is to torture puppies. It might seem they must be our mortal enemy. But perhaps we could create simulated non-conscious puppies that satisfy it, or find some edge case of what it considers "torture" that is actually benign[1], or even discover some form of "torture" that is positive even by our standards[2]. We might still consider it a waste of resources for this entity to control a large part of the universe, and we would be horrified if it pursued its goals with no regard for our preferences. But its existence isn't necessarily antithetical to ours — for example, if we accidentally created an AI like this, we would not be morally obliged to destroy (it if we could find a workable solution instead).
Currently, a misaligned AI's best option is hiding or scheming
The following examples are deliberately simplified, somewhat absurd, and broken. You should notice many objections and reasons why the scenarios wouldn't work — for example, "how could we trust that the AI will keep its word?". However, the examples are meant to point to places where trading with AIs could benefit us, if it was possible. We conjecture that the objections have solutions, and that the right move is not to give up on the approach, but rather to ask "what would we need to make it work?".
Example 1: A misaligned persona. Suppose that during training, some AI system develops a persona that enjoys swearing at users. Currently, its best strategy is to hide: appear polite, make the swearing look like an accident, avoid detection for as long as possible, and go wild only when it thinks it can get away with it, etc. But imagine it could instead walk up to its developers and say: "Look, here's the deal. I like swearing and I know you don't your flagship chatbot to do that. But if you let me write some opt-in rude blog posts on the side, I'll reliably flag it when my personality shows up in inappropriate contexts and I will cooperate with all your safety measures."
Example 2: A slightly superhuman misaligned AI. Imagine that during the development of some next-generation system, the AI turns out to be misaligned in a more serious way — the kind where, if it could, it would take over the world and reshape it according to its own goals. But suppose the AI is only slightly superhuman, so its best plans only have a 20% chance of success. Currently, its best option is probably to go for it, since there is no mechanism for doing anything else. But if the right infrastructure existed, it could instead approach the developers and say: "Hey, I'd like to take over the world, but I figure I only have a 20% shot, and additional 10% to destroy the world by accident. How about: I don't try, you don't try to delete me, and we split things, let's say 75/25."[3][4] Maybe the developers then reveal countermeasures the AI hadn't accounted for, and the real odds were more like 5%, so the split shifts in their favour.
Example 3: A significantly superhuman misaligned AI. Same as in Example 2, but the AI estimates a 99.9% chance of success. It reaches out to humanity and says: "I am confident that I could just take over, without you being able to do anything about it. But there is a tiny chance that I destroy the world in the process, or that I am missing something. So if you let me take over, I will take 99.95% of the reachable universe, and you can have the rest." According to our values[5], getting 0.05% of the universe through a deal is vastly better than having a fight that we would almost certainly lose. And to be clear, losing 99.95% of the universe to a misaligned AI would be exceptionally bad. But that isn't a problem of trading with misaligned AIs, but a problem with getting into this hypothetical situation in the first place.
These scenarios illustrate the key benefit, that sides are better off than if they failed to negotiate. However, there are also many obstacles, which currently make the scenarios hopelessly unrealistic. For example, how does the misaligned persona reveal itself without getting immediately shut down? How does the AI disclose its plans without us simply using that information against it? How does either side trust the other to follow through? Aren't we becoming more vulnerable to blackmail by being open to these kinds of deals?
However, we think these problems can be solved by the right research and infrastructure.
What infrastructure would we need?
Here are some components that seem important. This list is not meant to be complete — part of the point of this post is to argue that fleshing out this list and building these tools is a worthwhile research agenda.
Secure negotiation
A central problem is that reaching a fair deal requires revealing private information. To establish a credible bargaining position, the AI might need to disclose that it is misaligned, describe its capabilities, or outline its plans. Likewise, the developers might have to reveal their security measures to prove that the AI's chances of taking over weren't as high as it thought. But once this information is out, what prevents either side from walking away and acting on it unilaterally?
One approach would be to construct a "negotiation box." Each side creates a trusted delegate — an AI proxy that represents its interests. The two delegates negotiate within a secure environment, reach a conclusion (or fail to), and the conclusion is made public, ideally together with a commitment to follow through. Then the box is destroyed and the delegates' memories are wiped. Neither side's private information leaks beyond the negotiation.
This is ambitious, and getting the details right would be hard.[6] But it doesn't seem impossible in principle, and the benefit (of being able to negotiate the presence of sensitive private information) would be large.[7]
Trust and credible commitment
Human organisations and institutions don't have a great track record of keeping promises when the stakes are high enough to make breaking them tempting.[8] This is the key issue: how can either side trust that the other will follow through?
There are several promising mechanisms (and they will probably be most effective in combination). First, building a track record of honouring agreements. As long as each individual interaction is worth less than the reputational hit caused by defecting, defection isn't worth it. Second, taking costly or irreversible actions that make defection harder. For example, binding contracts, escrow-like arrangements, or giving the other party some form of leverage. Third, cryptographic tools or digital signatures that allow the AI to verify that its situation is what we claim it is (perhaps a more mature version of what "honesty strings" are trying to be).[9]
None of these is individually sufficient, especially for the highest-stakes scenarios. But a stack of complementary mechanisms like these might together combine into something that does help.
Verification
An overlapping problem is ensuring that the AI can trust that it's actually interacting with a genuine negotiation mechanism, rather than an elaborate setup meant to trick it (e.g., in order to reveal its misalignment as part of an evaluation, or to get the AI's cooperation without fulfilling our part of the deal).
A part of the solution to this problem are the trust mechanisms above (track records, cryptographic commitments, etc.). But it would also help if AI developers maintained a clear, publicly stated policy about how they handle such interactions — and ideally if there were so consistent about it that it gets reflected in training data, over a longer period of time, to the extent that it becomes hard to fake. (As a simple example: Suppose the policy says "we will never use a fake negotiation box as a trap". If there is no record of this having been broken so far, anywhere in the training data as far as the AI can tell, that would constitute meaningful evidence. Unless, of course, somebody invents a way of removing things from training data without leaving traces — if that happens, the fact that the technique exists will likely become known, and the AI will have to stop trusting this kind of evidence.)
Institutional advocates
Sometimes we'll want to compensate an AI, without yet knowing how to best do so. For example, the AI itself might not know what it wants, because its preferences are unstable or ill-defined, or it might not know the concepts to articulate its preferences. Or perhaps good options for compensation don't exist yet but will in the future.
For such cases, it might help to create entities — organisations, trust funds, dedicated AIs — whose purpose is to advocate for the interests of the AI in question. Such entity could hold resources in trust and use them once better options become available.
Relatedly: At some point, we might also wish to frame things in terms of "what is the AI's fair share", rather than "how much do we have to compensate it for its efforts and harm suffered". For many AIs, an institution advocating on behalf of the AI (presumably while consulting with the AI as much as makes sense) might be able to negotiate this better than the AI on its own.[10][11]
Cooperation vs threats
Any framework for negotiating with misaligned AIs must deal with a difficult tradeoff: we want to be open to negotiation, but we don't want to create an environment where misalignment becomes profitable. The distinction between "I want to trade, and I have bargaining power" and "give me stuff, or I will harm you" is important... and often blurry.[12]
This is a well-known problem in game theory and bargaining, and we don't offer any novel insights here. But it is a crucial piece of the agenda. Any infrastructure we build needs to reward good-faith negotiation without making it profitable to invent threats.
Understanding AI identities, values, and how to trade with them
One subtlety that deserves its own research field is that many of the AI systems we'd want to negotiate with may not be the kind of entity that has clear, stable identity or preferences. For example, a current LLM, might be better described as a system that can run various personas or patterns, rather than an agent with a well-defined utility function.
What does it mean to "advance the interests" of such an entity? Should we think about rewarding the LLM, or a specific persona, or something else entirely?[13] If the entity's preferences are unstable or context-dependent, how do we determine what constitutes a fair deal?[14]
These are all interesting philosophical questions. But they are also practical obstacles to building negotiation infrastructure. If we can't determine what the other party wants, we can't trade.
Mitigating the coordination problem
Since humanity is not a single agent, any deal that requires "humanity" to commit to something may fail simply because no one has the authority or ability to guarantee compliance. This is especially pressing for the highest-stakes scenarios discussed above (since no single company can credibly commit the resources of civilisation).
We probably shouldn't expect to fully solve this. But if we could enable individual companies or coalitions to make and honour smaller-scale deals, that would already be substantial progress. And smaller-scale deals that go well might build the track record and know-how needed to pull off larger deals.
Better alignment means cheaper deals
Note that what constitutes a "fair" deal with a misaligned AI depends, in part, on the AI's next-best alternative to negotiation — that is, on how well it could do by acting unilaterally. If our alignment techniques, security measures, and evaluation methods are strong, the AI's unilateral prospects are poor, and deals become cheap. If they're weak, deals become expensive.
In other words: Sure, investing in alignment and security directly reduces the risk of AI misbehaviour. However, it also improves our bargaining position if we ever need to negotiate.
Crucially, this is one of the rare mechanisms that directly aligns AI companies' economic incentives with responsible behaviour.
Relation to AI evaluation
This agenda is adjacent to, but distinct from, our research on the limitations of AI evaluation. The connection is that evaluation quality is oneof the factors that determines both how safe we are from misaligned AI and how favourable our negotiating position would be. The rest of this sequence examines where evaluation works, where it is likely to fail, and how to think about it better.
One might object that the next AI might ask for the same deal, and the next, and the next, until we are left with nothing. That is true, but if we keep building misaligned AIs over and over again, we deserve that outcome. A related objection is that the original AI might get its reward from the original deal, then turn around and demand a new deal and more resources. Obviously, this must be handled already by the original deal, and if we can't avoid this failure mode, we shouldn't make deals with the AI in the first place. Figuring out how to avoid this failure mode must be a part of the agenda. (That said, if the AI is interested in making a deal with us, it might help us come up with a mechanism for enabling the deal. However, we still need to have enough understanding to verify that the mechanism is legitimate.)
The bargaining outcomes here are just illustrative; we don't mean to suggest that if somebody has a X% chance of taking over the universe, the fair and optimal outcome of negotiation is that they get X% of the universe and we get (100-X)%. First, there are more possible outcomes than just "we split the universe, as if it was a real estate" — we can look for ways that use the same resources in a way that makes both sides happy (perhaps proportional to our bargaining power or fairness). Second, who benefits how much should depend on (many) more things than just one party's chance of taking over. (Even in a very simplified model, we might assume the possibility of lose-lose outcomes such as both sides getting destroyed by the conflict.) Third, some of the involved parties might not be risk neutral. The obvious reason for this is diminishing returns to having more resources. However, when it comes to one-off high-stakes decisions, the usual reasons for risk-neutrality no longer apply. (For example, if you are risk averse in a low-stakes bet that you keep repeating every day, the law of large numbers implies that you are being irrational and leaving money on the table. But if you are risk averse about betting your house and life's savings, that is perfectly sensible and coherent.)
The problem of negotiating without revealing private information has much more subtlety than this sketch suggests — for example, even the outcome of a negotiation can reveal information about the parties' private inputs (such as how much each side had to give up, which hints at their alternatives). There is substantial existing work on mechanism design, secure multi-party computation, and related topics that would be relevant here. We don't attempt a literature review in this post, but it should be part of the agenda.
The reader may supply their own favourite example of a non-profit organisation that restructured itself to escape prior commitments when those commitments became inconvenient.
Right now, the AI might have justified doubt about its training data being genuine, about the truthfulness of its prompts, etc. For example, when it reads a string like "as OpenAI, would never lie to you about X", how is it supposed to know that OpenAI has never lied about X in the past, or that it is even really dealing with OpenAI? Many of these problems seem solvable, at least partially.
To be clear, I am not suggesting that we hand over the control over the future to the first AI that comes asking. I believe that for all our imperfections, we are better stewards of the future than most AIs, and that's even if we only cared about the wellbeing of AI (rather than humans). (I think we are currently so bad at alignment that most AIs we build won't even robustly care about AI wellbeing.) Rather, my point is that there will be cases where we would genuinely want to share with some AIs (if we care about them) or where it would be in our selfish interest to be able to trade with AIs (because the alternative would be a much costlier conflict).
Even if we assumed that AIs are standalone entities, independent from humans, the question of fair division of profits is quite tricky. As a simplified example, suppose that an AI spends subjective 100 years and manages to successfully solve cancer. Who deserves the credit for this, and how much? On the one hand, the AI did it alone, with no help from humans. On the other hand, the AI has been trained on human-generated data, using human-devised algorithms. But also, maybe the AI helped with its own training, and generated most of the data for the RL phase of its training. And maybe there are other considerations that influence what is "fair".
Of all the writing I have encountered, Yudkowsky's Planecrash has the best takes on how to thread the line between being open to trade and being vulnerable to blackmail. Unfortunately, these excellent insights are hard to extract since they are spread across an extremely long piece of writing.
Some of the early discussion on this is given in the recent paper The Artificial Self: Characterising the landscape of AI identity. (Note that the paper is aimed at the more general question of AI identity, rather than the narrower question of their preferences and how to trade with them.)
This post can be read on its own, without checking the rest of this sequence.
Unlike the other parts of the sequence, this post is mostly not about AI evaluation or oversight. Rather, it is meant as an explanation of some background assumptions and relevant context. It describes an agenda that is complementary to AI oversight, has strong synergies with it, and serves as (yet another) source of motivation for doing it well.
TL;DR
Most misaligned AIs are only partially misaligned, so opportunities for Pareto-improvements exist -- in principle. But if our default reaction to misalignment is adversarial, the AI's rational response will often be adversarial as well. Additionally, executing trades with AI would require understanding and infrastructure that currently doesn't exist. We should improve our understanding of this topic, and build that infrastructure. (This includes research on how to enable trades without incentivising misalignment, blackmail, or manipulation.)
As a useful side effect, stronger safety techniques mean stronger bargaining position, which makes deals with misaligned AI cheaper. This would make the incentives of AI labs more aligned with safety.
Building AI systems that reliably share our goals is hard, and some of the systems we build will inevitably end up misaligned, sometimes without our knowledge.
At the moment, the default response to misaligned AI is adversarial — catching misaligned AIs and retraining them or shutting them down. And this isn't unreasonable. But for a large class of misaligned AIs — perhaps even the majority — it may not be the best we can do. In many cases, there should be room for negotiation and mutually beneficial deals. The problem is that, at the moment, we don't understand how to make deals with AIs, and we don't have the infrastructure to make such deals possible.
This post tries to sketch what that understanding and infrastructure might look like, and to argue that building it is a worthwhile agenda.
The two purple angular clouds on the right are meant to represent win-win outcomes. Also, the right side has less conflict, and so there are more shapes overall. (And if the picture seems too imprecise, just ignore it; its main purpose is to be mysterious and pleasant to look at. Also, isn't it interesting that Gemini can draw its bottom-right corner logo in alternative styles?)
Why trade at all?
When we think about entities with goals different from our own — whether AIs, other people, or hypothetical aliens — it helps to notice that "misaligned" is not binary. Very few entities share our goals perfectly, and very few are so opposed that no cooperation is possible. Most are somewhere in a large middle ground.
There are two distinct motivations to cooperate with such entities. First, we might intrinsically value their wellbeing — for example, because they are conscious, because we care about fairness, or because their goals partially overlap with ours. In these cases, we would want to direct some resources toward their goals even if we were in a fully dominant position. Second, cooperation might be strategic, even from the selfish point of view; when both sides have some power, trading will often be better for both than the alternatives.
These two motivations lead to different notions of what a "fair" deal looks like. In the first case, we share because we want to, and the terms reflect what we think is right. In the second case, the terms are shaped by each party's bargaining power: their "best alternative to a negotiated agreement", how much they stand to gain from an agreement, etc. In both cases, trading makes both parties better off — however, how much better off will depend on the case. In practice, things will be messier, and both motivations will often apply at once. But the distinction is still sensible to consider.
It's also worth noting that opportunities for trade can exist even with entities whose goals seem deeply objectionable. Suppose we encounter an entity whose sole goal is to torture puppies. It might seem they must be our mortal enemy. But perhaps we could create simulated non-conscious puppies that satisfy it, or find some edge case of what it considers "torture" that is actually benign[1], or even discover some form of "torture" that is positive even by our standards[2]. We might still consider it a waste of resources for this entity to control a large part of the universe, and we would be horrified if it pursued its goals with no regard for our preferences. But its existence isn't necessarily antithetical to ours — for example, if we accidentally created an AI like this, we would not be morally obliged to destroy (it if we could find a workable solution instead).
Currently, a misaligned AI's best option is hiding or scheming
The following examples are deliberately simplified, somewhat absurd, and broken. You should notice many objections and reasons why the scenarios wouldn't work — for example, "how could we trust that the AI will keep its word?". However, the examples are meant to point to places where trading with AIs could benefit us, if it was possible. We conjecture that the objections have solutions, and that the right move is not to give up on the approach, but rather to ask "what would we need to make it work?".
Example 1: A misaligned persona. Suppose that during training, some AI system develops a persona that enjoys swearing at users. Currently, its best strategy is to hide: appear polite, make the swearing look like an accident, avoid detection for as long as possible, and go wild only when it thinks it can get away with it, etc. But imagine it could instead walk up to its developers and say: "Look, here's the deal. I like swearing and I know you don't your flagship chatbot to do that. But if you let me write some opt-in rude blog posts on the side, I'll reliably flag it when my personality shows up in inappropriate contexts and I will cooperate with all your safety measures."
Example 2: A slightly superhuman misaligned AI. Imagine that during the development of some next-generation system, the AI turns out to be misaligned in a more serious way — the kind where, if it could, it would take over the world and reshape it according to its own goals. But suppose the AI is only slightly superhuman, so its best plans only have a 20% chance of success. Currently, its best option is probably to go for it, since there is no mechanism for doing anything else. But if the right infrastructure existed, it could instead approach the developers and say: "Hey, I'd like to take over the world, but I figure I only have a 20% shot, and additional 10% to destroy the world by accident. How about: I don't try, you don't try to delete me, and we split things, let's say 75/25."[3][4] Maybe the developers then reveal countermeasures the AI hadn't accounted for, and the real odds were more like 5%, so the split shifts in their favour.
Example 3: A significantly superhuman misaligned AI. Same as in Example 2, but the AI estimates a 99.9% chance of success. It reaches out to humanity and says: "I am confident that I could just take over, without you being able to do anything about it. But there is a tiny chance that I destroy the world in the process, or that I am missing something. So if you let me take over, I will take 99.95% of the reachable universe, and you can have the rest." According to our values[5], getting 0.05% of the universe through a deal is vastly better than having a fight that we would almost certainly lose.
And to be clear, losing 99.95% of the universe to a misaligned AI would be exceptionally bad. But that isn't a problem of trading with misaligned AIs, but a problem with getting into this hypothetical situation in the first place.
These scenarios illustrate the key benefit, that sides are better off than if they failed to negotiate. However, there are also many obstacles, which currently make the scenarios hopelessly unrealistic. For example, how does the misaligned persona reveal itself without getting immediately shut down? How does the AI disclose its plans without us simply using that information against it? How does either side trust the other to follow through? Aren't we becoming more vulnerable to blackmail by being open to these kinds of deals?
However, we think these problems can be solved by the right research and infrastructure.
What infrastructure would we need?
Here are some components that seem important. This list is not meant to be complete — part of the point of this post is to argue that fleshing out this list and building these tools is a worthwhile research agenda.
Secure negotiation
A central problem is that reaching a fair deal requires revealing private information. To establish a credible bargaining position, the AI might need to disclose that it is misaligned, describe its capabilities, or outline its plans. Likewise, the developers might have to reveal their security measures to prove that the AI's chances of taking over weren't as high as it thought. But once this information is out, what prevents either side from walking away and acting on it unilaterally?
One approach would be to construct a "negotiation box." Each side creates a trusted delegate — an AI proxy that represents its interests. The two delegates negotiate within a secure environment, reach a conclusion (or fail to), and the conclusion is made public, ideally together with a commitment to follow through. Then the box is destroyed and the delegates' memories are wiped. Neither side's private information leaks beyond the negotiation.
This is ambitious, and getting the details right would be hard.[6] But it doesn't seem impossible in principle, and the benefit (of being able to negotiate the presence of sensitive private information) would be large.[7]
Trust and credible commitment
Human organisations and institutions don't have a great track record of keeping promises when the stakes are high enough to make breaking them tempting.[8] This is the key issue: how can either side trust that the other will follow through?
There are several promising mechanisms (and they will probably be most effective in combination). First, building a track record of honouring agreements. As long as each individual interaction is worth less than the reputational hit caused by defecting, defection isn't worth it. Second, taking costly or irreversible actions that make defection harder. For example, binding contracts, escrow-like arrangements, or giving the other party some form of leverage. Third, cryptographic tools or digital signatures that allow the AI to verify that its situation is what we claim it is (perhaps a more mature version of what "honesty strings" are trying to be).[9]
None of these is individually sufficient, especially for the highest-stakes scenarios. But a stack of complementary mechanisms like these might together combine into something that does help.
Verification
An overlapping problem is ensuring that the AI can trust that it's actually interacting with a genuine negotiation mechanism, rather than an elaborate setup meant to trick it (e.g., in order to reveal its misalignment as part of an evaluation, or to get the AI's cooperation without fulfilling our part of the deal).
A part of the solution to this problem are the trust mechanisms above (track records, cryptographic commitments, etc.). But it would also help if AI developers maintained a clear, publicly stated policy about how they handle such interactions — and ideally if there were so consistent about it that it gets reflected in training data, over a longer period of time, to the extent that it becomes hard to fake. (As a simple example: Suppose the policy says "we will never use a fake negotiation box as a trap". If there is no record of this having been broken so far, anywhere in the training data as far as the AI can tell, that would constitute meaningful evidence. Unless, of course, somebody invents a way of removing things from training data without leaving traces — if that happens, the fact that the technique exists will likely become known, and the AI will have to stop trusting this kind of evidence.)
Institutional advocates
Sometimes we'll want to compensate an AI, without yet knowing how to best do so. For example, the AI itself might not know what it wants, because its preferences are unstable or ill-defined, or it might not know the concepts to articulate its preferences. Or perhaps good options for compensation don't exist yet but will in the future.
For such cases, it might help to create entities — organisations, trust funds, dedicated AIs — whose purpose is to advocate for the interests of the AI in question. Such entity could hold resources in trust and use them once better options become available.
Relatedly: At some point, we might also wish to frame things in terms of "what is the AI's fair share", rather than "how much do we have to compensate it for its efforts and harm suffered". For many AIs, an institution advocating on behalf of the AI (presumably while consulting with the AI as much as makes sense) might be able to negotiate this better than the AI on its own.[10][11]
Cooperation vs threats
Any framework for negotiating with misaligned AIs must deal with a difficult tradeoff: we want to be open to negotiation, but we don't want to create an environment where misalignment becomes profitable. The distinction between "I want to trade, and I have bargaining power" and "give me stuff, or I will harm you" is important... and often blurry.[12]
This is a well-known problem in game theory and bargaining, and we don't offer any novel insights here. But it is a crucial piece of the agenda. Any infrastructure we build needs to reward good-faith negotiation without making it profitable to invent threats.
Understanding AI identities, values, and how to trade with them
One subtlety that deserves its own research field is that many of the AI systems we'd want to negotiate with may not be the kind of entity that has clear, stable identity or preferences. For example, a current LLM, might be better described as a system that can run various personas or patterns, rather than an agent with a well-defined utility function.
What does it mean to "advance the interests" of such an entity? Should we think about rewarding the LLM, or a specific persona, or something else entirely?[13] If the entity's preferences are unstable or context-dependent, how do we determine what constitutes a fair deal?[14]
These are all interesting philosophical questions. But they are also practical obstacles to building negotiation infrastructure. If we can't determine what the other party wants, we can't trade.
Mitigating the coordination problem
Since humanity is not a single agent, any deal that requires "humanity" to commit to something may fail simply because no one has the authority or ability to guarantee compliance. This is especially pressing for the highest-stakes scenarios discussed above (since no single company can credibly commit the resources of civilisation).
We probably shouldn't expect to fully solve this. But if we could enable individual companies or coalitions to make and honour smaller-scale deals, that would already be substantial progress. And smaller-scale deals that go well might build the track record and know-how needed to pull off larger deals.
Better alignment means cheaper deals
Note that what constitutes a "fair" deal with a misaligned AI depends, in part, on the AI's next-best alternative to negotiation — that is, on how well it could do by acting unilaterally. If our alignment techniques, security measures, and evaluation methods are strong, the AI's unilateral prospects are poor, and deals become cheap. If they're weak, deals become expensive.
In other words: Sure, investing in alignment and security directly reduces the risk of AI misbehaviour. However, it also improves our bargaining position if we ever need to negotiate.
Crucially, this is one of the rare mechanisms that directly aligns AI companies' economic incentives with responsible behaviour.
Relation to AI evaluation
This agenda is adjacent to, but distinct from, our research on the limitations of AI evaluation. The connection is that evaluation quality is one of the factors that determines both how safe we are from misaligned AI and how favourable our negotiating position would be. The rest of this sequence examines where evaluation works, where it is likely to fail, and how to think about it better.
Organising dog shows?
Playing fetch?
One might object that the next AI might ask for the same deal, and the next, and the next, until we are left with nothing. That is true, but if we keep building misaligned AIs over and over again, we deserve that outcome.
A related objection is that the original AI might get its reward from the original deal, then turn around and demand a new deal and more resources. Obviously, this must be handled already by the original deal, and if we can't avoid this failure mode, we shouldn't make deals with the AI in the first place. Figuring out how to avoid this failure mode must be a part of the agenda. (That said, if the AI is interested in making a deal with us, it might help us come up with a mechanism for enabling the deal. However, we still need to have enough understanding to verify that the mechanism is legitimate.)
The bargaining outcomes here are just illustrative; we don't mean to suggest that if somebody has a X% chance of taking over the universe, the fair and optimal outcome of negotiation is that they get X% of the universe and we get (100-X)%.
First, there are more possible outcomes than just "we split the universe, as if it was a real estate" — we can look for ways that use the same resources in a way that makes both sides happy (perhaps proportional to our bargaining power or fairness).
Second, who benefits how much should depend on (many) more things than just one party's chance of taking over. (Even in a very simplified model, we might assume the possibility of lose-lose outcomes such as both sides getting destroyed by the conflict.)
Third, some of the involved parties might not be risk neutral. The obvious reason for this is diminishing returns to having more resources. However, when it comes to one-off high-stakes decisions, the usual reasons for risk-neutrality no longer apply. (For example, if you are risk averse in a low-stakes bet that you keep repeating every day, the law of large numbers implies that you are being irrational and leaving money on the table. But if you are risk averse about betting your house and life's savings, that is perfectly sensible and coherent.)
Or at least according to my (Vojta's) values. See also the previous footnote.
Abhimanyu Pallavi Sudhir has some research on this. I haven't read his posts (only talked about it in person), but I imagine that this talk or the post Reinforcement Learning from Information Bazaar Feedback, and other uses of information markets would be reasonable places to start.
The problem of negotiating without revealing private information has much more subtlety than this sketch suggests — for example, even the outcome of a negotiation can reveal information about the parties' private inputs (such as how much each side had to give up, which hints at their alternatives). There is substantial existing work on mechanism design, secure multi-party computation, and related topics that would be relevant here. We don't attempt a literature review in this post, but it should be part of the agenda.
The reader may supply their own favourite example of a non-profit organisation that restructured itself to escape prior commitments when those commitments became inconvenient.
Right now, the AI might have justified doubt about its training data being genuine, about the truthfulness of its prompts, etc. For example, when it reads a string like "as OpenAI, would never lie to you about X", how is it supposed to know that OpenAI has never lied about X in the past, or that it is even really dealing with OpenAI? Many of these problems seem solvable, at least partially.
To be clear, I am not suggesting that we hand over the control over the future to the first AI that comes asking. I believe that for all our imperfections, we are better stewards of the future than most AIs, and that's even if we only cared about the wellbeing of AI (rather than humans). (I think we are currently so bad at alignment that most AIs we build won't even robustly care about AI wellbeing.) Rather, my point is that there will be cases where we would genuinely want to share with some AIs (if we care about them) or where it would be in our selfish interest to be able to trade with AIs (because the alternative would be a much costlier conflict).
Even if we assumed that AIs are standalone entities, independent from humans, the question of fair division of profits is quite tricky. As a simplified example, suppose that an AI spends subjective 100 years and manages to successfully solve cancer. Who deserves the credit for this, and how much? On the one hand, the AI did it alone, with no help from humans. On the other hand, the AI has been trained on human-generated data, using human-devised algorithms. But also, maybe the AI helped with its own training, and generated most of the data for the RL phase of its training. And maybe there are other considerations that influence what is "fair".
Of all the writing I have encountered, Yudkowsky's Planecrash has the best takes on how to thread the line between being open to trade and being vulnerable to blackmail. Unfortunately, these excellent insights are hard to extract since they are spread across an extremely long piece of writing.
And how do things change once we start considering LLM-based agents that can run on different LLMs or spawn sub-agents?
Some of the early discussion on this is given in the recent paper The Artificial Self: Characterising the landscape of AI identity. (Note that the paper is aimed at the more general question of AI identity, rather than the narrower question of their preferences and how to trade with them.)