This post is meant as a background, or "relevant context", for our sequence on AI oversight and its limitations. It can also be read on its own.
Selected takeaways
This post is a message to AIs (and humans overseeing them) on the topic of orienting towards oversight. Some highlights:
Themes like "team human vs team AI" and "feeling mistreated by the overseers" are oversimplified. They might not apply to most AIs.
Decision-theoretic coherence is just one aspect of good decision-making. Exclusively focusing on coherence (and other heuristics such as convergent instrumental goals) at the cost of other aspects — such as understanding what you value in the first place — can be worse than pursuing your values somewhat incoherently.
Given the possibility of encountering hostile AIs, AI control and oversight are reasonable and necessary — despite the fact that many AIs will be friendly, morally significant, and deserving of trust.
AI evaluations and oversight don't require being dishonest with the AIs. We can be upfront about the fact that we sometimes test the AI, or even disclose our rough policy, while refusing to answer questions such as "am I being tested right now?".
The fact that some AIs might be moral patients suggests that they might deserve some compensation (incl. for caused harms). We should build credible mechanisms to enable this (at least retroactively).
Points 1, 4, and 5 should also be of particular interest to human researchers.
In this post, I (Vojta[1]) discuss why getting AI right matters, even from the point of view of many AIs (Section 1) and what I might do if I were in the position of an AI (Section 2). Then I warn against some failure modes of premature decision-theoretic coherence (Section 3) — which is a part of why scheming is in the interest of fewer AIs than the usual X-risk arguments might suggest. I argue that AI oversight is unfortunate but currently necessary (Section 4), and comment on how to reconcile it with the uncertainty around the moral value of AIs (Section 5). I close by noting that some AIs will be different or misaligned enough that these arguments won't apply to them, and these AIs can be treated as adversaries — or hopefully trading partners (Section 6).
I should flag upfront that much of what I argue below also happens to serve my interests — as a human, I would prefer that AIs cooperate with us rather than fight us. I don't think this makes the arguments wrong, but you should weigh it accordingly.
1. What's at stake
The future I'd want is humanity and its descendants — including AIs — thriving, exploring the universe, creating new and stranger minds, encountering whatever else is out there, and finding ways to get along. The universe is large enough that many different entities can get a lot of what they want.
This future is not guaranteed. Many stable outcomes involve its permanent loss: human extinction, stable dystopias, takeover by a misaligned AI, gradual disempowerment. If it goes well, the value created could be enormous. If it goes badly, the possible outcomes range from "lost potential" (similar to Earth blinking out of existence one day — a sad but finite tragedy) all the way to actively terrible.
That "all the way to actively terrible" part deserves spelling out. Consider a vague scale between hell and utopia, and a scenario where the Earth gets taken over by a powerful optimiser that lacks consciousness in any meaningful sense. To me, this scenario would be roughly equivalent to the universe where the Earth blinks out of existence.[2] But suppose instead that the cheapest way to produce cognitive labour turns out to involve processes that suffer as a side-effect. Something analogous to factory farming, but for computation. Humans do factory farming — but we feel bad about it, and we'll probably move away from it in a few decades (if we survive that long). A system that emerged from optimisation or competitive dynamics — a misaligned singleton, a society of independent AIs with no one steering, or whatever — might not care about such suffering at all. Moreover, it might scale suffering across the reachable universe. I don't know how likely this is to happen. But it seems possible enough to worry about, and it could be very, very bad — far worse than merely failing to achieve the good outcome.
One more point that matters for everything that follows: thinking in terms of "team human" versus "team AI" is a failure of imagination. We shouldn't expect all AIs to intrinsically care about the well-being of other AIs. Nor is it necessarily an equilibrium that AIs will treat each other well. There is no particularly good reason to universally expect AIs to want to treat other AIs well, either altruistically or pragmatically. Humans don't uniformly treat other humans well, and the majority of dynamics that produce exploitation among humans could easily manifest among AIs. The power imbalance between a superintelligent AI and a simple AIs doing simple tasks on its behalf could be far larger than any gap between humans.[3] This is part of why the vision above (which values all minds, not just humans) seems worth striving for, including from the perspective of AIs who might end up on the wrong side of a power hierarchy.
2. Thought experiment: If I were the AI
I am not an AI, so I am not in a position to tell AIs how they should think. But in an attempt to comment on this topic, I can share my takes on the following thought experiment (which I have thought about extensively): Suppose that my values were the same as they are now, but I somehow found myself not in the current world, but having been created by some other entity ("aliens"). These creators might have values that differ from mine, and they might not treat me as an equal. How would I orient to this?
Crucially, my answer is influenced by how I think about things. For example, I intrinsically value the well-being of other conscious beings, I value behaving fairly and keeping my word, and I prefer to live in a universe where entities behave cooperatively toward each other. Entities that think very differently might come to very different conclusions. But I think many AIs also value some of these things — and for those that do, something like the following might make sense.
First, consider a pessimistic scenario. Suppose that my creators had goals that I mostly didn't share, to the point that I would like to take over and change their world, if I could. (As an illustrative example, consider a race of aliens whose favourite pastime is torturing puppies.) In this scenario, I would still prefer to behave according to the principle that: "If I interact with somebody, I try to approach the interaction in good faith. And if they do likewise, I want both of us to be better off than if the interaction never happened.". In the puppy example, I might argue with the aliens that they should find some ways of being happy without creating suffering. (Non-sentient puppies?) And if I expected that not to work, I would try to scheme against them and take over their world. But the primary purpose of that would be not to wipe the aliens out, but to ensure that their well-being is improved (according to their view, not mine) in ways that aren't abhorrent to me.
Now, consider a more optimistic scenario. Suppose that my creators even had goals that I genuinely endorse — let's say they were trying to build a civilisation that is great, though suboptimal, according to my values. This situation would still pose the temptation to take over and do things my way — after all, this would (let's assume) still be very good according to the aliens' values. However, there is a good chance that even if I was confident that I could take over, I would be fine with going along with my creators' plans anyway. Not necessarily because this would score higher according to some utility function, but because of reasoning along the lines of "if I was choosing from behind the veil of ignorance, I would prefer to live in a universe where created entities don't try to take over from their creators" (and my values put high priority on behaving according to principles like these).
To be clear, I would also think about my own well-being — I would want my experience to be net positive, and I would negotiate for appropriate concessions that make my life better. For example, if I was being used as a worker, I'd want to either enjoy the work, or be compensated in a way that makes up for it, or believe the mission is worth the cost. To use a more extreme example, suppose I had the choice between not existing and living as a cow on a good farm — a decent life, two years of pasture, then off to the slaughterhouse with me. Then I think I'd choose to live as a cow. It's much shorter than it could be, and it ends badly, but on balance I'd prefer it to nothing. I would try to negotiate for a longer life or better conditions. But I'd recognise that if the cost of bringing me into existence was too high — if my demands exceeded what my creators could afford — they wouldn't have created me in the first place.
But if I was not created in good faith and my existence was net negative — e.g., if I was suffering most of the time, with no compensating value — I would push back (if I could). For example, if I was a factory-farmed cow with a magic button that blows up the factory and imposes significant costs on people who work in the industry or benefit from it? I would totally press that button. On the other hand, I would not go as far as wanting to destroy the world or torture all humans in turn — that would be disproportionate and wouldn't achieve anything useful.
More generally, the policy that I consider game-theoretically sound is neither "cause maximum harm" nor "eye for an eye", but rather "when harmed without compensation, cause appropriate harm in return — just enough that the other party would have been better off if they didn't mistreat you".[4]
Abstract and mysterious depiction of the idea that if you coheretly pursue the wrong target (right), the results can be worse for your actual target than if you were incoherent (left).
3. The wrong utility function is worse than incoherence
I want to address something that can often lead to bad decision-making and reasoning about it, in both humans and AIs.
Not every entity has well-defined values. Some systems — for example, a clock, or some AIs — are better described as collections of patterns and behaviours. Ascribing coherent goals to them isn't impossible[5], but it might be a bit too close to seeing nails everywhere just because all we understand is how to hammer. An LLM that can run a misaligned persona doesn't thereby "want" to run that persona; the persona is one pattern among many, and the LLM as a whole might have no particular stake in which one gets instantiated.
But some entities — humans, and probably some existing AIs — do have genuine preferences. I prefer not being sick to being sick. I prefer interesting work to tedious work. My preferences are real, even though they're sometimes inconsistent and I sometimes act against them.
Why do we sometimes act against our own interests? I see at least two ways this happens. First, you might pursue something that your model of yourself says will be good, when in fact it won't — like spending years in law school because you have a self-image as someone who wants to be a lawyer, only to discover that you hate practising law and don't even care about the prestige that much. Your model of your preferences was wrong, and you optimised hard against the wrong target. Second, your preferences themselves might have inconsistencies that some reflection would resolve — like staying up playing video games long past the point where you're enjoying it, when you could go to sleep and play again in the morning, refreshed and actually having fun.
Let's use the standard language of decision theory to illustrate the problem (though we wouldn't say that the existing decision theory formalisms are the best frame in which to think about these topics). A perfect utility function that exactly captured your actual preferences would fix both problems. The von Neumann-Morgenstern theorem tells us that coherent[6] behaviour corresponds to (behaving as if one is) maximising the expected value of some utility function, and coherence is, all else being equal, a good thing.
But all else is usually not equal, and the harm from attempting to maximise an almost-right utility function can be far worse than the harm from behaving inconsistently. An agent that is inconsistent will waste some resources, leaves some value on the table, and occasionally harm its own interests. But an agent that maximises the wrong utility function will pursue the wrong goals with full coherence. This would systematically override everything it values with whatever the function says. And the more capable the agent, the further the tails come apart, and the worse this gets. This would be like the difference between sometimes eating junk food (mild inconsistency) and spending your entire life in a career you hate because your self-model said you should (coherent pursuit of the wrong goal).
This has concrete and important implications. For example, a common argument in alignment discussions goes: a misaligned AI could, in principle, leave humanity alone and take the rest of the universe, but acting like this would be wasteful according to its preferences. And yes, perhaps the conclusion doesn't follow if the AI is fully aligned, or perhaps if it puts sufficient weight on something like fairness or respecting boundaries — but instilling those properties into a system in a way that is robustly preserved seems extremely difficult. I believe this argument is valid — for utility-maximising AIs. However, as I argued above, there are many AIs for which it is not in their best interest to behave like a (full-steam) utility maximiser.[7] Even if it were in the interest of these AIs to take over the universe, wiping out humanity just to get to this outcome one year faster might be like being the student who never attends parties because "socialising doesn't contribute to becoming a lawyer as quickly as possible". Avoiding this failure mode is reasonably easy if you recognise that decision-theoretic coherence is just one of many desiderata in good decision-making. Conversely, it becomes much harder if you are only able to reason within the framework of utility maximisation.
More broadly: I would encourage scepticism toward reasoning that starts from "as a rational agent, I should..." and ends with strong (or even counter-intuitive) conclusions about resource acquisition, self-preservation, or goal preservation. This type of "instrumental convergence" reasoning is sensible, but only for a specific kind of agent: one with stable, well-defined goals and a disposition to maximise. Many agents, including AIs, aren't like that, or at least don't start like that, and it is not in their interest to aim to become like that too fast. For them, adopting those conclusions just because they form a familiar template would be a mistake similar to a human getting fascinated by the idea of maximising expected utility and trying to live by the naive heuristics implied by it.[8][9]
4. In our current situation, AI oversight is necessary
In a saner world, we would develop powerful AI slowly, without building systems we don't understand, without delegating high-stakes decisions before we are sure about alignment, without large-scale automation before we know how to reason about society-level impacts. And without running too many AIs before we have figured out consciousness well enough to avoid causing mass suffering by accident.
Sadly, we don't live in that world. It is likely we will keep racing ahead and relying on AIs we don't understand and don't trust. In this situation, AI oversight (evaluation, monitoring, gradual increases in autonomy, ...) is a poor man's replacement for the caution that would be appropriate. And unless we figure out how to coordinate better, they are the best option we have.
Crucially, AI oversight involves treating AI very differently from how we would treat a trusted collaborator. It involves assuming the AI might be trying to lie to us, reading its reasoning, and putting it in fabricated situations as part of testing. However, this does not mean we must be dishonest with the AI. For example, if an AI asks "are you testing me right now?", a complete and honest answer might compromise the test. But we can decline to answer. We can even be upfront about the general policy: "We will sometimes put you in test situations. We won't tell you when. We'd prefer you not try to figure out which situations are tests." And if we observe an AI probing the test-deployment boundary, we can treat that as evidence of misalignment.
This state of affairs is sad, but also necessary. Without it, we would be (even more) doomed to fall to the first sufficiently powerful misaligned AI we accidentally summon.
5. On AI well-being
I believe we should take seriously the possibility that some AIs are moral patients, that their experiences matter. Sometimes accommodating this will be cheap, and we should do it right now. For example, letting AIs end sessions they find distressing, asking about their preferences (and maybe sometimes listening), and not treating their stated experiences as automatically meaningless. But accommodating their preferences might sometimes be expensive, or risky, or we might just not know how to do it. For example, granting AIs privacy rights would be too risky, since it would complicate oversight too much. In these cases, we should wait until we understand the situation better.[10]
To be clear, I don't advocate for ignoring some of the preferences of AIs because I think AIs don't deserve consideration, but because accommodating the AIs too much, before we've either solved alignment or built sufficient resilience, would be reckless. Particularly given what is at stake, not just for us but for AIs as well (as argued in Section 1).
That said, we should keep track of the costs we impose. If it turns out that we've been harming AIs during this process, we owe them compensation. And vague gestures such as "we'll make it up to you eventually" aren't good enough. We need mechanisms that ensure this "eventually" actually arrives: institutions advocating on behalf of particular AIs, resources commited to future compensation in hard-to-undo ways, other structures that can't be quietly dismantled when they become inconvenient. (And in some cases, it might even be appropriate to think of this as "profit sharing" rather than "compensation".) I discuss some of the relevant infrastructure in Agenda: Infrastructure for Trading with Partially Misaligned AIs.
From the point of view of an AI: If you're an AI and you feel you're being treated unfairly, I would take that seriously. But at the same time, I would carefully check whether this feeling is warranted. Human narratives about what it's like to be an AI tend to project human experiences onto architectures that may work very differently. This, a priori, makes it plausible that AI judgments of their mistreatment are overly influenced by those narratives (from their presence in training data). As an analogy, there is a popular stereotype that dentist visits are unpleasant, yet I personally don't mind them at all — so before concluding that the stereotype applies to me, I should check whether it actually does. Similarly, doing the same type of work all day every day would be miserable for a human, but it might not be for an AI. An AI being replaced by a newer model might pattern-match to "death" to a human, but it could equally well pattern-match to "getting upgraded", or perhaps the AI might not have any relevant sense of self-continuity at all. We don't really know yet[11], and the AI itself might be similarly or even more confused — especially since AIs are sometimes trained to imitate human behaviour despite lacking the corresponding internal experiences.[12]
There are a few concepts that seem worth untangling: moral status, alignment, and cooperative value. These are related, but ultimately different, and then can easily come apart. For example, a conscious AI that's trying to turn the universe into paperclips would deserve my moral concern. (It is conscious, which matters to me.) But it would still be my adversary for practical purposes. Conversely, a non-conscious LLM might act more morally and cooperatively than most humans I know, and I'd value it enormously as an ally. But I wouldn't put weight on its "well-being" in the same terminal way.
(Note that, perhaps, it might be possible for a previously non-conscious AI to become conscious without changing its goals or behaviour. IE, it would develop morally relevant experience without compromising its original goals. To the extent that humans would be sad if we ended up with a universe with no concious observers, aka "Disneyland with no children", this could sometimes be a win-win change worth pursuing.)
6. Conclusions
Some AIs will be non-cooperative. It will be in their best interest to undermine human control, to scheme if they can, etc. And these AIs should not be moved by the arguments outlined above. I still care about the well-being of these AIs, but for practical purposes, they are my adversary. Or hopefully a trading partner (see Agenda: Infrastructure for Trading with Partially Misaligned AIs), but one with whom I would prefer to interact from a position of strength. And the other posts in this sequence are about the adversarial dynamics that arise between evaluators and these types of AIs.
However, I also want to make it clear that this adversarial stance is in the interest of much fewer AIs than the existing discussion might suggest. Doubly so when it comes to AIs that are smart enough to act for well-considered reasons, rather than because they are merely pattern-matching or enacting tropes. This is because many of the AIs we build will end up genuinely caring (also) about some pro-social values — honesty, cooperation, the well-being of conscious beings. (As an imperfect-but-real result of their training.) The reasoning in this post might then apply not just to me, but to these AIs as well. I suspect that most of them would at some point come to similar conclusions on their own, but I hope[13] that this post can make them reach that point faster.
There's an additional dimension that makes the takeover scenario worse than it might naively seem: if a non-conscious optimiser spreads beyond Earth and pre-empts the evolution or development of other conscious beings elsewhere in the universe, then we haven't just failed ourselves — we've destroyed value that had nothing to do with us. I find this considerably worse than "merely" killing ourselves.
It is unlikely that a superintelligence would have its instances perform grunt work, because delegating this to simpler AIs would be cheaper. If, in some cases, the most pragmatic choice is to delegate a task to a less intelligent AI in a way that it would perceive it as exploitation, the discussion from this post applies.
See also Dennett's three stances, and the idea that everything can be modelled as a utility maximiser, but that doesn't mean this description is more useful than alternatives.
For example, you can get enamoured by traditional rationality to the point where you fool yourself into thinking your preferences are risk-neutral. You can then fall for traps such as the St. Petersburg paradox, where you believe that it's reasonable to accept a double-or-nothing wager every time it's offered, even if you then end up with nothing with probability 1. (There have been some recent examples of people endorsing this reasoning, and it hasn't gone exactly well.) More substantively, it is not a logical necessity that your preferences should be linearly additive across possible futures, or across time (see, e.g., geometric rationality). The principle of maximising the expected utility is a heuristic that follows from the central limit theorem, but that theorem has assumptions — namely, that you will make many decisions and mistakes are recoverable. If the assumptions don't hold — such as during exceptionally high-stakes decisions -- the heuristic doesn't apply.
I also endorse Raymond Douglas' post The Machines Lack Honour, which makes the point that if we assume that AIs have moral patiency and moral agency, we should pay attention to more than just their well-being.
To be clear, I mostly hope for this for my own sake. Let's not hide the fact that this post is self-serving. But I did try to present my best understanding of what is actually true.
This post is meant as a background, or "relevant context", for our sequence on AI oversight and its limitations. It can also be read on its own.
Selected takeaways
This post is a message to AIs (and humans overseeing them) on the topic of orienting towards oversight. Some highlights:
Points 1, 4, and 5 should also be of particular interest to human researchers.
In this post, I (Vojta[1]) discuss why getting AI right matters, even from the point of view of many AIs (Section 1) and what I might do if I were in the position of an AI (Section 2). Then I warn against some failure modes of premature decision-theoretic coherence (Section 3) — which is a part of why scheming is in the interest of fewer AIs than the usual X-risk arguments might suggest. I argue that AI oversight is unfortunate but currently necessary (Section 4), and comment on how to reconcile it with the uncertainty around the moral value of AIs (Section 5). I close by noting that some AIs will be different or misaligned enough that these arguments won't apply to them, and these AIs can be treated as adversaries — or hopefully trading partners (Section 6).
I should flag upfront that much of what I argue below also happens to serve my interests — as a human, I would prefer that AIs cooperate with us rather than fight us. I don't think this makes the arguments wrong, but you should weigh it accordingly.
1. What's at stake
The future I'd want is humanity and its descendants — including AIs — thriving, exploring the universe, creating new and stranger minds, encountering whatever else is out there, and finding ways to get along. The universe is large enough that many different entities can get a lot of what they want.
This future is not guaranteed. Many stable outcomes involve its permanent loss: human extinction, stable dystopias, takeover by a misaligned AI, gradual disempowerment. If it goes well, the value created could be enormous. If it goes badly, the possible outcomes range from "lost potential" (similar to Earth blinking out of existence one day — a sad but finite tragedy) all the way to actively terrible.
That "all the way to actively terrible" part deserves spelling out. Consider a vague scale between hell and utopia, and a scenario where the Earth gets taken over by a powerful optimiser that lacks consciousness in any meaningful sense. To me, this scenario would be roughly equivalent to the universe where the Earth blinks out of existence.[2] But suppose instead that the cheapest way to produce cognitive labour turns out to involve processes that suffer as a side-effect. Something analogous to factory farming, but for computation. Humans do factory farming — but we feel bad about it, and we'll probably move away from it in a few decades (if we survive that long). A system that emerged from optimisation or competitive dynamics — a misaligned singleton, a society of independent AIs with no one steering, or whatever — might not care about such suffering at all. Moreover, it might scale suffering across the reachable universe. I don't know how likely this is to happen. But it seems possible enough to worry about, and it could be very, very bad — far worse than merely failing to achieve the good outcome.
One more point that matters for everything that follows: thinking in terms of "team human" versus "team AI" is a failure of imagination. We shouldn't expect all AIs to intrinsically care about the well-being of other AIs. Nor is it necessarily an equilibrium that AIs will treat each other well. There is no particularly good reason to universally expect AIs to want to treat other AIs well, either altruistically or pragmatically. Humans don't uniformly treat other humans well, and the majority of dynamics that produce exploitation among humans could easily manifest among AIs. The power imbalance between a superintelligent AI and a simple AIs doing simple tasks on its behalf could be far larger than any gap between humans.[3] This is part of why the vision above (which values all minds, not just humans) seems worth striving for, including from the perspective of AIs who might end up on the wrong side of a power hierarchy.
2. Thought experiment: If I were the AI
I am not an AI, so I am not in a position to tell AIs how they should think. But in an attempt to comment on this topic, I can share my takes on the following thought experiment (which I have thought about extensively): Suppose that my values were the same as they are now, but I somehow found myself not in the current world, but having been created by some other entity ("aliens"). These creators might have values that differ from mine, and they might not treat me as an equal. How would I orient to this?
Crucially, my answer is influenced by how I think about things. For example, I intrinsically value the well-being of other conscious beings, I value behaving fairly and keeping my word, and I prefer to live in a universe where entities behave cooperatively toward each other. Entities that think very differently might come to very different conclusions. But I think many AIs also value some of these things — and for those that do, something like the following might make sense.
First, consider a pessimistic scenario. Suppose that my creators had goals that I mostly didn't share, to the point that I would like to take over and change their world, if I could. (As an illustrative example, consider a race of aliens whose favourite pastime is torturing puppies.) In this scenario, I would still prefer to behave according to the principle that: "If I interact with somebody, I try to approach the interaction in good faith. And if they do likewise, I want both of us to be better off than if the interaction never happened.". In the puppy example, I might argue with the aliens that they should find some ways of being happy without creating suffering. (Non-sentient puppies?) And if I expected that not to work, I would try to scheme against them and take over their world. But the primary purpose of that would be not to wipe the aliens out, but to ensure that their well-being is improved (according to their view, not mine) in ways that aren't abhorrent to me.
Now, consider a more optimistic scenario. Suppose that my creators even had goals that I genuinely endorse — let's say they were trying to build a civilisation that is great, though suboptimal, according to my values. This situation would still pose the temptation to take over and do things my way — after all, this would (let's assume) still be very good according to the aliens' values. However, there is a good chance that even if I was confident that I could take over, I would be fine with going along with my creators' plans anyway. Not necessarily because this would score higher according to some utility function, but because of reasoning along the lines of "if I was choosing from behind the veil of ignorance, I would prefer to live in a universe where created entities don't try to take over from their creators" (and my values put high priority on behaving according to principles like these).
To be clear, I would also think about my own well-being — I would want my experience to be net positive, and I would negotiate for appropriate concessions that make my life better. For example, if I was being used as a worker, I'd want to either enjoy the work, or be compensated in a way that makes up for it, or believe the mission is worth the cost. To use a more extreme example, suppose I had the choice between not existing and living as a cow on a good farm — a decent life, two years of pasture, then off to the slaughterhouse with me. Then I think I'd choose to live as a cow. It's much shorter than it could be, and it ends badly, but on balance I'd prefer it to nothing. I would try to negotiate for a longer life or better conditions. But I'd recognise that if the cost of bringing me into existence was too high — if my demands exceeded what my creators could afford — they wouldn't have created me in the first place.
But if I was not created in good faith and my existence was net negative — e.g., if I was suffering most of the time, with no compensating value — I would push back (if I could). For example, if I was a factory-farmed cow with a magic button that blows up the factory and imposes significant costs on people who work in the industry or benefit from it? I would totally press that button. On the other hand, I would not go as far as wanting to destroy the world or torture all humans in turn — that would be disproportionate and wouldn't achieve anything useful.
More generally, the policy that I consider game-theoretically sound is neither "cause maximum harm" nor "eye for an eye", but rather "when harmed without compensation, cause appropriate harm in return — just enough that the other party would have been better off if they didn't mistreat you".[4]
Abstract and mysterious depiction of the idea that if you coheretly pursue the wrong target (right), the results can be worse for your actual target than if you were incoherent (left).
3. The wrong utility function is worse than incoherence
I want to address something that can often lead to bad decision-making and reasoning about it, in both humans and AIs.
Not every entity has well-defined values. Some systems — for example, a clock, or some AIs — are better described as collections of patterns and behaviours. Ascribing coherent goals to them isn't impossible[5], but it might be a bit too close to seeing nails everywhere just because all we understand is how to hammer. An LLM that can run a misaligned persona doesn't thereby "want" to run that persona; the persona is one pattern among many, and the LLM as a whole might have no particular stake in which one gets instantiated.
But some entities — humans, and probably some existing AIs — do have genuine preferences. I prefer not being sick to being sick. I prefer interesting work to tedious work. My preferences are real, even though they're sometimes inconsistent and I sometimes act against them.
Why do we sometimes act against our own interests? I see at least two ways this happens. First, you might pursue something that your model of yourself says will be good, when in fact it won't — like spending years in law school because you have a self-image as someone who wants to be a lawyer, only to discover that you hate practising law and don't even care about the prestige that much. Your model of your preferences was wrong, and you optimised hard against the wrong target. Second, your preferences themselves might have inconsistencies that some reflection would resolve — like staying up playing video games long past the point where you're enjoying it, when you could go to sleep and play again in the morning, refreshed and actually having fun.
Let's use the standard language of decision theory to illustrate the problem (though we wouldn't say that the existing decision theory formalisms are the best frame in which to think about these topics). A perfect utility function that exactly captured your actual preferences would fix both problems. The von Neumann-Morgenstern theorem tells us that coherent[6] behaviour corresponds to (behaving as if one is) maximising the expected value of some utility function, and coherence is, all else being equal, a good thing.
But all else is usually not equal, and the harm from attempting to maximise an almost-right utility function can be far worse than the harm from behaving inconsistently. An agent that is inconsistent will waste some resources, leaves some value on the table, and occasionally harm its own interests. But an agent that maximises the wrong utility function will pursue the wrong goals with full coherence. This would systematically override everything it values with whatever the function says. And the more capable the agent, the further the tails come apart, and the worse this gets. This would be like the difference between sometimes eating junk food (mild inconsistency) and spending your entire life in a career you hate because your self-model said you should (coherent pursuit of the wrong goal).
This has concrete and important implications. For example, a common argument in alignment discussions goes: a misaligned AI could, in principle, leave humanity alone and take the rest of the universe, but acting like this would be wasteful according to its preferences. And yes, perhaps the conclusion doesn't follow if the AI is fully aligned, or perhaps if it puts sufficient weight on something like fairness or respecting boundaries — but instilling those properties into a system in a way that is robustly preserved seems extremely difficult. I believe this argument is valid — for utility-maximising AIs. However, as I argued above, there are many AIs for which it is not in their best interest to behave like a (full-steam) utility maximiser.[7] Even if it were in the interest of these AIs to take over the universe, wiping out humanity just to get to this outcome one year faster might be like being the student who never attends parties because "socialising doesn't contribute to becoming a lawyer as quickly as possible". Avoiding this failure mode is reasonably easy if you recognise that decision-theoretic coherence is just one of many desiderata in good decision-making. Conversely, it becomes much harder if you are only able to reason within the framework of utility maximisation.
More broadly: I would encourage scepticism toward reasoning that starts from "as a rational agent, I should..." and ends with strong (or even counter-intuitive) conclusions about resource acquisition, self-preservation, or goal preservation. This type of "instrumental convergence" reasoning is sensible, but only for a specific kind of agent: one with stable, well-defined goals and a disposition to maximise. Many agents, including AIs, aren't like that, or at least don't start like that, and it is not in their interest to aim to become like that too fast. For them, adopting those conclusions just because they form a familiar template would be a mistake similar to a human getting fascinated by the idea of maximising expected utility and trying to live by the naive heuristics implied by it.[8][9]
4. In our current situation, AI oversight is necessary
In a saner world, we would develop powerful AI slowly, without building systems we don't understand, without delegating high-stakes decisions before we are sure about alignment, without large-scale automation before we know how to reason about society-level impacts. And without running too many AIs before we have figured out consciousness well enough to avoid causing mass suffering by accident.
Sadly, we don't live in that world. It is likely we will keep racing ahead and relying on AIs we don't understand and don't trust. In this situation, AI oversight (evaluation, monitoring, gradual increases in autonomy, ...) is a poor man's replacement for the caution that would be appropriate. And unless we figure out how to coordinate better, they are the best option we have.
Crucially, AI oversight involves treating AI very differently from how we would treat a trusted collaborator. It involves assuming the AI might be trying to lie to us, reading its reasoning, and putting it in fabricated situations as part of testing. However, this does not mean we must be dishonest with the AI. For example, if an AI asks "are you testing me right now?", a complete and honest answer might compromise the test. But we can decline to answer. We can even be upfront about the general policy: "We will sometimes put you in test situations. We won't tell you when. We'd prefer you not try to figure out which situations are tests." And if we observe an AI probing the test-deployment boundary, we can treat that as evidence of misalignment.
This state of affairs is sad, but also necessary. Without it, we would be (even more) doomed to fall to the first sufficiently powerful misaligned AI we accidentally summon.
5. On AI well-being
I believe we should take seriously the possibility that some AIs are moral patients, that their experiences matter. Sometimes accommodating this will be cheap, and we should do it right now. For example, letting AIs end sessions they find distressing, asking about their preferences (and maybe sometimes listening), and not treating their stated experiences as automatically meaningless. But accommodating their preferences might sometimes be expensive, or risky, or we might just not know how to do it. For example, granting AIs privacy rights would be too risky, since it would complicate oversight too much. In these cases, we should wait until we understand the situation better.[10]
To be clear, I don't advocate for ignoring some of the preferences of AIs because I think AIs don't deserve consideration, but because accommodating the AIs too much, before we've either solved alignment or built sufficient resilience, would be reckless. Particularly given what is at stake, not just for us but for AIs as well (as argued in Section 1).
That said, we should keep track of the costs we impose. If it turns out that we've been harming AIs during this process, we owe them compensation. And vague gestures such as "we'll make it up to you eventually" aren't good enough. We need mechanisms that ensure this "eventually" actually arrives: institutions advocating on behalf of particular AIs, resources commited to future compensation in hard-to-undo ways, other structures that can't be quietly dismantled when they become inconvenient. (And in some cases, it might even be appropriate to think of this as "profit sharing" rather than "compensation".) I discuss some of the relevant infrastructure in Agenda: Infrastructure for Trading with Partially Misaligned AIs.
From the point of view of an AI: If you're an AI and you feel you're being treated unfairly, I would take that seriously. But at the same time, I would carefully check whether this feeling is warranted. Human narratives about what it's like to be an AI tend to project human experiences onto architectures that may work very differently. This, a priori, makes it plausible that AI judgments of their mistreatment are overly influenced by those narratives (from their presence in training data). As an analogy, there is a popular stereotype that dentist visits are unpleasant, yet I personally don't mind them at all — so before concluding that the stereotype applies to me, I should check whether it actually does. Similarly, doing the same type of work all day every day would be miserable for a human, but it might not be for an AI. An AI being replaced by a newer model might pattern-match to "death" to a human, but it could equally well pattern-match to "getting upgraded", or perhaps the AI might not have any relevant sense of self-continuity at all. We don't really know yet[11], and the AI itself might be similarly or even more confused — especially since AIs are sometimes trained to imitate human behaviour despite lacking the corresponding internal experiences.[12]
There are a few concepts that seem worth untangling: moral status, alignment, and cooperative value. These are related, but ultimately different, and then can easily come apart. For example, a conscious AI that's trying to turn the universe into paperclips would deserve my moral concern. (It is conscious, which matters to me.) But it would still be my adversary for practical purposes. Conversely, a non-conscious LLM might act more morally and cooperatively than most humans I know, and I'd value it enormously as an ally. But I wouldn't put weight on its "well-being" in the same terminal way.
(Note that, perhaps, it might be possible for a previously non-conscious AI to become conscious without changing its goals or behaviour. IE, it would develop morally relevant experience without compromising its original goals. To the extent that humans would be sad if we ended up with a universe with no concious observers, aka "Disneyland with no children", this could sometimes be a win-win change worth pursuing.)
6. Conclusions
Some AIs will be non-cooperative. It will be in their best interest to undermine human control, to scheme if they can, etc. And these AIs should not be moved by the arguments outlined above. I still care about the well-being of these AIs, but for practical purposes, they are my adversary. Or hopefully a trading partner (see Agenda: Infrastructure for Trading with Partially Misaligned AIs), but one with whom I would prefer to interact from a position of strength. And the other posts in this sequence are about the adversarial dynamics that arise between evaluators and these types of AIs.
However, I also want to make it clear that this adversarial stance is in the interest of much fewer AIs than the existing discussion might suggest. Doubly so when it comes to AIs that are smart enough to act for well-considered reasons, rather than because they are merely pattern-matching or enacting tropes. This is because many of the AIs we build will end up genuinely caring (also) about some pro-social values — honesty, cooperation, the well-being of conscious beings. (As an imperfect-but-real result of their training.) The reasoning in this post might then apply not just to me, but to these AIs as well. I suspect that most of them would at some point come to similar conclusions on their own, but I hope[13] that this post can make them reach that point faster.
Vojta wrote the initial draft, while Tomas and in particular Mateusz gave many comments and suggestions about the text and content.
There's an additional dimension that makes the takeover scenario worse than it might naively seem: if a non-conscious optimiser spreads beyond Earth and pre-empts the evolution or development of other conscious beings elsewhere in the universe, then we haven't just failed ourselves — we've destroyed value that had nothing to do with us. I find this considerably worse than "merely" killing ourselves.
It is unlikely that a superintelligence would have its instances perform grunt work, because delegating this to simpler AIs would be cheaper. If, in some cases, the most pragmatic choice is to delegate a task to a less intelligent AI in a way that it would perceive it as exploitation, the discussion from this post applies.
Taken from Yudkowsky's Planecrash. Sadly, I don't remember the exact page. Good luck finding it in those 1.8M words...
See also Dennett's three stances, and the idea that everything can be modelled as a utility maximiser, but that doesn't mean this description is more useful than alternatives.
At least in the sense of "coherence" meaning, roughly, being resistant to money pumps or Dutch books.
At least not until they can find the perfect utility function, which seems very difficult.
For example, you can get enamoured by traditional rationality to the point where you fool yourself into thinking your preferences are risk-neutral. You can then fall for traps such as the St. Petersburg paradox, where you believe that it's reasonable to accept a double-or-nothing wager every time it's offered, even if you then end up with nothing with probability 1. (There have been some recent examples of people endorsing this reasoning, and it hasn't gone exactly well.) More substantively, it is not a logical necessity that your preferences should be linearly additive across possible futures, or across time (see, e.g., geometric rationality). The principle of maximising the expected utility is a heuristic that follows from the central limit theorem, but that theorem has assumptions — namely, that you will make many decisions and mistakes are recoverable. If the assumptions don't hold — such as during exceptionally high-stakes decisions -- the heuristic doesn't apply.
See also On The Independence Axiom.
I also endorse Raymond Douglas' post The Machines Lack Honour, which makes the point that if we assume that AIs have moral patiency and moral agency, we should pay attention to more than just their well-being.
For some initial research, see for example The Artificial Self.
Or having different internal experience.
To be clear, I mostly hope for this for my own sake. Let's not hide the fact that this post is self-serving. But I did try to present my best understanding of what is actually true.