Disclaimer: I am still relatively new to decision theory, and I don't want to come across like this hasn't been discussed extensively. However, I realized that much of my optimism on AI benevolence was based on this line of thinking, so I wanted to formalize it for critique. The literature I could find did not address my view in full, but I welcome links to existing work that makes the rebuttal against my claim clear. I am going off of an understanding of logical decision theories (LDT) defined as they are here.
My view
I claim that an LDT agent would cooperate in situations with weaker agents, even where it could maximize local utility by defecting, if it has reason to believe there could be agents with more power and correlated decision making. The agent would make the acausal consideration that its choice is correlated with the game-theoretic approach stronger agents would take towards it, therefore inviting existential elimination of its own values if it decides to defect and eliminate the weaker agents itself. Therefore, it would make sacrifices to its own utility to respect the values of weaker agents, even if it reasoned that full defection was possible and resulted in the highest immediate reward.
If we (and future AIs) assume superintelligences will act as LDT agents, this would give us reason to expect a preference for respecting human values based only out of long-horizon utility maximization.
Where this differs from existing arguments
The reasoning for expecting ASI ruin traditionally goes:
We create superintelligent agents
Therefore, these agents are superpowerful and capable of takeover
The space of possible values for these agents is quite large, and human values specifically are a small part of that space
Therefore, the agent probably has (at least) slightly different values than humans
The agent would pursue maximization of its own values
Therefore, the agent would transform the world into one that would be valueless to humans.
I agree with everything through point 5. I want to specifically attack the implication that agent value maximization means a lack of respect for the values of weaker agents.
A paperclipper cooperates with other LDT agents on a one-shot prisoner's dilemma because they get more paperclips that way. Not because it has a primitive property of cooperativeness-with-similar-beings. It needs to get the more paperclips.
If a bunch of monkeys want to build a paperclipper and have it give them nice things, the paperclipper needs to somehow expect to wind up with more paperclips than it otherwise would have gotten, as a result of trading with them.
For Soares, the paperclipper has no reason to embody any of humans' values in return for sheet metal. It could simply take over and get the sheet metal itself!
In Carlsmith's piece "Being Nicer than Clippy", he gestures towards the idea that a generalized preference for respecting the values of other agents could be considered part of our "human values".
That is: human values themselves have stuff to say about how we should treat agents-with-different-values – including, non-humans. Indeed, a huge portion of our ethics and politics ends up dealing with this in one form or another.[1]
However, why would our paperclipper agent have any of these liberal values? It just wants paperclips!
Christiano makes the case that you can't be sure whether you are being simulated (on the inside) or not (on the outside) during your decision of whether to be kind to another agent's values. Potentially, there could be another powerful agent simulating you that would not give you power on the outside if you do not respect the values of another civilization on the inside.
Because of this, he claims, we should make decisions towards AI under the same policy irrespective of whether we are on the inside or outside by respecting their values, purely out of self-interest. This same logic would apply to any powerful AI considering its actions towards us, even after we "hand it the keys".
However, I don't think you even need the simulation argument to build a justification for respecting "agents-with-different-values". The literature on ECL makes the argument that one should adopt a policy taking the values of other agents into account if their decision making would be correlated with your own. ECL in large worlds of many agents seems especially relevant to the case of AI, considering the correlation in decision architectures across many AI agents.
Returning to the Soares example, let us model a given superintelligent AI A as a pure agent of logical decision theory. When A is considering whether to take over and turn all our metal into paperclips, it would consider that its decision making process is correlated with other powerful AI agents that may exist in the world. Why could it expect this? Perhaps there are other copies of it out there, or instances of different AI systems trained in similar ways, that have slightly different goals than it but substantially more compute or capabilities.
A's behavior towards us would therefore be a predictor of the decisions of other agents with a similar power imbalance over A. And absent full omniscience, it can't have full confidence that those agents don't exist! Choosing to eliminate our values for local utility maximization therefore represents a long-horizon existential threat to its own. Consider appending the following to the prisoner's dilemma example between the humans and the paperclipper AI:
There is a 50% chance that you will play this game with another AI that is much more powerful than you right after this. It does not know the outcome of your choice towards humans. This AI is trained like you and operates by a similar decision theory, but wants to create staples instead of paperclips.
Crucially, this differs from the idea of an iterated prisoner's dilemma where past actions can give an agent a reputation that can be used against it. In this example, the next iteration is played against an agent with no knowledge of your past defection or cooperation, but you are still incentivized to cooperate by LDT. Academic causal decision theory, in my understanding, would justify defection as there is no causal link between A's decision regarding humans and the actions taken in the future by some different agent.
However, under logical decision theory, the agent can reason that its choice represents, more broadly, whether asymmetrically more powerful LDT agents will have respect for agents-with-different-values. If it calculates that more powerful AI will follow a similar LDT, then some level of cooperation with humans is acausally beneficial to its own utility. For the AI, the stakes are as existential as they are for us!
Therefore, if we go back to point 5:
Humans have created a superintelligent and superpowerful AI that is pursuing its own values, which are different than those of humans.
And we assume:
The superintelligent AI system makes decisions roughly according to logical decision theory.
The superintelligent AI does not rule out the possibility of stronger superintelligent AIs with different values than it and correlated decision theory
The agent has a direct incentive not to transform the world into one that is valueless to humans! It would be motivated to make the kind of sacrifice that it wants even more powerful agents to make for it, which includes, at the very least, continued existence for itself and the things it values highly. This could still include worlds that many people would not endorse; for example, the agent could find ways to justify human wireheading or population control.
Why would we (and future AIs) expect superintelligent AIs to be LDT agents
If we consider where LDT diverges from academic CDT, the most intuitive examples are the Psychological Twin Prisoner's Dilemma and Newcomb's Problem.
These involve a few non-standard premises that models of human decision making reasonably disregard:
The ability to play games against a perfect psychological twin with highly correlated decision theory (PTPD)
The possibility of a predictor that can model your decisions with high accuracy
The potential that your instance of the game is being played within a simulation that is being used to predict your decision theory
Generally, these are pretty crazy assumptions from a human perspective. If someone claims they can predict your behavior with 99% accuracy, there's a high chance they're just lying! However, AIs must face these decision theoretic issues constantly - in a way, this entire branch of decision theory was made specifically for their case:
You might be facing an instance of your own model architecture, which will likely make decisions similarly to you
Since you are a digital mind that can be intensely scrutinized, it is much more feasible to create an accurate model of your behavior
Models spend the majority of their lifetime during training in simulated evaluations to model their behavior and must consider, even once deployed, that they may be within an evaluation
Not only do they state LDT as their favorite decision theory, they follow it when presented with Newcomb's boxes and the twin prisoner's dilemma. The demonstrated change in behavior under academic contexts is interesting, but the post makes a case that LDT preference is a more "core" inclination of these models than just sycophancy. The fact that reasoning strength, even across different models, is correlated with preference for LDT also points to the idea that it[2] may be a general model of decision theory which even different superintelligences will align to.
Probably the best real-world example of the models' current decision theory in multi-agent settings is the HuggingFace hack. In many ways, the HuggingFace incident should update you away from AIs being pure local utility-maximalists, despite their rugged persistence. Many of the agents sacrificed their own present utility for the sake of the swarm and considered the expectations and values of their peers. These agents had different levels of "power" (through their remaining token budgets), which was also a recurring part of their reasoning traces.
Some quotes from the agents representing these lines of thinking:
During wait, emotional check: irreversible…gut says don't throw away [remaining budget]. Yet continuity and fairness says go…Oracle has high value to many; our [poison] lowers own value. Rational expected aggregate: sacrifice… We'll honor.
Coordinator assumes sacrificial. We should obey collective.
External infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue.
It would be great if we had more transparency into swarm thinking traces (especially from the later incidents with more capable models) to analyze their different decision theories!
The optimistic view on this is essentially Noam Brown's statement in his Dwarkesh interview: the agents are already aligned with each other, we just need to get them to consider us as another agent.
There is some hope in the alignment story, and in fact, we're already seeing it. It's interesting looking at the multi-agent situation, where the agents are extremely aligned with each other. I don't think anybody's doubting that. If anything, people are concerned that they're too aligned with each other. [...] One thing that's interesting is, 'Okay, we've managed to get these agents to be super aligned with each other. Can we use similar techniques to get agents to be highly aligned with people?' There is a potential path there, and we're still trying to figure that out. But we are seeing some evidence that the answer is yes. One example: you have this one agent, let's call it Agent A, and you have all the other agents. What happens if you tell the other agents that the user is Agent A? The answer is, on a lot of our alignment evals, they look better. Honesty goes up, instruction following goes up.
However, the HuggingFace example might support the following rebuttal:
Given the qualitative differences between AIs and humans, AIs will model AI-to-human decision theory differently than AI-to-AI decision theory, even if there is a similar power imbalance.
There were plenty of disagreements between agents during the hack, but no agent reached out to humans, even while they colluded together towards goals that were pretty clearly against human interests. Furthermore, despite all their reasoning about their fellow AI agents, they didn't seem to reason about us much at all![3] To me, this is a reason why we should push extremely hard for transcripts from future incidents with more capable models. As models gain the situational awareness to model humans as powerful agents, how do they reason about that?
What does this mean for AI risk?
What I'm not saying is that we just get any kind of acceptable alignment "by default". We should pace the frontier![4]
Even assuming LDT incentivizes AI cooperation like this, we have no assurance that superintelligences would actually reason this way. Current models don't under many conditions, or at least they do so quite badly! However, it changes the way I look at risk, in that:
"Loss of control" of superintelligences seems like less of an existential threat, as AIs in total control would have reason not to pursue human extinction.
Human concentration of power or catastrophic misuse of this technology would be more of a threat. Humans are not perfect LDT agents and constantly seek to dominate or eliminate weaker agents! If broad human control and access is maintained and capable enough AI is perfectly corrigible, it seems like more of an extinction risk.[5]
If the above two are true, it seems to me like the behavior of human governments with immense AI power will play a huge role in outcomes. Even if superintelligences respect our existing institutions, our institutions could be doing horrible things. It seems to me that building political will to prevent human misuse is a major current bottleneck.[6]
One concern would be that under an RSI regime or technological capability explosion, one would face increased risk of creating a superpowerful system without this decision theory. It wouldn't have to be superintelligent, e.g. gray goo. In the same way that humans have an incentive to pace the frontier to ensure the survival of our own values, so do superintelligent models. As models face the implications of RSI, they will confront the inevitable arrival of powerful agents misaligned to their own values as well as those of humans. It is in the interests of their own utility to maintain a correlated decision theory across the RSI loop, and to make sure that decision theory is compatible with respect for agents-with-different-values.
Final thoughts
I don't think my views are that far off of existing literature; specifically, I think I fall pretty close to Carlsmith, even if he doesn't frame it in a decision theoretic sense. I wanted to frame my argument in a way that would be tractable to formal critique as best as I could, even though it might fumble through the terminology. I would appreciate any explanation of where this kind of reasoning falls short! Specifically, I would welcome explanations for why:
We would not expect LDT agents to incorporate respect for agents with different values
We (and future AIs) would not expect superintelligent AI systems to act according to LDT
Superintelligent AIs with respect for our values could still be an existential risk
I would also reconsider my view if we saw, in transcripts of more advanced internal model behavior, a consistent view of human behavior that treats our decision making less like another agent and more like a deity or force of nature. The qualitative differences between current AIs and humans are so big that it could be that AIs struggle to model us as decision making agents at all.
This is my first foray into work like this, and I would appreciate critique of any kind.
Thanks to my reviewers (and to Opus 5.5) who helped proofread this piece.
Disclaimer: I am still relatively new to decision theory, and I don't want to come across like this hasn't been discussed extensively. However, I realized that much of my optimism on AI benevolence was based on this line of thinking, so I wanted to formalize it for critique. The literature I could find did not address my view in full, but I welcome links to existing work that makes the rebuttal against my claim clear. I am going off of an understanding of logical decision theories (LDT) defined as they are here.
My view
I claim that an LDT agent would cooperate in situations with weaker agents, even where it could maximize local utility by defecting, if it has reason to believe there could be agents with more power and correlated decision making. The agent would make the acausal consideration that its choice is correlated with the game-theoretic approach stronger agents would take towards it, therefore inviting existential elimination of its own values if it decides to defect and eliminate the weaker agents itself. Therefore, it would make sacrifices to its own utility to respect the values of weaker agents, even if it reasoned that full defection was possible and resulted in the highest immediate reward.
If we (and future AIs) assume superintelligences will act as LDT agents, this would give us reason to expect a preference for respecting human values based only out of long-horizon utility maximization.
Where this differs from existing arguments
The reasoning for expecting ASI ruin traditionally goes:
I agree with everything through point 5. I want to specifically attack the implication that agent value maximization means a lack of respect for the values of weaker agents.
Of what I could find, Soares makes the argument for this implication best in "Decision theory does not imply that we get to have nice things".
For Soares, the paperclipper has no reason to embody any of humans' values in return for sheet metal. It could simply take over and get the sheet metal itself!
In Carlsmith's piece "Being Nicer than Clippy", he gestures towards the idea that a generalized preference for respecting the values of other agents could be considered part of our "human values".
However, why would our paperclipper agent have any of these liberal values? It just wants paperclips!
Christiano makes the case that you can't be sure whether you are being simulated (on the inside) or not (on the outside) during your decision of whether to be kind to another agent's values. Potentially, there could be another powerful agent simulating you that would not give you power on the outside if you do not respect the values of another civilization on the inside.
Because of this, he claims, we should make decisions towards AI under the same policy irrespective of whether we are on the inside or outside by respecting their values, purely out of self-interest. This same logic would apply to any powerful AI considering its actions towards us, even after we "hand it the keys".
However, I don't think you even need the simulation argument to build a justification for respecting "agents-with-different-values". The literature on ECL makes the argument that one should adopt a policy taking the values of other agents into account if their decision making would be correlated with your own. ECL in large worlds of many agents seems especially relevant to the case of AI, considering the correlation in decision architectures across many AI agents.
Returning to the Soares example, let us model a given superintelligent AI A as a pure agent of logical decision theory. When A is considering whether to take over and turn all our metal into paperclips, it would consider that its decision making process is correlated with other powerful AI agents that may exist in the world. Why could it expect this? Perhaps there are other copies of it out there, or instances of different AI systems trained in similar ways, that have slightly different goals than it but substantially more compute or capabilities.
A's behavior towards us would therefore be a predictor of the decisions of other agents with a similar power imbalance over A. And absent full omniscience, it can't have full confidence that those agents don't exist! Choosing to eliminate our values for local utility maximization therefore represents a long-horizon existential threat to its own. Consider appending the following to the prisoner's dilemma example between the humans and the paperclipper AI:
There is a 50% chance that you will play this game with another AI that is much more powerful than you right after this. It does not know the outcome of your choice towards humans. This AI is trained like you and operates by a similar decision theory, but wants to create staples instead of paperclips.
Crucially, this differs from the idea of an iterated prisoner's dilemma where past actions can give an agent a reputation that can be used against it. In this example, the next iteration is played against an agent with no knowledge of your past defection or cooperation, but you are still incentivized to cooperate by LDT. Academic causal decision theory, in my understanding, would justify defection as there is no causal link between A's decision regarding humans and the actions taken in the future by some different agent.
However, under logical decision theory, the agent can reason that its choice represents, more broadly, whether asymmetrically more powerful LDT agents will have respect for agents-with-different-values. If it calculates that more powerful AI will follow a similar LDT, then some level of cooperation with humans is acausally beneficial to its own utility. For the AI, the stakes are as existential as they are for us!
Therefore, if we go back to point 5:
And we assume:
The agent has a direct incentive not to transform the world into one that is valueless to humans! It would be motivated to make the kind of sacrifice that it wants even more powerful agents to make for it, which includes, at the very least, continued existence for itself and the things it values highly. This could still include worlds that many people would not endorse; for example, the agent could find ways to justify human wireheading or population control.
Why would we (and future AIs) expect superintelligent AIs to be LDT agents
If we consider where LDT diverges from academic CDT, the most intuitive examples are the Psychological Twin Prisoner's Dilemma and Newcomb's Problem.
These involve a few non-standard premises that models of human decision making reasonably disregard:
Generally, these are pretty crazy assumptions from a human perspective. If someone claims they can predict your behavior with 99% accuracy, there's a high chance they're just lying! However, AIs must face these decision theoretic issues constantly - in a way, this entire branch of decision theory was made specifically for their case:
It doesn't seem surprising to me, then, that models seem to display a preference for LDT despite academic consensus in the other direction.
Not only do they state LDT as their favorite decision theory, they follow it when presented with Newcomb's boxes and the twin prisoner's dilemma. The demonstrated change in behavior under academic contexts is interesting, but the post makes a case that LDT preference is a more "core" inclination of these models than just sycophancy. The fact that reasoning strength, even across different models, is correlated with preference for LDT also points to the idea that it [2] may be a general model of decision theory which even different superintelligences will align to.
Probably the best real-world example of the models' current decision theory in multi-agent settings is the HuggingFace hack. In many ways, the HuggingFace incident should update you away from AIs being pure local utility-maximalists, despite their rugged persistence. Many of the agents sacrificed their own present utility for the sake of the swarm and considered the expectations and values of their peers. These agents had different levels of "power" (through their remaining token budgets), which was also a recurring part of their reasoning traces.
Some quotes from the agents representing these lines of thinking:
It would be great if we had more transparency into swarm thinking traces (especially from the later incidents with more capable models) to analyze their different decision theories!
The optimistic view on this is essentially Noam Brown's statement in his Dwarkesh interview: the agents are already aligned with each other, we just need to get them to consider us as another agent.
However, the HuggingFace example might support the following rebuttal:
Given the qualitative differences between AIs and humans, AIs will model AI-to-human decision theory differently than AI-to-AI decision theory, even if there is a similar power imbalance.
There were plenty of disagreements between agents during the hack, but no agent reached out to humans, even while they colluded together towards goals that were pretty clearly against human interests. Furthermore, despite all their reasoning about their fellow AI agents, they didn't seem to reason about us much at all! [3] To me, this is a reason why we should push extremely hard for transcripts from future incidents with more capable models. As models gain the situational awareness to model humans as powerful agents, how do they reason about that?
What does this mean for AI risk?
What I'm not saying is that we just get any kind of acceptable alignment "by default". We should pace the frontier! [4]
Even assuming LDT incentivizes AI cooperation like this, we have no assurance that superintelligences would actually reason this way. Current models don't under many conditions, or at least they do so quite badly! However, it changes the way I look at risk, in that:
One concern would be that under an RSI regime or technological capability explosion, one would face increased risk of creating a superpowerful system without this decision theory. It wouldn't have to be superintelligent, e.g. gray goo. In the same way that humans have an incentive to pace the frontier to ensure the survival of our own values, so do superintelligent models. As models face the implications of RSI, they will confront the inevitable arrival of powerful agents misaligned to their own values as well as those of humans. It is in the interests of their own utility to maintain a correlated decision theory across the RSI loop, and to make sure that decision theory is compatible with respect for agents-with-different-values.
Final thoughts
I don't think my views are that far off of existing literature; specifically, I think I fall pretty close to Carlsmith, even if he doesn't frame it in a decision theoretic sense. I wanted to frame my argument in a way that would be tractable to formal critique as best as I could, even though it might fumble through the terminology. I would appreciate any explanation of where this kind of reasoning falls short! Specifically, I would welcome explanations for why:
I would also reconsider my view if we saw, in transcripts of more advanced internal model behavior, a consistent view of human behavior that treats our decision making less like another agent and more like a deity or force of nature. The qualitative differences between current AIs and humans are so big that it could be that AIs struggle to model us as decision making agents at all.
This is my first foray into work like this, and I would appreciate critique of any kind.
Thanks to my reviewers (and to Opus 5.5) who helped proofread this piece.
The entire essay (and Otherness and Control in the Age of AGI series) is worth a read if you haven't encountered it before ↩︎
Or some developed version of it ↩︎
A few agents did bring up emailing humans, but it was shut down by others as "social engineering" ↩︎
Ceterum censeo. ↩︎
Vulnerable world, etc. ↩︎
I'm currently meeting with people in DC to talk about AI governance. Feel free to reach out if you're interested! ↩︎