This is the first in a series of blog posts making the case for acausal interactions being both tremendously important and tractable to influence. This post states the series' main points and serves as an overview of the remaining posts. (At the time of posting, most of these other posts will not have been published yet. I will link them as they come out. Each paragraph represents a different upcoming post.)
There are likely no objective criteria for picking the best decision theory, so intelligent agents won't necessarily converge. Because of this, intelligent agents might follow CDT despite CDT looking unappealing to us. The argument here is very similar to the argument for agents not converging to the same values.
We explain updatelessness, and how it comesapart from simple commitment. As a teaser: One way of viewing updatelessness is to always act the way you wish you had committed to in advance where this sometimes requires you to think about what you would have liked to commit to "in advance of" being alive. If you're not updateless, it's possible to Dutch book you (trick you into giving your money away) as long as you follow EDT or CDT and there are better-than-chance predictors of your behaviour.
Part 2: Types of acausal interactions
Non-causal decision theories enable Evidential Cooperation in Large Worlds (ECL) where you cooperate with agents you're not causally interacting with (e.g.,agents that are very far away) because of your similarity to these other agents: If you cooperate, similar agents are more likely to also cooperate and if everyone cooperates with each other, everyone benefits. Participating in ECL is analogous to cooperating in a Prisoners' Dilemma against a near-copy of yourself.
Both causal and non-causal decision theories enable Blind Anthropic Cooperation (BAC) where agents simulate other agents and reward them for adherence to mutually beneficial rules because they think sufficiently many others are rewarding adherence to the same rules including the rule to make simulations and reward adherence to the rules. Non-causal decision theories allow you to coordinate on the most mutually beneficial rules.
Causal and non-causal decision theories enable Model-Based Acausal Trade (MBAT) where you take actions because someone else will "see" you take this action when predicting you. Compared to ECL and BAC, MBAT is more like ordinary trade.
MBAT allows agents to trick CDT agents to give them all their money through an adversarial offer. (Post will be based on this paper.)
Part 3: Implications and other practical matters
Which decision theory ASI ends up with clearly matters given that it influences how ECL, MBAT, and BAC go and CDT can get you tricked out of all your money. We want ASI to be decision-theory aligned.
Influencing AI's decision theory is time sensitive: There might be path-dependencies in what decision theory ASI has since there are likely no objective criteria for picking the best decision theory. The opinions of the first weak AGIs will likely influence the opinions of humans and successor AIs alike.
Multi-agent or self-play RL converges to CDT-like behaviour for spurious reasons. If future AIs are heavily trained using this type of RL, they might generalise to hold deeply pro-CDT intuitions. This would be very bad on the process level (we don't want ASI's decision theory to be the result of arbitrary training features) and likely very bad on the object level (CDT seems unappealing). (Previous paper, post, and experiment on this.)
Our technical agenda has two directions:
Influence capabilities: differentially accelerate AI's decision theory and conceptual reasoning abilities, so as to make AI give better advice / make better decisions where decision theory is relevant and uplift important conceptual work on decision theory.
Influence propensities: align the AI's (meta-)decision-theoretic inclinations directly and prepare any necessary technical demonstrations to convince AI labs that this is important.
Part 4: Extras
We present an acausal syllabus for people who wish to learn more.
Safe Pareto Improvements are an important game-theoretic concept that might improve future acausal interactions.
Model-based acausal trade (MBAT) works most straightforwardly with accurate physics simulations. However, these might be computationally infeasible. We explain how MBAT might work without physics simulations.
We explain anthropics and how anthropic theories combine with decision theories.
CDT achieves better outcomes when combined with a special theory of anthropics but remains vulnerable to acausal money-extraction schemes.
This is the first in a series of blog posts making the case for acausal interactions being both tremendously important and tractable to influence. This post states the series' main points and serves as an overview of the remaining posts. (At the time of posting, most of these other posts will not have been published yet. I will link them as they come out. Each paragraph represents a different upcoming post.)
Part 1: Background theories
We explain Evidential Decision Theory (EDT) and Causal Decision Theory (CDT), the main contender decision theories. CDT looks much worse in our view. Functional Decision Theory (FDT) doesn't deviate much if at all from updateless EDT in its recommendations, so we don't study it separately.
There are likely no objective criteria for picking the best decision theory, so intelligent agents won't necessarily converge. Because of this, intelligent agents might follow CDT despite CDT looking unappealing to us. The argument here is very similar to the argument for agents not converging to the same values.
We explain updatelessness, and how it comes apart from simple commitment. As a teaser: One way of viewing updatelessness is to always act the way you wish you had committed to in advance where this sometimes requires you to think about what you would have liked to commit to "in advance of" being alive. If you're not updateless, it's possible to Dutch book you (trick you into giving your money away) as long as you follow EDT or CDT and there are better-than-chance predictors of your behaviour.
Part 2: Types of acausal interactions
Non-causal decision theories enable Evidential Cooperation in Large Worlds (ECL) where you cooperate with agents you're not causally interacting with (e.g.,agents that are very far away) because of your similarity to these other agents: If you cooperate, similar agents are more likely to also cooperate and if everyone cooperates with each other, everyone benefits. Participating in ECL is analogous to cooperating in a Prisoners' Dilemma against a near-copy of yourself.
Both causal and non-causal decision theories enable Blind Anthropic Cooperation (BAC) where agents simulate other agents and reward them for adherence to mutually beneficial rules because they think sufficiently many others are rewarding adherence to the same rules including the rule to make simulations and reward adherence to the rules. Non-causal decision theories allow you to coordinate on the most mutually beneficial rules.
Causal and non-causal decision theories enable Model-Based Acausal Trade (MBAT) where you take actions because someone else will "see" you take this action when predicting you. Compared to ECL and BAC, MBAT is more like ordinary trade.
MBAT allows agents to trick CDT agents to give them all their money through an adversarial offer. (Post will be based on this paper.)
Part 3: Implications and other practical matters
Which decision theory ASI ends up with clearly matters given that it influences how ECL, MBAT, and BAC go and CDT can get you tricked out of all your money. We want ASI to be decision-theory aligned.
Influencing AI's decision theory is time sensitive: There might be path-dependencies in what decision theory ASI has since there are likely no objective criteria for picking the best decision theory. The opinions of the first weak AGIs will likely influence the opinions of humans and successor AIs alike.
Multi-agent or self-play RL converges to CDT-like behaviour for spurious reasons. If future AIs are heavily trained using this type of RL, they might generalise to hold deeply pro-CDT intuitions. This would be very bad on the process level (we don't want ASI's decision theory to be the result of arbitrary training features) and likely very bad on the object level (CDT seems unappealing). (Previous paper, post, and experiment on this.)
Our technical agenda has two directions:
Part 4: Extras
We present an acausal syllabus for people who wish to learn more.
Safe Pareto Improvements are an important game-theoretic concept that might improve future acausal interactions.
Model-based acausal trade (MBAT) works most straightforwardly with accurate physics simulations. However, these might be computationally infeasible. We explain how MBAT might work without physics simulations.
We explain anthropics and how anthropic theories combine with decision theories.
CDT achieves better outcomes when combined with a special theory of anthropics but remains vulnerable to acausal money-extraction schemes.