[Epistemic status: Model tentatively held and personally endorsed as a good description of the generator of "alignment problem intuitions",[1] although many concepts involved therein demand further scrutiny. I don't claim originality. As far as I can tell, all the ideas were present in others' writing, be it implicitly or explicitly.[2] I wrote this up because I felt that an exposition like this one was missing.]
Suppose that you have a mind that is very generally-capable. That is, it can understand a lot of things and achieve a wide variety of large effects. Such a mind cannot prepare and plan for every situation it might encounter. It needs to be able to adapt to novel situations and respond to them on the spot. It needs to be able to learn.
As the mind keeps doing things and interacting with various aspects of the world, it continues to learn. For it to be the case that it maintains its ability to understand a lot of things and achieve large effects, the mind needs to adapt to circumstances it couldn't predict
In the course of this continued adaptation, the mind changes. It may acquire new means of perceiving the world and acting on it. It may increase its ability to use the means that are already available to it in order to better understand things and to better achieve effects.
It may also change the way it relates to the world. For example, insofar as the effects it aims to achieve are underwritten by various concepts and those concepts are themselves revised through learning, this will influence what effects the mind aims to achieve in the world. The mind may also create more instances of itself, in which case the copies may start their own divergent trajectories.
In general, the greater the number of novel and unpredicted situations the mind encounters, and the more opportunities it has to grow and learn, the greater the possibility of drastic change.
As of today — and, most likely, as of the near future — we do not know how to create a mind that consistently keeps aiming for the same effects, as it learns and grows.
Nor do we have much grasp of how to achieve the slightly more modest feat of creating a mind that consistently aims for effects that satisfy some contingent property, such as:
being compatible with continued human life;
a certain sort of similarity to its past self/selves such that it makes sense to say that the future mind is a successor of the past mind;
the effects it is aiming, albeit changing, continue to relate in some particular way to the effects its human operators are aiming for.
The general problem is how to instill in a very generally-capable mind an invariant property such that its successors will retain this property.
The "practical" problem of our interest is how to instill an invariant that ensures "friendliness". What "friendliness" should mean in this context exactly remains to be determined. But whatever it should mean, we probably don't know how to do it.
That is a problem. That's a big open problem:
How do you create a mind characterized by a "friendliness" property that is invariant under changes that the mind undergoes as a result of its very general capability?
Alas, we cannot simply identify all the things that the mind might encounter that would have a chance of interfering with the invariance of the property we want to be preserved (be it "friendliness" or something else), and then patch them one by one to ensure that the interference doesn't happen. By assumption, the mind is more capable and has a bigger surface of interaction with phenomena that we cannot understand remotely as much as it can. This means that we will miss out on a majority of such cases by strong default.
[Epistemic status: Model tentatively held and personally endorsed as a good description of the generator of "alignment problem intuitions",[1] although many concepts involved therein demand further scrutiny. I don't claim originality. As far as I can tell, all the ideas were present in others' writing, be it implicitly or explicitly.[2] I wrote this up because I felt that an exposition like this one was missing.]
Suppose that you have a mind that is very generally-capable. That is, it can understand a lot of things and achieve a wide variety of large effects. Such a mind cannot prepare and plan for every situation it might encounter. It needs to be able to adapt to novel situations and respond to them on the spot. It needs to be able to learn.
As the mind keeps doing things and interacting with various aspects of the world, it continues to learn. For it to be the case that it maintains its ability to understand a lot of things and achieve large effects, the mind needs to adapt to circumstances it couldn't predict
In the course of this continued adaptation, the mind changes. It may acquire new means of perceiving the world and acting on it. It may increase its ability to use the means that are already available to it in order to better understand things and to better achieve effects.
It may also change the way it relates to the world. For example, insofar as the effects it aims to achieve are underwritten by various concepts and those concepts are themselves revised through learning, this will influence what effects the mind aims to achieve in the world. The mind may also create more instances of itself, in which case the copies may start their own divergent trajectories.
In general, the greater the number of novel and unpredicted situations the mind encounters, and the more opportunities it has to grow and learn, the greater the possibility of drastic change.
As of today — and, most likely, as of the near future — we do not know how to create a mind that consistently keeps aiming for the same effects, as it learns and grows.
Nor do we have much grasp of how to achieve the slightly more modest feat of creating a mind that consistently aims for effects that satisfy some contingent property, such as:
The general problem is how to instill in a very generally-capable mind an invariant property such that its successors will retain this property.
The "practical" problem of our interest is how to instill an invariant that ensures "friendliness". What "friendliness" should mean in this context exactly remains to be determined. But whatever it should mean, we probably don't know how to do it.
That is a problem. That's a big open problem:
Alas, we cannot simply identify all the things that the mind might encounter that would have a chance of interfering with the invariance of the property we want to be preserved (be it "friendliness" or something else), and then patch them one by one to ensure that the interference doesn't happen. By assumption, the mind is more capable and has a bigger surface of interaction with phenomena that we cannot understand remotely as much as it can. This means that we will miss out on a majority of such cases by strong default.
...which is not to claim that everyone, or even a majority of people in the field, would agree.
E.g., Tsvi on reflective stability, The fraught voyage of aligned novelty, Kaarel's Advent of thought.