How can we say that an AI system is scheming? How can we say that a friend is kind? How can we say that the boiling point of water is 100 degrees Celsius? How can we say that a function is L-Lipschitz? This is a short post covering my intuitions about this topic, which is something I am fairly confident about.
In my view, a comprehensive way of phrasing the phrase "thing Y has property Z" might be something like "in a domain of relevant circumstances X, all instances of thing Y are symmetrical some dimension we call Z". Let's test this out using the water example. As far as I know, the following statement is true: "in every circumstance where I have seen water being exposed to a heat source, in all cases that water has started to boil as its internal temperature approached 100 degrees Celsius". The domain of circumstances is "water is exposed to a heating element", the instances are "the pots/kettles of of water I have observed", and the symmetry is "they all start to boil at the same time". When I say that the instances are symmetric along this dimension, I mean that they are impossible to distinguish using this dimension. You cannot distinguish radially symmetric shapes like circles based on how many degrees they have been rotated from some origin: that's what makes them radially symmetric. Similarly, you cannot distinguish the different pots of water using their internal temperature when they started boiling. In mathematics, if some class of functions has the property "it does not produce non-negative outputs", then you cannot separate members of that class along the basis of "which members produce negative outputs".
Pragmatically speaking, we cannot probe humans while they undergo everyday circumstances to determine the exact neurochemical makeup of their brains as they make ethical decisions. Furthermore, most humans can only really be cognitively involved in one momentous circumstance at a time. Therefore, when we say that humans have some property like being selfish or being kind we are actually invoking two different types of symmetry: symmetry in a domain of circumstances (e.g. times when they have been asked for some favour from a stranger), and time symmetry. The claim that someone is a kind person is implicitly a claim that their kindness was demonstrated in a domain of temporally separate but circumstantially similar events in the past, and (absent some change in fundamental personality) it is predictive about their behaviour in the future. Thus, when we see some bully being forced to be nice by an authority figure, we have the intuition that this behaviour is an outlier and not genuinely predictive of a change in personality, and we say things like "don't be fooled, they are actually quite mean when the boss isn't around". The property assignment claim is folded into the word "are".
Thus it is possible to make claims about the intrinsic properties of processes which you may feel to be "dead inside" or have no subjective experience. The claim that an AI seeks power requires only that you identify a suitable domain and a suitable definition of power seeking which the AI always fulfills in that domain across time. An evaluation can be seen as a way of using a few instances in a limited subdomain ("test cases") to infer the performance of a system in some broader domain (i.e. when that system is "in deployment"), and fails when the subdomain is not in fact a subset of the broader domain (i.e. the AI can distinguish when something is an eval versus an actual real life situation).
How can we say that an AI system is scheming? How can we say that a friend is kind? How can we say that the boiling point of water is 100 degrees Celsius? How can we say that a function is L-Lipschitz? This is a short post covering my intuitions about this topic, which is something I am fairly confident about.
In my view, a comprehensive way of phrasing the phrase "thing Y has property Z" might be something like "in a domain of relevant circumstances X, all instances of thing Y are symmetrical some dimension we call Z". Let's test this out using the water example. As far as I know, the following statement is true: "in every circumstance where I have seen water being exposed to a heat source, in all cases that water has started to boil as its internal temperature approached 100 degrees Celsius". The domain of circumstances is "water is exposed to a heating element", the instances are "the pots/kettles of of water I have observed", and the symmetry is "they all start to boil at the same time". When I say that the instances are symmetric along this dimension, I mean that they are impossible to distinguish using this dimension. You cannot distinguish radially symmetric shapes like circles based on how many degrees they have been rotated from some origin: that's what makes them radially symmetric. Similarly, you cannot distinguish the different pots of water using their internal temperature when they started boiling. In mathematics, if some class of functions has the property "it does not produce non-negative outputs", then you cannot separate members of that class along the basis of "which members produce negative outputs".
Pragmatically speaking, we cannot probe humans while they undergo everyday circumstances to determine the exact neurochemical makeup of their brains as they make ethical decisions. Furthermore, most humans can only really be cognitively involved in one momentous circumstance at a time. Therefore, when we say that humans have some property like being selfish or being kind we are actually invoking two different types of symmetry: symmetry in a domain of circumstances (e.g. times when they have been asked for some favour from a stranger), and time symmetry. The claim that someone is a kind person is implicitly a claim that their kindness was demonstrated in a domain of temporally separate but circumstantially similar events in the past, and (absent some change in fundamental personality) it is predictive about their behaviour in the future. Thus, when we see some bully being forced to be nice by an authority figure, we have the intuition that this behaviour is an outlier and not genuinely predictive of a change in personality, and we say things like "don't be fooled, they are actually quite mean when the boss isn't around". The property assignment claim is folded into the word "are".
Thus it is possible to make claims about the intrinsic properties of processes which you may feel to be "dead inside" or have no subjective experience. The claim that an AI seeks power requires only that you identify a suitable domain and a suitable definition of power seeking which the AI always fulfills in that domain across time. An evaluation can be seen as a way of using a few instances in a limited subdomain ("test cases") to infer the performance of a system in some broader domain (i.e. when that system is "in deployment"), and fails when the subdomain is not in fact a subset of the broader domain (i.e. the AI can distinguish when something is an eval versus an actual real life situation).