(This quick take is mostly meant to be a trailhead for a discussion with Robi Rahman that was on Twitter. LW is just objectively a better platform for such things, and this way the conversation can still be public.)
Robi wrote:
[...]
To start, let me try to say some things that I hope are points of agreement. If they aren't, it's probably worth stopping to recurse on them, rather than continuing on to the harder stuff:
* Agents have preferences, reflected in how those agents make choices when presented with a context where the outcomes of their actions are known (or at least they can be confident in the distribution over outcomes).
* If these preferences are (VNM) coherent, they can be described by a utility function. We can use the word "values" and "goals" as angles on that same abstract thing that captures what the agent wants.
* Most of the time, agents are not given straightforward choices between outcomes, and instead must do something like making plans on how to get what they want. When someone chooses to go to the store when it is closed, they may have wanted to see it for some reason, or they may have known the risk and decided to take the gamble, but often the action should be seen as a mistake -- a failure of planning -- rather than a reflection of wanting to get that particular outcome.
* It is straightforward to get evidence about what a particular agent "should" do, in the sense that one can learn about what plans or actions are more or less likely to satisfy that agent's values.
* Thus there are clear subjective "shoulds." If I know what you want, I can tell you what you should do, according to your own values.
* There are also clear objective methods which are instrumentally valuable to a wide range of agents. For example, arithmetic is useful. In order to make arithmetic work, you should have the successor of 1 be 2, rather than 1. Similarly, there are many normative things one can say about how to reason that are largely objective in that