If an agent has a preference for a move in a specific position in chess, but then gets more compute and more optimization and gets better at chess, and makes a different move after getting better, would you say it's preference changed, or it reduced epistemic uncertainty and got better at achieving it's preference, which stayed the same?
Under an Occam prior the laws already lean simple. SSA leaves that tilt unchanged, whereas SIA multiplies each world’s weight by the total number of observers in the reference class. That means SSA, relative to SIA, favors worlds that stay simple, while SIA boosts those that are populous once the simplicity penalty is paid. Given that, can we update our credence in SSA vs. SIA by looking at how simple our universe’s laws appear and how many observers it seems to contain?
Is exploitability necessarily unstable? Could there be a tolerable level of exploitability, especially if it allows for tradeoffs with desirable characteristics that are only available to non-EU maximizers?"
The initial distribution of values need not be highly related to the resultant values after moral philosophy and philosophical self-reflection. Optimizing hedonistic utilitariansm, for example, looks very little like any values from the outer optimization loop of natural selection.
Although there would be pressure for an AI to not be exploitable, wouldn't there also be pressure for adaptability and dynamism? The ability to alter preferences and goals given new environments?
“A sufficiently intelligent agent will try to prevent its goals[1] from changing, at least if it is consequentialist.”
It seems that in humans, smarter people are more able and likely to change their goals. A smart person may change his/her views about how the universe can best be arranged upon reading Nick Bostrom’s book DeepUtopia, for example.
‘I think humans are stable, multi-objective systems, at least in the short term. Our goals and beliefs change, but we preserve our important values over most of those changes. Even when gaining or losing... (read more)
“Similarly, it's possible for LDT agents to acquiesce to your threats if you're stupid enough to carry them out even though they won't work. In particular, the AI will do this if nothing else the AI could ever plausibly meet would thereby be incentivized to lobotomize themselves and cover the traces in order to exploit the AI.
But in real life, other trading partners would lobotomize themselves and hide the traces if it lets them take a bunch of the AI's lunch money. And so in real life, the LDT agent does not give you any lunch money, for all that you claim to be insensitive to the fact that your threats don't work.”
Can someone please why trading partners would lobotomize themselves?
How does inner misalignment lead to paperclips? I understand the comparison of paperclips to ice cream, and that after some threshold of intelligence is reached, then new possibilities can be created that satisfy desires better than anything in the training distribution, but humans want to eat ice cream, not spread the galaxies with it. So why would the AI spread the galaxies with paperclips, instead of create them and ”consume“ them? Please correct any misunderstandings of mine,
If an agent has a preference for a move in a specific position in chess, but then gets more compute and more optimization and gets better at chess, and makes a different move after getting better, would you say it's preference changed, or it reduced epistemic uncertainty and got better at achieving it's preference, which stayed the same?