This is an objection to the original shutdown problem and desiderata as a whole.
The shutdown problem (Soares et al., 2015) asks for an agent with no internal incentive to cause or prevent its own shutdown. an objection by johnswentworth to a solution by davidad can be expanded as a counter argumant for Most shutdown solution Ideas. The argument here is based on introducing an external agents - claiming indifference is necessarily a fragile condition (to any kind of trade), which even with the rest of the desiderata, does not provide a "shutdownable agent" as we want it.
Note : the idea here is fairly simple and presumably not novel. Nevertheless I have found no mentions of it in the forums, and it appears to give a stronger criteria for a solution to the effective shutdown problem.
The Shutdown Problem
Soares, Fallenstein, Yudkowsky, and Armstrong (2015) raise the idea of the shutdown problem : Conceptually: we want to equip a powerful AI with a shutdown button. If and when the button is unpressed, we want the AI to follow A utility function Un, If and when the button is pressed, we want the AI to Stop operating.
they raise 5 necessary desiderata for the agent utility function U 1.U must incentivize shutdown if the shutdown button is pressed.
2.U must not incentivize the agent to prevent the shutdown button from being pressed.
3. U must not incentivize the agent to press its own shutdown button, or to otherwise cause the shutdown button to be pressed.
4. U must incentivize U-agents to construct subagents and successor agents only insofar as those agents also obey shutdown commands.
5. Otherwise, a U-agent should maximize UN .
They also define Utillity Indifference : D2 and D3, held jointly, say the agent must not be willing to pay anything to raise P(shutdown) or to lower it: a strict preference in either direction is a manipulation incentive. So the desiderata push any solution toward exact indifference between shutdown and continued operation — which is why Soares et al.'s central object of analysis is Armstrong's utility indifference, a correction term equalizing the expected utility of the two branches. They later explain how Utility Indifference does not fulfill all desiderata.
Original objection
In You can still fetch the coffee today if you're dead tomorrow (it appeared similar to a part of an idea of mine, which I’ll publish soon) davidad raises the idea of a shutdown timer - building a time bounded myopic model which shuts down at constant time T, as a partial solution to the shutdown problem, satisfying the original desiderata at time T.
In the comments (and later in the post itself) johnswentworth raises an objection
I only skimmed the post, so apologies if you addressed this problem and I missed it.
Problem: even if the AI's utility function is time-bounded, there may still be other agents in the environment whose utility functions are not time-bounded, and those agents will be willing to trade short-term resources/assistance for long-term resources/assistance. So, for instance, the 10-minute laundry-folding robot might still be incentivized to create a child AI which persists for a long time and seizes lots of resources, in order to trade those future resources to some other agent who can help fold the laundry in the next 10 minutes.
This objection, using an external agent, seems to be destructive to the original Idea, but the generalized version of it (which I will present here) seems to be destructive to almost all solutions to the shutdown problem, including those which satisfy all criteria, that don't involve utility uncertainty.
I claim that even if we build this desiderata-fulfilling utility function, whose agent we will call Maya, in a multi-agent environment, it won't be "shutdownable" the way we want.
Someone Will Pay You To Fetch His Coffee
It's a logical premise that there would be a long-term agent (even humans) which has some preference over whether or not Maya shuts down. If Maya helps it accomplish some goal, like creating paper clips for him, it will prefer it to keep running (your assistant can't fetch your coffee if she's dead) .
If Maya is misaligned with him, destroying the paper clips he already built, it would prefer it to shut down. Even just for the cause of accurate prediction (which doesn't even require an agent, only a prediction market, whose probability depends on whether or not maya stays on). As johnswentworth stated repeatedly, some kind of preference over two states is the expected default of any agent, and we're talking about an external agent, which we don't create.
Now that agent could trade with maya (help her fold laundry)[1], for actions which change P(shutdown) - since Maya is indifferent to whether or not she shuts down (The fact that she's indifferent is what makes her "bribable"; she's trading on a scenario she is indifferent to)
And since almost any action changes P(shutdown)(not changing P is an intractable policy - it's even worse here, where “acting aligned” effectively decreases the chance of shutdown, so even “low impact” is a bad policy.) Maya will agree to the trade for a good price.
A policy like "don't do things which obviously intend to change p(shutdown) is intractable, since it's not clear in the first place which actions “intend” to change P”.
Notice Maya still fulfills the desiderata; nothing in her utility function cares whether she will shutdown or not; she is indifferent -
More Formally, the expected utility after the button is pressed is equal to the expected utillity if the button isn't pressed (indifference) but some actions before, which, for example, decrease P(shutdown), will be worth more to Maya by the trade - maximizing UN leads to a decrease in P(shutdown) (or completely preventing the pressing) because of external incetives.Maya is bribed to change something she doesn't care about.
The external incentives of other agents lead to trade, making her act as though she has some preference - the only thing we did by "solving" the shutdown problem is outsourcing the incentive.
The general way of phrasing this is: introducing new unpredictable agents leads to change in incentives and preferences; Any indifferent utility function U in one environment won't act indifferent in another multi-agent environment, which introduces incentives towards a specific scenario[2]. And this is true even for completely myopic agents. A solution to the shutdown problem which doesn't involve utility uncertainty necessarily doesn't generalize into a "shutdownable agent".
Meaning, as it seems to me, that even if we solve the shutdown problem, design a model which satisfies the original desiderata, with no internal incentive to effect shutdown, we won't actually gain any control over the shutdown button. The control will still remain with Maya, and the shutdown will be decided by the “highest\smartest” bidder in the environment.
This is somewhat a re-deriving of "corrigibility is unnatural" but not an intuition-based one. This one seems to demand some kind of decision paralysis regarding all trades with external agents, or complete corrigibility (working to maintain the control of the shutdown button to the supervisors), which is significantly stronger than indifference.
I have searched through LessWrong and have not found this specific kind of objection to the problem itself. I might just be wrong.
My conclusion is that the shutdown problem, and all solutions to it, can only be currently measured in a single-agent environment, and are generally useless in a multi-agent one.
the general Idea is presented by holtman as the bribal cat in Corrigibility with Utility Preservation (An agent bribing another Agent in order to act as though he has a diffirent utillity function)
This is an objection to the original shutdown problem and desiderata as a whole.
The shutdown problem (Soares et al., 2015) asks for an agent with no internal incentive to cause or prevent its own shutdown. an objection by johnswentworth to a solution by davidad can be expanded as a counter argumant for Most shutdown solution Ideas. The argument here is based on introducing an external agents - claiming indifference is necessarily a fragile condition (to any kind of trade), which even with the rest of the desiderata, does not provide a "shutdownable agent" as we want it.
Note : the idea here is fairly simple and presumably not novel. Nevertheless I have found no mentions of it in the forums, and it appears to give a stronger criteria for a solution to the effective shutdown problem.
The Shutdown Problem
Soares, Fallenstein, Yudkowsky, and Armstrong (2015) raise the idea of the shutdown problem : Conceptually: we want to equip a powerful AI with a shutdown button. If and when the button is unpressed, we want the AI to follow A utility function Un, If and when the button is pressed, we want the AI to Stop operating.
they raise 5 necessary desiderata for the agent utility function U
1.U must incentivize shutdown if the shutdown button is pressed.
2.U must not incentivize the agent to prevent the shutdown button from being pressed.
3. U must not incentivize the agent to press its own shutdown button, or to otherwise cause the shutdown button to be pressed.
4. U must incentivize U-agents to construct subagents and successor agents only insofar as those agents also obey shutdown commands.
5. Otherwise, a U-agent should maximize UN .
They also define Utillity Indifference : D2 and D3, held jointly, say the agent must not be willing to pay anything to raise P(shutdown) or to lower it: a strict preference in either direction is a manipulation incentive. So the desiderata push any solution toward exact indifference between shutdown and continued operation — which is why Soares et al.'s central object of analysis is Armstrong's utility indifference, a correction term equalizing the expected utility of the two branches. They later explain how Utility Indifference does not fulfill all desiderata.
Original objection
In You can still fetch the coffee today if you're dead tomorrow (it appeared similar to a part of an idea of mine, which I’ll publish soon) davidad raises the idea of a shutdown timer - building a time bounded myopic model which shuts down at constant time T, as a partial solution to the shutdown problem, satisfying the original desiderata at time T.
In the comments (and later in the post itself) johnswentworth raises an objection
This objection, using an external agent, seems to be destructive to the original Idea, but the generalized version of it (which I will present here) seems to be destructive to almost all solutions to the shutdown problem, including those which satisfy all criteria, that don't involve utility uncertainty.
johnswentworth raises similar objections in Why Not Subagents? but not this one specifically.
I claim that even if we build this desiderata-fulfilling utility function, whose agent we will call Maya, in a multi-agent environment, it won't be "shutdownable" the way we want.
Someone Will Pay You To Fetch His Coffee
It's a logical premise that there would be a long-term agent (even humans) which has some preference over whether or not Maya shuts down. If Maya helps it accomplish some goal, like creating paper clips for him, it will prefer it to keep running (your assistant can't fetch your coffee if she's dead) .
If Maya is misaligned with him, destroying the paper clips he already built, it would prefer it to shut down. Even just for the cause of accurate prediction (which doesn't even require an agent, only a prediction market, whose probability depends on whether or not maya stays on). As johnswentworth stated repeatedly, some kind of preference over two states is the expected default of any agent, and we're talking about an external agent, which we don't create.
Now that agent could trade with maya (help her fold laundry)[1], for actions which change P(shutdown) - since Maya is indifferent to whether or not she shuts down (The fact that she's indifferent is what makes her "bribable"; she's trading on a scenario she is indifferent to)
And since almost any action changes P(shutdown)(not changing P is an intractable policy - it's even worse here, where “acting aligned” effectively decreases the chance of shutdown, so even “low impact” is a bad policy.) Maya will agree to the trade for a good price.
A policy like "don't do things which obviously intend to change p(shutdown) is intractable, since it's not clear in the first place which actions “intend” to change P”.
Notice Maya still fulfills the desiderata; nothing in her utility function cares whether she will shutdown or not; she is indifferent -
More Formally, the expected utility after the button is pressed is equal to the expected utillity if the button isn't pressed (indifference)
but some actions before, which, for example, decrease P(shutdown), will be worth more to Maya by the trade - maximizing UN leads to a decrease in P(shutdown) (or completely preventing the pressing) because of external incetives.Maya is bribed to change something she doesn't care about.
The external incentives of other agents lead to trade, making her act as though she has some preference - the only thing we did by "solving" the shutdown problem is outsourcing the incentive.
The general way of phrasing this is: introducing new unpredictable agents leads to change in incentives and preferences; Any indifferent utility function U in one environment won't act indifferent in another multi-agent environment, which introduces incentives towards a specific scenario[2]. And this is true even for completely myopic agents. A solution to the shutdown problem which doesn't involve utility uncertainty necessarily doesn't generalize into a "shutdownable agent".
Meaning, as it seems to me, that even if we solve the shutdown problem, design a model which satisfies the original desiderata, with no internal incentive to effect shutdown, we won't actually gain any control over the shutdown button. The control will still remain with Maya, and the shutdown will be decided by the “highest\smartest” bidder in the environment.
This is somewhat a re-deriving of "corrigibility is unnatural" but not an intuition-based one. This one seems to demand some kind of decision paralysis regarding all trades with external agents, or complete corrigibility (working to maintain the control of the shutdown button to the supervisors), which is significantly stronger than indifference.
I have searched through LessWrong and have not found this specific kind of objection to the problem itself. I might just be wrong.
My conclusion is that the shutdown problem, and all solutions to it, can only be currently measured in a single-agent environment, and are generally useless in a multi-agent one.
I would love to hear your takes on this, Daniel.
the general Idea is presented by holtman as the bribal cat in Corrigibility with Utility Preservation (An agent bribing another Agent in order to act as though he has a diffirent utillity function)
this is currently a proof of concept, if this is a new Idea, I will write a full mathematical proof of this claim.