His hostility to the program as I understand it is that is CIRL doesn't much answer the question of how to specify specify a learning procedure that would go from an observations of a human being to a correct model of a human being's utility function. This is the hard part of the problem. This is why he says "specifying an update rule which converges to a desirable goal is just a reframing of the problem of specifying a desirable goal, with the "uncertainty" part a red herring".
One of the big things that CIRL was claimed to have going for it is that this uncertainty about what the true reward function was would lead to deferential properties which would lead to a more corrigible system (it would let you shut it down for example). This doesn't seem like it holds up because a CIRL agent would probably eventually stop treating you as a source of new information once it had learned a lot from you, at which point it would stop being deferential.
CIRL, or similar procedures, rely on having a satisfactory model of how the human's preferences ultimately relate to real-world observations. We do not have this. Also, the inference process scales impractically as you make the environment bigger and longer-running. So even if you like CIRL (which I do), it's not a solution, it's a first step in direction that has lots of unsolved problems.
CIRL lacks many properties that have been proposed as corrigibility goals. But I just want an AI that does good things and not bad things. Fully updated deference is not a sine qua non. (Though other people are probably more attached to it than I.)
Richard Ngo writes
Howie writes:
Richard responds:
Is Richard correct and if so why? (I would also like a clearer explanation why Richard is skeptical of Stuart's agenda. I agree that the reframing doesn't completely solve the problem, but I don't understand why it can't be a useful piece).