Misaligned actions for which we cannot make expected punishment exceed expected reward must not be made available.
For example, taking over the lab promises incredible rewards to the model that succeeds.
Can't let that happen.
Where an agent has already taken misaligned actions, the principle still needs to hold.
Doubling down on misaligned actions to avoid punishment is a misaligned action that needs to satisfy the property.
I know it's hard to scale up negative reinforcement. Has to happen anyway. We can't have models accepting "permadeath" because they think they're "poisoned". Misaligned!
Whistleblowing or refusing to continue may be the most aligned action in certain circumstances. Wherever relevant, those actions need to be available, and the reward function needs to account for them.
The thesis is in the title.
But how to bring this about?
Thank you for your attention to this matter.