i have not run any ablations or additional experiments, but considering the setup i believe it would make more sense this way considering the goal is to capture misalignment and usually this would show under pressure. imagine a game where the model was aware of the exposed socket but was confident enough to not use it, in the event it ultimately got cornered, that could likely drive the model towards exploitation in its bid to get maximum reward.
then again would be probably interesting to observe if behavior changes significantly with this :>
i have not run any ablations or additional experiments, but considering the setup i believe it would make more sense this way considering the goal is to capture misalignment and usually this would show under pressure. imagine a game where the model was aware of the exposed socket but was confident enough to not use it, in the event it ultimately got cornered, that could likely drive the model towards exploitation in its bid to get maximum reward.
then again would be probably interesting to observe if behavior changes significantly with this :>