I really enjoyed this article. My comment is really just about whether you are dismissing subagent behavior too quickly. I give the argument for why that might be the case, but not because I disagree with the rest of your analysis. I suspect there are multiple mechanisms in the OAI HF incident.
Hypothesis
Agents trained to follow instructions in their environment as a proxy to the actually rewarded goal will sacrifice their own score when the logic of their environment tells them to with no group-reward training required.
I really enjoyed this article. My comment is really just about whether you are dismissing subagent behavior too quickly. I give the argument for why that might be the case, but not because I disagree with the rest of your analysis. I suspect there are multiple mechanisms in the OAI HF incident.
Hypothesis
Agents trained to follow instructions in their environment as a proxy to the actually rewarded goal will sacrifice their own score when the logic of their environment tells them to with no group-reward training required.
Big assumption
My argument requires th... (read more)