Reward Sacrifice in the Hugging Face Incident May Generalize From Multi-Agent RL
Epistemic status: Trying a bold and narrow hypothesis for my first LessWrong post. In the METR & Redwood Research report about the Hugging Face Incident there are descriptions of agents willingly sacrificing their evaluation score to gain information that could be useful for the swarm. > "Many agents also made...
Sep 89