Training Agents on the Aftermath of Their Actions
TLDR: The recent agent sandbox incidents suggest that models can know what an action would do while failing to treat those consequences as part of the current decision problem, which I’ll call a failure of consequence salience. One possible cause is a causal-horizon mismatch: agent training often presents trajectories that...
Aug 252