agree that if P(success|hack)>P(success|no hack) that the model will learn to be more eager to hack, but not convinced that corresponds to actually being better at hacking.
Why would Claude learn more hacking by hacking its sandbox than simply solving eg cybergym or exploitgym problems? I don't have any reason to believe that hacking the sandbox was any more difficult or required substantially different techniques than completing the intended RL tasks.
Nit: We don't need the plants to take all the carbon dioxide we breath out in a day in order to keep carbon dioxide levels in check, right? Rather than the amount of CO2 we breath out in a day, we really just need the plants to absorb CO2 at a rate equal to human CO2 production rate - rate of CO2 outflow due to air circulation at the desired steady-state CO2 concentration
Very fun article anyway!