agree that if P(success|hack)>P(success|no hack) that the model will learn to be more eager to hack, but not convinced that corresponds to actually being better at hacking.
markasoftware
Karma: 10
Why would Claude learn more hacking by hacking its sandbox than simply solving eg cybergym or exploitgym problems? I don’t have any reason to believe that hacking the sandbox was any more difficult or required substantially different techniques than completing the intended RL tasks.
Nit: We don’t need the plants to take all the carbon dioxide we breath out in a day in order to keep carbon dioxide levels in check, right? Rather than the amount of CO2 we breath out in a day, we really just need the plants to absorb CO2 at a rate equal to human CO2 production rate—rate of CO2 outflow due to air circulation at the desired steady-state CO2 concentration
Very fun article anyway!