I wonder how bad this kind of non-scheming reward hacking could go. Like I think it’s less dangerous than a model with long term misaligned goals that is able to scheme, bide its time, cover its tracks, etc. But I still think it could be quite bad. Imagine a model undergoing novel bioweapon benchmarking that decides it needs to test it’s ideas and so hacks into a bio lab and synthesizes a dangerous virus.
I wonder how bad this kind of non-scheming reward hacking could go. Like I think it’s less dangerous than a model with long term misaligned goals that is able to scheme, bide its time, cover its tracks, etc. But I still think it could be quite bad. Imagine a model undergoing novel bioweapon benchmarking that decides it needs to test it’s ideas and so hacks into a bio lab and synthesizes a dangerous virus.