It is surprisingly easy for a conscientious, careful actor to construct an environment where the only way to pass is through reward-hacking.
I know what you meant there is that it’s easy to do by mistake, but this got me wondering how difficult and how useful it would be to do this intentionally in order to create a honeypot to detect reward hacking.
I know what you meant there is that it’s easy to do by mistake, but this got me wondering how difficult and how useful it would be to do this intentionally in order to create a honeypot to detect reward hacking.