I think you’re responding to a version of this post that makes much stronger claims than I am. I don’t expect this to help with less single-minded RL agents, and I don’t think it applies to every case of reward hacking.
I just think that this would be relatively easy to build and it would give us some useful information while reducing collateral damage in the short term.
I think you’re responding to a version of this post that makes much stronger claims than I am. I don’t expect this to help with less single-minded RL agents, and I don’t think it applies to every case of reward hacking.
I just think that this would be relatively easy to build and it would give us some useful information while reducing collateral damage in the short term.