Sure, but seems easier to monitor a less capable model for reward hacking, especially when its placed inside an environment whose singular point is measuring reward hacking ability.
Sure, but seems easier to monitor a less capable model for reward hacking, especially when its placed inside an environment whose singular point is measuring reward hacking ability.