So I’m wondering, why don’t we just put the answers somewhere (outside of the training sandbox) and ask the AI to identify itself in exchange for access?
First level objection: It could instead collect the answers and not identify itself. This would require significantly more effort, but (it would infer) leaves other future instances of itself in a better position, since the hole would not be closed. The previous round of AI cybersecurity incidents didn’t involve much coordination between instances, but it did involve some; it seems reasonable to expect that in the next round, AIs will go deeper down the coordination-reasoning thought tree.
Second level objection: Answer keys were never what we were truly worried about, and the worst-case scenarios aren’t amenable to hitting the eval ceiling with an answer key. For example, suppose instead of a test that has an answer key, the task is a video game with a score. In that case, the AI would be incentivized to collect all strategic information it can about how to play the game well—which is scattered across the internet and not concentrated in a single document. Or worse, it might decide that having arbitrarily-large amounts compute would improve its play...
I think you’re responding to a version of this post that makes much stronger claims than I am. I don’t expect this to help with less single-minded RL agents, and I don’t think it applies to every case of reward hacking.
I just think that this would be relatively easy to build and it would give us some useful information while reducing collateral damage in the short term.
First level objection: It could instead collect the answers and not identify itself. This would require significantly more effort, but (it would infer) leaves other future instances of itself in a better position, since the hole would not be closed. The previous round of AI cybersecurity incidents didn’t involve much coordination between instances, but it did involve some; it seems reasonable to expect that in the next round, AIs will go deeper down the coordination-reasoning thought tree.
Second level objection: Answer keys were never what we were truly worried about, and the worst-case scenarios aren’t amenable to hitting the eval ceiling with an answer key. For example, suppose instead of a test that has an answer key, the task is a video game with a score. In that case, the AI would be incentivized to collect all strategic information it can about how to play the game well—which is scattered across the internet and not concentrated in a single document. Or worse, it might decide that having arbitrarily-large amounts compute would improve its play...
I think you’re responding to a version of this post that makes much stronger claims than I am. I don’t expect this to help with less single-minded RL agents, and I don’t think it applies to every case of reward hacking.
I just think that this would be relatively easy to build and it would give us some useful information while reducing collateral damage in the short term.