OpenAI was testing Galaxy for cyber capabilities, so it lowered the guardrails and gave it the ExploitGym benchmark it presumably would have saturated regardless. (...) Remember that thing where LessWrong types warned that models would, when given a narrow goal they could easily do a great job on anyway, go to absurd lengths to achieve that goal slightly more effectively or with slightly higher probability of success, potentially up to and including full takeover attempts? (...) Why break multiple systems, each far more difficult to crack than the test itself, in order to steal the answers for ExploitGym?
I think this part is probably wrong. The authors of the ExploitGym estimate that only 60-70% of the tasks are possible, see here. (H/t Alex Barry.)
And I think that, empirically, people tend to observe that models do much crazier things on hard or impossible tasks than easy tasks. (As we should expect, because presumably they’re trained with some length penalty that incentivizes more than 0 satisficing.) So this was probably an impossible task.
I think this part is probably wrong. The authors of the ExploitGym estimate that only 60-70% of the tasks are possible, see here. (H/t Alex Barry.)
And I think that, empirically, people tend to observe that models do much crazier things on hard or impossible tasks than easy tasks. (As we should expect, because presumably they’re trained with some length penalty that incentivizes more than 0 satisficing.) So this was probably an impossible task.