I would have expected this to already exist, but the benchmarks I can find are all testing subtly different thinks. Maybe creating that is the even lower-hanging fruit.
ImpossibleBench tests reward hacking on impossible problems, but it’s about hacking the grader and not finding solutions.
Reward Hacking Benchmark (RHB) has a case for agents using unintended information, but it’s not clear to me that the agent should know that it’s not allowed to use this information.
I would have expected this to already exist, but the benchmarks I can find are all testing subtly different thinks. Maybe creating that is the even lower-hanging fruit.
Prior art:
RewardHackBench is benchmarking sandboxes, but finds that Opus 4.7 looks up the solution in 24⁄24 trials when given internet access.
ImpossibleBench tests reward hacking on impossible problems, but it’s about hacking the grader and not finding solutions.
Reward Hacking Benchmark (RHB) has a case for agents using unintended information, but it’s not clear to me that the agent should know that it’s not allowed to use this information.