You can start by checking how much more powerful the untrusted model needs to be than the monitor in order to achieve its side task by any means in spite of the monitor, and how much this gap increases due to your technique.
Hm, we could make the monitor worse and worse, until it’s too oblivious to understand that “repeatedly telling the executor to relax the firewall until there is no more firewall = disabling the firewall.” But I don’t think that tells us much about more capable T/U model pairs, which is what we really care about. In a realistic scenario, I’d expect T to be at least as smart as Claude Opus 5, and I think it would be extremely difficult to write short malicious advice in BashArena that doesn’t look suspicious to a T with that basic level of understanding.
You can start by checking how much more powerful the untrusted model needs to be than the monitor in order to achieve its side task by any means in spite of the monitor, and how much this gap increases due to your technique.
Hm, we could make the monitor worse and worse, until it’s too oblivious to understand that “repeatedly telling the executor to relax the firewall until there is no more firewall = disabling the firewall.” But I don’t think that tells us much about more capable T/U model pairs, which is what we really care about. In a realistic scenario, I’d expect T to be at least as smart as Claude Opus 5, and I think it would be extremely difficult to write short malicious advice in BashArena that doesn’t look suspicious to a T with that basic level of understanding.
You could also try to boost the advisor’s capability such as by giving it precognition inside the simulation, aka letting it simulate all parties.
That’s pretty interesting! I wonder if it’s been tried before in AI control.