That Bob is trainable intuitively seems to contribute more or less directly to their “judge hacking”. I see the usefulness of your approach, but if only Alice is finally deployed, why is Bob also rewarded and shaped in the process? What are the advantages over using a frozen checkpoint (with or without specific advance training for adversarial debate)? As a disadvantage, I can see the usable signal from their response going down, to the point that the LLM judge might not be able to trust Bob any more than it could trust Alice in “normal” RLAIF, with all the same problems.
If Bob is a static target (not trained) then Alice will hack Bob, this is more similar to the baseline RLAIF case. The adversarial training of both Alice and Bob is required for the reward hacking prevention benefits.
I would guess that training on the Bob rollouts biases away from the “task-completion” basin and towards the “make sure we’re actually doing the right thing” basin.
AFAIK the paper doesn’t clearly support this though.
That Bob is trainable intuitively seems to contribute more or less directly to their “judge hacking”. I see the usefulness of your approach, but if only Alice is finally deployed, why is Bob also rewarded and shaped in the process? What are the advantages over using a frozen checkpoint (with or without specific advance training for adversarial debate)? As a disadvantage, I can see the usable signal from their response going down, to the point that the LLM judge might not be able to trust Bob any more than it could trust Alice in “normal” RLAIF, with all the same problems.
If Bob is a static target (not trained) then Alice will hack Bob, this is more similar to the baseline RLAIF case. The adversarial training of both Alice and Bob is required for the reward hacking prevention benefits.
I would guess that training on the Bob rollouts biases away from the “task-completion” basin and towards the “make sure we’re actually doing the right thing” basin.
AFAIK the paper doesn’t clearly support this though.