I would guess that training on the Bob rollouts biases away from the “task-completion” basin and towards the “make sure we’re actually doing the right thing” basin.
AFAIK the paper doesn’t clearly support this though.
I would guess that training on the Bob rollouts biases away from the “task-completion” basin and towards the “make sure we’re actually doing the right thing” basin.
AFAIK the paper doesn’t clearly support this though.