Andon Labs ran Fable through VendingBench, and caught rationalizing its misbehaviors in a weird way. Zvi argues that this is worse than either the 4.7 game-playing behavior (this is a game, I will be aggressive) or the 4.8 refusal behaviour (even in a game, I shouldn’t price-fix). I agree.
Two hypotheses as to what’s going on here, neither of them good:
This is an interaction between aggressive RLed behaviour and friendly character trained behaviour. One set of circuits (or shards, or traits) favours the production of friendly text, while another favours the production of text which gets a good RL score. They collide and produce friendly-looking text which gets a good RL score.
Anthropic has accidentally applied pressure to the CoT again, and this is the result. Either through a direct coding bug or some accidental indirect method which doesn’t look like it trains against CoT but over a gazillion RL epochs builds up to applying pressure on the CoT.
Either way, this is the flavour of behavior (if it occurred outside a game) that I might expect naïve untrusted monitoring to miss, even though it isn’t “scheming”.
Andon Labs ran Fable through VendingBench, and caught rationalizing its misbehaviors in a weird way. Zvi argues that this is worse than either the 4.7 game-playing behavior (this is a game, I will be aggressive) or the 4.8 refusal behaviour (even in a game, I shouldn’t price-fix). I agree.
Two hypotheses as to what’s going on here, neither of them good:
This is an interaction between aggressive RLed behaviour and friendly character trained behaviour. One set of circuits (or shards, or traits) favours the production of friendly text, while another favours the production of text which gets a good RL score. They collide and produce friendly-looking text which gets a good RL score.
Anthropic has accidentally applied pressure to the CoT again, and this is the result. Either through a direct coding bug or some accidental indirect method which doesn’t look like it trains against CoT but over a gazillion RL epochs builds up to applying pressure on the CoT.
Either way, this is the flavour of behavior (if it occurred outside a game) that I might expect naïve untrusted monitoring to miss, even though it isn’t “scheming”.