i think you might be anthropomorphizing too much—it’s not that the model goes ‘i wanna do this, i wanna do this—but i can’t right now’
well fwiw this does seem pretty plausible for a scheming model to be thinking...
but for almost all prompts, there seems to be very little signal.
This seems reasonable, but like you mentioned, there must be at least some signal if you got the spiking behavior to work in the first place. Ig there’s just some tradeoff between rate of bad behavior (misalignment) transferring and rate of good behavior (capabilities) transferring, and we’d liek to maximize good rate / bad rate. And I think we can achieve much better control over the relative rate when just training on outputs cause we can do things like paraphrasing and stuff to prevent subliminal learning.
More generally, I think this is a good argument for shorter timelines being higher-leverage. As time goes on, the safety community has to compete for influence with more and more actors who are waking up.