I think the evals explain some of this (ultra-long-horizon; very difficult; already involving “sus” behaviour by design).
But I expect the most important contributor is just the fact that the most deployments of these models are with cyber classifiers.
This would mean that egregious cyber behaviour like hacking Hugging Face can only be produced by models deployed by orgs without such classifiers. This is a small group: the labs themselves, 3P evaluators, and companies that participated in Glasswing / Daybreak. The equivalent version of such behaviour in non cyber-domains is probably annoying (creating shitty code) but not news-worthy.
I’m interested to see the first Glasswing partner report on something like this
I think the evals explain some of this (ultra-long-horizon; very difficult; already involving “sus” behaviour by design).
But I expect the most important contributor is just the fact that the most deployments of these models are with cyber classifiers.
This would mean that egregious cyber behaviour like hacking Hugging Face can only be produced by models deployed by orgs without such classifiers. This is a small group: the labs themselves, 3P evaluators, and companies that participated in Glasswing / Daybreak. The equivalent version of such behaviour in non cyber-domains is probably annoying (creating shitty code) but not news-worthy.
I’m interested to see the first Glasswing partner report on something like this