One thing I don’t understand about your claims here:
“I don’t think we’ve gotten close to regulation that requires specific ambitious risk mitigations, pausing under certain circumstances, etc.”
Isn’t this literally what the EU Code of Practice requires though?
Eg Measure 4.2 of the Code of Practice requires that companies “will only proceed with the development, the making available on the market, and/or the use of the model, if the systemic risks stemming from the model are determined to be acceptable”.
Note that the literal way you “determine if systemic risks are acceptable” is to “use RSP methodology”. Just because it’s in EU-speak doesn’t mean it’s not regulation. Every major developer has said they’ll comply with the Code (even xAI!!).
Similarly, SB-53 and RAISE also require developers to have safety frameworks and follow them. Anthropic is the first company to change its RSP such that “following it” does not mean binding itself to commitments.
What’s the delta between this regulation and the kind of regulation that would have prevented you from dropping your commitments?
(To be clear, I’m not convinced any of this is a mistake on Anthropic’s part, I just don’t understand the claim about regulation).
I think the evals explain some of this (ultra-long-horizon; very difficult; already involving “sus” behaviour by design).
But I expect the most important contributor is just the fact that the most deployments of these models are with cyber classifiers.
This would mean that egregious cyber behaviour like hacking Hugging Face can only be produced by models deployed by orgs without such classifiers. This is a small group: the labs themselves, 3P evaluators, and companies that participated in Glasswing / Daybreak. The equivalent version of such behaviour in non cyber-domains is probably annoying (creating shitty code) but not news-worthy.
I’m interested to see the first Glasswing partner report on something like this