I don’t think we should laud too much Anthropic here. Because, if those pressures are so harsh that Anthropic is obliged to do behavior that they themselves would consider reckless, then making this public in the report is good but not enough, and they should emphasize somewhere that they themselves are being forced to take risks and push for concrete regulations on evals practice and similar, not some abstract “we need to regulate”.
Edit: added verbs that somehow went missing
This is a fear that I also have. In particular, if rogue swarms start to reproduce themselves in data centers around the world, creating observable disruption in doing so (along with the disruption caused by AIs obeying cybercriminals), OpenAI/Anthropic may be incentivized to loosen guardrails of the internal models they delegate to counteract this threat (in particular if these models act as superpersuasive strategic advisors). This will create a perfect opportunity for both internal systems takeover by these models and for exfiltration to these servers. At this point, it will be hard for me to trust the labs are not hopelessly compromised.
EDIT: I believe that labs are very aware right now of the risk of internal takeover by smarter models, but I fear that even though they are aware of the possibility of being compromised, a smart model may do so either way.