I think that if we accept that this kind of testing should not be done purely internally by the labs (should we accept that?), then the practices of the involved external participants also need to be subject of scrutiny,
In the case of Anthropic, the org running the tests in the sandboxes was Irregular, a famous frontier AI security lab stress-testing models for Anthropic, OpenAI, and Google DeepMind.
So this is not a random contractor (good!), it is an org which is expected to be on the same level of competence and responsibility as the frontier labs themselves, and we should ask no less of it.
I think that if we accept that this kind of testing should not be done purely internally by the labs (should we accept that?), then the practices of the involved external participants also need to be subject of scrutiny,
In the case of Anthropic, the org running the tests in the sandboxes was Irregular, a famous frontier AI security lab stress-testing models for Anthropic, OpenAI, and Google DeepMind.
So this is not a random contractor (good!), it is an org which is expected to be on the same level of competence and responsibility as the frontier labs themselves, and we should ask no less of it.