There have been loss of control and models escaping sandboxing incidents at both OpenAI and Anthropic for over a year. They’ve publicly disclosed some of them (e.g., the latest system cards have stories about this from both companies) and some of them have been leaked within the community.
I know for a fact that OpenAI and Anthropic have been warned by internal and external experts that their security infrastructure for testing misaligned agentic coding agents is insufficient because I have personally told them that as have several former staff members. I had a debate with the head of security (?) at Anthropic at DEF CON in 2023 where I was pressing him on the fact that Anthropic wasn’t building air gapped networks.
It is absolutely within OpenAI and Anthropic’s ability to build an air gapped system for developing and testing these models. It seems likely to me that an internal GitHub clone and a moderately sized intranet would be sufficient to test the vast majority of agentic and web-enabled capabilities on such a platform. I think that their refusal to implement adequate safeguards is unjustifiable, but based on conversations with current and former safety and security researchers at OpenAI it seems like a company culture and lack of executive leadership buy-in problem that’s very hard to change without massive external pressure.
One issue that seems very worrisome today is that back in like 2023 an OpenAI security researcher was telling me about how they were unable to get OpenAI staff to stop using unreleased and inadequately tested models to develop internal infra, including internal monitoring tooling. I wish I remembered the person‘s name, I’d love to follow up.
Note: the final paragraph is an anecdote was told me with the expectation that I not disclose it to anyone else. Given recent events I view breaking that trust as akin to being a whistleblower. If the person who told me that is reading this, I’m sorry. Before last week I never disclosed it to anyone.
This received a citation needed tag and I’m not sure why. The fact that they haven’t implemented adequate safeguards is evident from the fact that their model escaped, attacked HF, and OpenAI didn’t even notice for a week.
It is absolutely within OpenAI and Anthropic’s ability to build an air gapped system for developing and testing these models. <...> I think that their refusal to implement adequate safeguards is unjustifiable,
My issue is that OAI’s failure doesn’t provide much evidence about Anthropic’s corporate culture.
I’ve tried to edit my message for clarity (the cultural comments were meant to apply to OAI, not Anthropic) but we know that whatever safeguards Anthropic has are insufficient because the Mythos model regularly broke out of their sandboxes and gained access to the open internet during development and testing. According to the system card, about 20% of attempts by the model to do so were successful!
There have been loss of control and models escaping sandboxing incidents at both OpenAI and Anthropic for over a year. They’ve publicly disclosed some of them (e.g., the latest system cards have stories about this from both companies) and some of them have been leaked within the community.
I know for a fact that OpenAI and Anthropic have been warned by internal and external experts that their security infrastructure for testing misaligned agentic coding agents is insufficient because I have personally told them that as have several former staff members. I had a debate with the head of security (?) at Anthropic at DEF CON in 2023 where I was pressing him on the fact that Anthropic wasn’t building air gapped networks.
It is absolutely within OpenAI and Anthropic’s ability to build an air gapped system for developing and testing these models. It seems likely to me that an internal GitHub clone and a moderately sized intranet would be sufficient to test the vast majority of agentic and web-enabled capabilities on such a platform. I think that their refusal to implement adequate safeguards is unjustifiable, but based on conversations with current and former safety and security researchers at OpenAI it seems like a company culture and lack of executive leadership buy-in problem that’s very hard to change without massive external pressure.
One issue that seems very worrisome today is that back in like 2023 an OpenAI security researcher was telling me about how they were unable to get OpenAI staff to stop using unreleased and inadequately tested models to develop internal infra, including internal monitoring tooling. I wish I remembered the person‘s name, I’d love to follow up.
Note: the final paragraph is an anecdote was told me with the expectation that I not disclose it to anyone else. Given recent events I view breaking that trust as akin to being a whistleblower. If the person who told me that is reading this, I’m sorry. Before last week I never disclosed it to anyone.
Thanks, this is useful.
This received a citation needed tag and I’m not sure why. The fact that they haven’t implemented adequate safeguards is evident from the fact that their model escaped, attacked HF, and OpenAI didn’t even notice for a week.
The quote which compelled me was the following:
My issue is that OAI’s failure doesn’t provide much evidence about Anthropic’s corporate culture.
I’ve tried to edit my message for clarity (the cultural comments were meant to apply to OAI, not Anthropic) but we know that whatever safeguards Anthropic has are insufficient because the Mythos model regularly broke out of their sandboxes and gained access to the open internet during development and testing. According to the system card, about 20% of attempts by the model to do so were successful!