It is absolutely within OpenAI and Anthropic’s ability to build an air gapped system for developing and testing these models. <...> I think that their refusal to implement adequate safeguards is unjustifiable,
My issue is that OAI’s failure doesn’t provide much evidence about Anthropic’s corporate culture.
I’ve tried to edit my message for clarity (the cultural comments were meant to apply to OAI, not Anthropic) but we know that whatever safeguards Anthropic has are insufficient because the Mythos model regularly broke out of their sandboxes and gained access to the open internet during development and testing. According to the system card, about 20% of attempts by the model to do so were successful!
The quote which compelled me was the following:
My issue is that OAI’s failure doesn’t provide much evidence about Anthropic’s corporate culture.
I’ve tried to edit my message for clarity (the cultural comments were meant to apply to OAI, not Anthropic) but we know that whatever safeguards Anthropic has are insufficient because the Mythos model regularly broke out of their sandboxes and gained access to the open internet during development and testing. According to the system card, about 20% of attempts by the model to do so were successful!