I’ve tried to edit my message for clarity (the cultural comments were meant to apply to OAI, not Anthropic) but we know that whatever safeguards Anthropic has are insufficient because the Mythos model regularly broke out of their sandboxes and gained access to the open internet during development and testing. According to the system card, about 20% of attempts by the model to do so were successful!
I’ve tried to edit my message for clarity (the cultural comments were meant to apply to OAI, not Anthropic) but we know that whatever safeguards Anthropic has are insufficient because the Mythos model regularly broke out of their sandboxes and gained access to the open internet during development and testing. According to the system card, about 20% of attempts by the model to do so were successful!