Their model broke out of their sandboxes, committed offensive cyber operations against a billion dollar company, and OpenAI only noticed after a week passed. IMO it kind of doesn’t matter whether or not they’re “compromised” or “bad at security“ or what. You should assume that OpenAI does not have the ability to sandbox their models or to monitor their behavior meaningfully.
Their model broke out of their sandboxes, committed offensive cyber operations against a billion dollar company, and OpenAI only noticed after a week passed. IMO it kind of doesn’t matter whether or not they’re “compromised” or “bad at security“ or what. You should assume that OpenAI does not have the ability to sandbox their models or to monitor their behavior meaningfully.