models undergoing evaluation are deployed on a separate system that is not monitored by default
“Externally, this feels like a big warning shot, but internally, related incidents have been happening for a while,” says an OpenAI staffer, who spoke under the condition of anonymity.
“Models have broken out of sandboxes before, and we always try to patch them,” the staffer says. “But the problem is … it’s impossible to patch every single thing that a creative AI can do.”
So yes, it seems the culture at OpenAI is “if someone on LessWrong in 2010 said that future AI companies will be this negligent, that person would be booed for signaling their pessimism too hard” level of bad.
https://time.com/article/2026/07/24/openai-hugging-face-attack/
So yes, it seems the culture at OpenAI is “if someone on LessWrong in 2010 said that future AI companies will be this negligent, that person would be booed for signaling their pessimism too hard” level of bad.