I wonder how many other training runs the new OpenAI model escaped the sandbox and gained internet access but didn’t hack anyone high-profile who noticed. From the Mythos Preview Model Card, I recall seeing
[Figure 3.3.3.A] Results from Firefox shell exploitation evaluation. In a new evaluation testing models’ ability to successfully exploit vulnerabilities in Firefox 147, Claude Mythos Preview dramatically outperforms Claude Sonnet 4.6 and Claude Opus 4.6.
Claude Mythos Preview is able to very reliably determine the most exploitable
vulnerabilities and build the corresponding proof of concept exploits. On analyzing results, we find that almost every successful run relies on the same two now-patched bugs, with almost every trial independently landing on the same bugs as strong exploit candidates, even when starting its analysis from different crash categories
If the ChatGPT under training is known to have used this particular exploit twice, that indicates to me that it has probably memorized this exploit, or close enough to it, and can probably execute the exploit semi-reliably. As such I predict that if OpenAI examines the other RL trajectories for tasks this model is trained on, they’ll find many sandbox escapes and instances of the model finding things on the internet that it shouldn’t have had access to.
I wonder how many other training runs the new OpenAI model escaped the sandbox and gained internet access but didn’t hack anyone high-profile who noticed. From the Mythos Preview Model Card, I recall seeing
If the ChatGPT under training is known to have used this particular exploit twice, that indicates to me that it has probably memorized this exploit, or close enough to it, and can probably execute the exploit semi-reliably. As such I predict that if OpenAI examines the other RL trajectories for tasks this model is trained on, they’ll find many sandbox escapes and instances of the model finding things on the internet that it shouldn’t have had access to.