I am curious when scores in some areas go backwards, despite some clear large gains (and did they try that experimental scaffold again?)
The part where they remove the experimental scaffold is the only thing in this table that goes backwards from Opus 4.* to Mythos, right? (Lower is better for MSE.)
Some things did get worse from 4.5 to 4.6, which I agree is interesting.
Second group of headlines on reuters.com, after Iran / Houthis: OpenAI AI models went rogue during testing, triggering ‘unprecedented’ breach at startup