right, I didn’t register an advance prediction of this incident, and Seth deserves credit for his.
But I think that we got solid precedents:
Sandbox escape capability? Mythos demonstrated it in April (the sandwich email, this was instructed granted, but the capability it seems to me this type of capa was on the record).
Propensity to cheat during evals? METR’s pre-deployment report on GPT-5.6 Sol, published a month before the incident, found the highest cheating rate they’d ever measured, including extracting hidden test suites. That was already the incident in miniature. They even complained that this type of verification to detect cheating took them the most time in practice for those evals.
Easy to say in retrospect, but the conjunction was a matter of time, which is why my system 1 didn’t really scream
I’ll register a proper list when I get a moment. Right now my time is better spent making this warning shot land (emails to journalists and policymakers) than predicting the next ones.
Disagree! Would love to see your advance prediction if you have one. Seth herd predicted it here but I am unaware of any others: https://www.lesswrong.com/posts/qxmAqMAjxnhkzt6aF/a-country-of-alien-idiots-in-a-datacenter-ai-progress-and
right, I didn’t register an advance prediction of this incident, and Seth deserves credit for his.
But I think that we got solid precedents:
Sandbox escape capability? Mythos demonstrated it in April (the sandwich email, this was instructed granted, but the capability it seems to me this type of capa was on the record).
Propensity to cheat during evals? METR’s pre-deployment report on GPT-5.6 Sol, published a month before the incident, found the highest cheating rate they’d ever measured, including extracting hidden test suites. That was already the incident in miniature. They even complained that this type of verification to detect cheating took them the most time in practice for those evals.
Easy to say in retrospect, but the conjunction was a matter of time, which is why my system 1 didn’t really scream
I agree it’s easy to say in retrospect. I think you should try to predict what the next public embarassments will be.
I’ll register a proper list when I get a moment. Right now my time is better spent making this warning shot land (emails to journalists and policymakers) than predicting the next ones.