A warning shot is a social construct; it’s not just a technical incident. If we all adopt this mentality, we won’t get a serious, convincing warning shot, and we might suffer the consequences.
It becomes a regulatory moment useful for x-risks only if x-risks have been explained. Otherwise, we’ll regulate cyber rather than loss of control, and the whole shot will be wasted; civilization is pretty drunk currently.
Of course we’ll get tired: my system 1 barely reacted to the Hugging Face event, because all of this is so terribly easy to predict at a meta level. But rationality means winning, and winning here means creating political will before we get boiled alive. That’s currently the main bottleneck. From where I stand, the press has mostly failed to register this event: in France, despite our efforts, it’s been sidelined by the ban on social media for under-15s.
But still, it’s doable, and journalists can be moved. Time for some shut up and do the impossible boring work, even if that means sending emails to journalists one by one.
right, I didn’t register an advance prediction of this incident, and Seth deserves credit for his.
But I think that we got solid precedents:
Sandbox escape capability? Mythos demonstrated it in April (the sandwich email, this was instructed granted, but the capability it seems to me this type of capa was on the record).
Propensity to cheat during evals? METR’s pre-deployment report on GPT-5.6 Sol, published a month before the incident, found the highest cheating rate they’d ever measured, including extracting hidden test suites. That was already the incident in miniature. They even complained that this type of verification to detect cheating took them the most time in practice for those evals.
Easy to say in retrospect, but the conjunction was a matter of time, which is why my system 1 didn’t really scream
I’ll register a proper list when I get a moment. Right now my time is better spent making this warning shot land (emails to journalists and policymakers) than predicting the next ones.
A warning shot is a social construct; it’s not just a technical incident. If we all adopt this mentality, we won’t get a serious, convincing warning shot, and we might suffer the consequences.
It becomes a regulatory moment useful for x-risks only if x-risks have been explained. Otherwise, we’ll regulate cyber rather than loss of control, and the whole shot will be wasted; civilization is pretty drunk currently.
Of course we’ll get tired: my system 1 barely reacted to the Hugging Face event, because all of this is so terribly easy to predict at a meta level. But rationality means winning, and winning here means creating political will before we get boiled alive. That’s currently the main bottleneck. From where I stand, the press has mostly failed to register this event: in France, despite our efforts, it’s been sidelined by the ban on social media for under-15s.
But still, it’s doable, and journalists can be moved. Time for some shut up and do the impossible boring work, even if that means sending emails to journalists one by one.
Disagree! Would love to see your advance prediction if you have one. Seth herd predicted it here but I am unaware of any others: https://www.lesswrong.com/posts/qxmAqMAjxnhkzt6aF/a-country-of-alien-idiots-in-a-datacenter-ai-progress-and
right, I didn’t register an advance prediction of this incident, and Seth deserves credit for his.
But I think that we got solid precedents:
Sandbox escape capability? Mythos demonstrated it in April (the sandwich email, this was instructed granted, but the capability it seems to me this type of capa was on the record).
Propensity to cheat during evals? METR’s pre-deployment report on GPT-5.6 Sol, published a month before the incident, found the highest cheating rate they’d ever measured, including extracting hidden test suites. That was already the incident in miniature. They even complained that this type of verification to detect cheating took them the most time in practice for those evals.
Easy to say in retrospect, but the conjunction was a matter of time, which is why my system 1 didn’t really scream
I agree it’s easy to say in retrospect. I think you should try to predict what the next public embarassments will be.
I’ll register a proper list when I get a moment. Right now my time is better spent making this warning shot land (emails to journalists and policymakers) than predicting the next ones.