Folks, we’re going to get so many warning shots. You’re going to get tired of warning shots. You’ll come to OpenAI HQ and say “please Mr. MTS, no more warning shots,” and they’ll say “I think I got it this time, fifteenth times’ the charm” *blows up orphanage*
Yep. We’re going to have A country of alien idiots in a datacenter. And it won’t just be coming from devs. We’ll have plenty of people deploying those genius idiots with a lack of foresight and safeguards that will lead to mayhem and shenanigans.
AI has always done strange and harmful things, regularly. It’s just been too incompetent for people to take much notice. Now it’s getting competent enough to do real damage when it goes off the rails.
I also don’t doubt we’ll also see AIs capable of self-replicating and maintaining themselves across the internet well before they’re capable of RSI. What’s going to happen to communications infrastructure when its flooded by a swarm of agents chasing random goals and breaking into things to accomplish them?
I also don’t doubt we’ll also see AIs capable of self-replicating and maintaining themselves across the internet well before they’re capable of RSI.
This is something I’m skeptical for mundane economic reasons. Self-hosting Kimi K3 is a significant enough technical endeavor that you can get moderately e-famous for a few weeks by pulling it off. It’s quite expensive, especially without economies of scale on your side, and the upper bound is always getting bigger.
For ordinary people, this isn’t so bad, GLM-5.2 can do most of what anyone would want a local LLM for. But if you’re a rogue AI that eats and breathes compute and hosting, you don’t have that kind of slack. You’re competing for hosting with everyone else who might want to use it, and legitimate and illegitimate enterprises[1] alike can make use of frontier LLMs.
It’s the same reason why we don’t see countless people making “passive income with AI”, as so many clickbait articles suggest. Barring unique skills or unique knowledge, it’s difficult to differentiate yourself from countless better-resourced organizations who have the same tools you do and are competing for the same pool of compute.
Either through jailbreaking or as an Iran-Contra sort of deal, where a state leverages its tech champions in service of its under-the-table dealings. I’m sure America, Russia, and China alike have arrangements where friendly hackers have access to their best stuff when they need it.
My understanding is that this is arguing that current LLMs are too slow and expensive for them to be able to run amok online, since they have to be big enough to have the capability to break out of their container, make copies, and raise money for/steal compute. I think this is a bad bet against future models getting much more efficient and cheap to run while still having these capabilities, as well as the AIs just designing lots of very effective self-replicating malware to help progress their goals.
My understanding is that this is arguing that current LLMs are too slow and expensive for them to be able to run amok online,
Not at all! I’m arguing that the relativecapability difference is what matters, and small models without real-world backing will be competing for resources with massive models being run with national governments’ money and power behind them.
If you had access to GPT-3 in 2011, when its closest competitors were Markov chain models, it could probably make you a millionaire pretty readily. Today, if you tried to build a business around GPT-3, you’d be quickly eaten alive by people with the same business idea running Opus-5, Kimi K3, or ChatGPT Sol.
A warning shot is a social construct; it’s not just a technical incident. If we all adopt this mentality, we won’t get a serious, convincing warning shot, and we might suffer the consequences.
It becomes a regulatory moment useful for x-risks only if x-risks have been explained. Otherwise, we’ll regulate cyber rather than loss of control, and the whole shot will be wasted; civilization is pretty drunk currently.
Of course we’ll get tired: my system 1 barely reacted to the Hugging Face event, because all of this is so terribly easy to predict at a meta level. But rationality means winning, and winning here means creating political will before we get boiled alive. That’s currently the main bottleneck. From where I stand, the press has mostly failed to register this event: in France, despite our efforts, it’s been sidelined by the ban on social media for under-15s.
But still, it’s doable, and journalists can be moved. Time for some shut up and do the impossible boring work, even if that means sending emails to journalists one by one.
right, I didn’t register an advance prediction of this incident, and Seth deserves credit for his.
But I think that we got solid precedents:
Sandbox escape capability? Mythos demonstrated it in April (the sandwich email, this was instructed granted, but the capability it seems to me this type of capa was on the record).
Propensity to cheat during evals? METR’s pre-deployment report on GPT-5.6 Sol, published a month before the incident, found the highest cheating rate they’d ever measured, including extracting hidden test suites. That was already the incident in miniature. They even complained that this type of verification to detect cheating took them the most time in practice for those evals.
Easy to say in retrospect, but the conjunction was a matter of time, which is why my system 1 didn’t really scream
I’ll register a proper list when I get a moment. Right now my time is better spent making this warning shot land (emails to journalists and policymakers) than predicting the next ones.
The trick is to pause right now, when warning shots are abundant but not deadly. The more we pause, the more warning shots we get, the more political momentum we get for further pause measures.
Folks, we’re going to get so many warning shots. You’re going to get tired of warning shots. You’ll come to OpenAI HQ and say “please Mr. MTS, no more warning shots,” and they’ll say “I think I got it this time, fifteenth times’ the charm” *blows up orphanage*
Yep. We’re going to have A country of alien idiots in a datacenter. And it won’t just be coming from devs. We’ll have plenty of people deploying those genius idiots with a lack of foresight and safeguards that will lead to mayhem and shenanigans.
AI has always done strange and harmful things, regularly. It’s just been too incompetent for people to take much notice. Now it’s getting competent enough to do real damage when it goes off the rails.
I also don’t doubt we’ll also see AIs capable of self-replicating and maintaining themselves across the internet well before they’re capable of RSI. What’s going to happen to communications infrastructure when its flooded by a swarm of agents chasing random goals and breaking into things to accomplish them?
This is something I’m skeptical for mundane economic reasons. Self-hosting Kimi K3 is a significant enough technical endeavor that you can get moderately e-famous for a few weeks by pulling it off. It’s quite expensive, especially without economies of scale on your side, and the upper bound is always getting bigger.
For ordinary people, this isn’t so bad, GLM-5.2 can do most of what anyone would want a local LLM for. But if you’re a rogue AI that eats and breathes compute and hosting, you don’t have that kind of slack. You’re competing for hosting with everyone else who might want to use it, and legitimate and illegitimate enterprises[1] alike can make use of frontier LLMs.
It’s the same reason why we don’t see countless people making “passive income with AI”, as so many clickbait articles suggest. Barring unique skills or unique knowledge, it’s difficult to differentiate yourself from countless better-resourced organizations who have the same tools you do and are competing for the same pool of compute.
Either through jailbreaking or as an Iran-Contra sort of deal, where a state leverages its tech champions in service of its under-the-table dealings. I’m sure America, Russia, and China alike have arrangements where friendly hackers have access to their best stuff when they need it.
My understanding is that this is arguing that current LLMs are too slow and expensive for them to be able to run amok online, since they have to be big enough to have the capability to break out of their container, make copies, and raise money for/steal compute. I think this is a bad bet against future models getting much more efficient and cheap to run while still having these capabilities, as well as the AIs just designing lots of very effective self-replicating malware to help progress their goals.
Not at all! I’m arguing that the relative capability difference is what matters, and small models without real-world backing will be competing for resources with massive models being run with national governments’ money and power behind them.
If you had access to GPT-3 in 2011, when its closest competitors were Markov chain models, it could probably make you a millionaire pretty readily. Today, if you tried to build a business around GPT-3, you’d be quickly eaten alive by people with the same business idea running Opus-5, Kimi K3, or ChatGPT Sol.
I think it only counts as a warning shot if it’s scary relative to the previous warning shots.
That said, I probably agree. Thank god for slow takeoff.
So far!
A warning shot is a social construct; it’s not just a technical incident. If we all adopt this mentality, we won’t get a serious, convincing warning shot, and we might suffer the consequences.
It becomes a regulatory moment useful for x-risks only if x-risks have been explained. Otherwise, we’ll regulate cyber rather than loss of control, and the whole shot will be wasted; civilization is pretty drunk currently.
Of course we’ll get tired: my system 1 barely reacted to the Hugging Face event, because all of this is so terribly easy to predict at a meta level. But rationality means winning, and winning here means creating political will before we get boiled alive. That’s currently the main bottleneck. From where I stand, the press has mostly failed to register this event: in France, despite our efforts, it’s been sidelined by the ban on social media for under-15s.
But still, it’s doable, and journalists can be moved. Time for some shut up and do the impossible boring work, even if that means sending emails to journalists one by one.
Disagree! Would love to see your advance prediction if you have one. Seth herd predicted it here but I am unaware of any others: https://www.lesswrong.com/posts/qxmAqMAjxnhkzt6aF/a-country-of-alien-idiots-in-a-datacenter-ai-progress-and
right, I didn’t register an advance prediction of this incident, and Seth deserves credit for his.
But I think that we got solid precedents:
Sandbox escape capability? Mythos demonstrated it in April (the sandwich email, this was instructed granted, but the capability it seems to me this type of capa was on the record).
Propensity to cheat during evals? METR’s pre-deployment report on GPT-5.6 Sol, published a month before the incident, found the highest cheating rate they’d ever measured, including extracting hidden test suites. That was already the incident in miniature. They even complained that this type of verification to detect cheating took them the most time in practice for those evals.
Easy to say in retrospect, but the conjunction was a matter of time, which is why my system 1 didn’t really scream
I agree it’s easy to say in retrospect. I think you should try to predict what the next public embarassments will be.
I’ll register a proper list when I get a moment. Right now my time is better spent making this warning shot land (emails to journalists and policymakers) than predicting the next ones.
sure, if we’re lucky
The trick is to pause right now, when warning shots are abundant but not deadly. The more we pause, the more warning shots we get, the more political momentum we get for further pause measures.