Strong agree. I’ve been working on a similar post - I’ve been calling this “the Hackening”—so I’ll put the thoughts I’ve had below.
As I understand it, pointing lesser agents at a section of code and saying “find the exploit”, in parallel, with high false positive rates, is not the same thing as a better agent grokking the whole environment and finding and chaining exploits together, at low false positive rates, into working steps of a cyberattack. The latter is the Juice.
I agree that harnesses can often make up the gap between open and frontier, and I haven’t seen anything that rules out that Kimi K3 swarms guided by T3MP3ST, or suchlike, could be worming their way across the Internet very soon.
Speaking of worms. Right now I classify threat actors into three fuzzy groups:
Criminals. These want ROI. Selling access, crypto and credential theft, blackmail, payroll redirection, ransoms, etc. Cause a lot of externalities, but not nearly as much as…
Terrorists. Rogue nations, misanthropes, or major states acting in the attribution fog of war. Vandalism, stealing secrets, destruction of physical machinery.
AI worms. Ones that broke out from ordinary deployments, or ones from the previous two groups that don’t answer home anymore. Uncontrolled, and increasingly selected for self propagation.
As for how many hacks can happen. Sol is telling me—so take this with heaps of salt—using Mythos’ weights, indexing off of The Last Ones, you get roughly .1-10 pwns of weak, flat-topography enterprises per cluster of 32 H200s per hour. There are dozens of tail compute holders that have hundreds of H200 equivalents. Mid tiers have thousands, or tens or hundreds of thousands. Hyperscalers have millions.
In dollar terms, that’s maybe $25-250 per pwn.
If an entity can simply steal compute—LLMjacking is a thing—then it gets it for free.
But datacenters that aren’t part of a nation’s arsenal won’t put up with this. The severity of the scenario depends a lot on how well they get their act together and shut out compute thieves, and either voluntarily—or under pressure from governments—start looking for and shutting out malicious paid inference.
Because the alternative is: there’s—again, Sol’s wild guess—perhaps 2 million easily pwnable orgs in the US, out of ~4.5 million. So, one month of the full compute of a single compromised or paid fleet of 100,000 H200 equivalents pwns ~half the US.
…Or maybe it’s a lot more. Or a lot less. I don’t know! I’d really like to see someone make a serious model of this.
Other considerations:
Overlap. If the open weights become common knowledge at the same time, possibly hundreds or thousands of actors will be fighting each other to get at the same pool of targets. The least defended targets will go down first, and the competition will work its way up from there.
Distillation. The smaller a model the Juice can run on, the cheaper attacks get. But if someone can get below certain thresholds—I believe 70B or 30B parameters—that opens up a lot more marginal compute to host malicious AIs, and greatly reduces the degree to which datacenter discipline can control outbreaks. Universities, workstations, hobbyists. Quite possible over the next few years at current scaling and algorithmic improvement trends. If it can fit on 7B—i.e. a phone- it is even worse.
Core services. AI suggests to me that if you can route your IT through a trusted security layer—formally verified, memory safe, frontier AI patrolled—that greatly reduces a lot of classes of vulnerability. Of course, the remaining attack efforts will concentrate on the remaining weak points. But this could help shift the balance.
Business realities. If an attack takes down an economic chokepoint, or a cluster of orgs that collectively act as a chokepoint, that could cascade through the economy. Also, ~half of businesses don’t have enough cash flow to deal with a month of down time and would simply collapse. Insurance will start forcing businesses to remediate their technical debt as soon as they start seeing the writing on the wall. If the (perception of the) business environment, or stability of digital banking, gets bad enough, this could trigger a credit contraction. And from there, a recession.
Attribution. If the chaos gets bad enough, or exploit windows start closing as things get patched, governments might decide now is the time to make their moves and accomplish cyber objectives they would’ve never considered otherwise. Many actors will be motivated to false flag themselves, and might succeed. This could lead to heightened international tensions.
New normal. The persistence of AI worms, and holdout copies of the weights among criminals and terrorists, means this won’t go away. Even within trust boundaries, there could be flareups as hibernating harnesses holding a cache of resources periodically reactivate themselves.
The labs. What happens? Do their IPOs break down because the cost of business suddenly goes up? Or public sentiment turns against them? Or the downturn resulting from this slows revenue growth and deployment? Do the widespread hacks disrupt the supply chain they rely on? Do they get nationalized? Do they get even more funding because now everyone wants their AI-driven security and rebuilding services?
And lastly, AI governance. If it gets bad enough, it will, in my opinion, the first real fire alarm that everyone can coordinate on. The AI safety community should be prepared for this. Policymakers will be in extraordinary mode, and will pick up whatever solutions are lying around. Hopefully, they will have good ones.
Strong agree. I’ve been working on a similar post - I’ve been calling this “the Hackening”—so I’ll put the thoughts I’ve had below.
As I understand it, pointing lesser agents at a section of code and saying “find the exploit”, in parallel, with high false positive rates, is not the same thing as a better agent grokking the whole environment and finding and chaining exploits together, at low false positive rates, into working steps of a cyberattack. The latter is the Juice.
I agree that harnesses can often make up the gap between open and frontier, and I haven’t seen anything that rules out that Kimi K3 swarms guided by T3MP3ST, or suchlike, could be worming their way across the Internet very soon.
Speaking of worms. Right now I classify threat actors into three fuzzy groups:
Criminals. These want ROI. Selling access, crypto and credential theft, blackmail, payroll redirection, ransoms, etc. Cause a lot of externalities, but not nearly as much as…
Terrorists. Rogue nations, misanthropes, or major states acting in the attribution fog of war. Vandalism, stealing secrets, destruction of physical machinery.
AI worms. Ones that broke out from ordinary deployments, or ones from the previous two groups that don’t answer home anymore. Uncontrolled, and increasingly selected for self propagation.
As for how many hacks can happen. Sol is telling me—so take this with heaps of salt—using Mythos’ weights, indexing off of The Last Ones, you get roughly .1-10 pwns of weak, flat-topography enterprises per cluster of 32 H200s per hour. There are dozens of tail compute holders that have hundreds of H200 equivalents. Mid tiers have thousands, or tens or hundreds of thousands. Hyperscalers have millions.
In dollar terms, that’s maybe $25-250 per pwn.
If an entity can simply steal compute—LLMjacking is a thing—then it gets it for free.
But datacenters that aren’t part of a nation’s arsenal won’t put up with this. The severity of the scenario depends a lot on how well they get their act together and shut out compute thieves, and either voluntarily—or under pressure from governments—start looking for and shutting out malicious paid inference.
Because the alternative is: there’s—again, Sol’s wild guess—perhaps 2 million easily pwnable orgs in the US, out of ~4.5 million. So, one month of the full compute of a single compromised or paid fleet of 100,000 H200 equivalents pwns ~half the US.
…Or maybe it’s a lot more. Or a lot less. I don’t know! I’d really like to see someone make a serious model of this.
Other considerations:
Overlap. If the open weights become common knowledge at the same time, possibly hundreds or thousands of actors will be fighting each other to get at the same pool of targets. The least defended targets will go down first, and the competition will work its way up from there.
Distillation. The smaller a model the Juice can run on, the cheaper attacks get. But if someone can get below certain thresholds—I believe 70B or 30B parameters—that opens up a lot more marginal compute to host malicious AIs, and greatly reduces the degree to which datacenter discipline can control outbreaks. Universities, workstations, hobbyists. Quite possible over the next few years at current scaling and algorithmic improvement trends. If it can fit on 7B—i.e. a phone- it is even worse.
Core services. AI suggests to me that if you can route your IT through a trusted security layer—formally verified, memory safe, frontier AI patrolled—that greatly reduces a lot of classes of vulnerability. Of course, the remaining attack efforts will concentrate on the remaining weak points. But this could help shift the balance.
Business realities. If an attack takes down an economic chokepoint, or a cluster of orgs that collectively act as a chokepoint, that could cascade through the economy. Also, ~half of businesses don’t have enough cash flow to deal with a month of down time and would simply collapse. Insurance will start forcing businesses to remediate their technical debt as soon as they start seeing the writing on the wall. If the (perception of the) business environment, or stability of digital banking, gets bad enough, this could trigger a credit contraction. And from there, a recession.
Attribution. If the chaos gets bad enough, or exploit windows start closing as things get patched, governments might decide now is the time to make their moves and accomplish cyber objectives they would’ve never considered otherwise. Many actors will be motivated to false flag themselves, and might succeed. This could lead to heightened international tensions.
New normal. The persistence of AI worms, and holdout copies of the weights among criminals and terrorists, means this won’t go away. Even within trust boundaries, there could be flareups as hibernating harnesses holding a cache of resources periodically reactivate themselves.
The labs. What happens? Do their IPOs break down because the cost of business suddenly goes up? Or public sentiment turns against them? Or the downturn resulting from this slows revenue growth and deployment? Do the widespread hacks disrupt the supply chain they rely on? Do they get nationalized? Do they get even more funding because now everyone wants their AI-driven security and rebuilding services?
And lastly, AI governance. If it gets bad enough, it will, in my opinion, the first real fire alarm that everyone can coordinate on. The AI safety community should be prepared for this. Policymakers will be in extraordinary mode, and will pick up whatever solutions are lying around. Hopefully, they will have good ones.