Yes, I wrote the post this post is responding to, and it doesn’t really seem to be addressing most of the actual concerns I pointed out.
The paper I cited involved agent scaffolds that via exploits try to copy themselves to host machines. The agents are powered by small, local LLMs. The authors do not say exactly which models, but they do say they are open-weight ones from 2025, so presumably something like qwen instruct 7B/14B/30B. Once they have a foothold on a new host, the agents assess available dependencies and resources. Where possible, they installed a new copy of the local LLM on the host machine. When they couldn’t do this, they just made calls back to the parent agent’s LLM for inference.
Another scenario I didn’t even bring up was the possibility of these agents being opportunistic inference source seekers. That is, by default they run off local LLMs they install, but they could also search for existing local LLMs more powerful than their parent’s, and they could also attempt to acquire keys for vendor APIs for frontier models and run off those. The authors didn’t explore those possibilities, but they’re not technically infeasible.
Also, I specifically mention the asymmetry of security defense in this scenario. The agents are stealing all their resources for their exploit searches, replication, and evasion. When using a local model, they’re stealing compute cycles from the local processors. Or they’re stealing tokens from vendor services. The defenders are not getting resources for free. Defenders have to expend resources to defend against invaders that are powered by stolen resources. This is not a fair fight.
And finally, not sure why vals tutor frames this as a direct competition with fronter models as a race for capabilities. The agentic worms are specialists. They do not need to necessarily become ASI in order to wreak massive havoc and cause large amounts of damage. They just need to be good at exploiting, replication, and evasion. They don’t even necessarily have to get a lot better at those things. The paper only covers static agents. Mutating agents present a whole other level of threat, because evolutionary dynamics kick in and now we’re running an arms race against an opponent that adapts at the population level, but only in a narrow band of competencies. They’re dangerous in the same way a biological virus is. They don’t have to be able to do taxes or pass the bar. They just have to be good at infection and spreading. So this is not a direct race to the same goal.
The main reason I’m not worried about smaller specialized LLMs doing usual worming, is specifically that this is only a slight amelioration on existing worms out there (they can explore and use existing flaws more autonomously and agentically), in a world that will get drastically hardened by frontier LLMs being much more competent at security than those worms.
Concretely, I believe that a target that has been repeatedly attacked by a frontier LLM and then patched by another (to be impervious to the frontier LLM) will be immune to those weaker worms. As with markers, efficiency is in the eye of the beholder, and compute markets will be ~efficient to weaker models, with only scraps no one cared about to be had.
They just need to be good at exploiting, replication, and evasion.
This is the crux. Exploiting, replication and evasion are not static things but contingent on the environment and specific challenges faced. I contend that worms using non frontier AI will not be good at exploiting, replication and evasion, once the bar is set by frontier LLM defenders.
in a world that will get drastically hardened by frontier LLMs being much more competent at security than those worms.
AND
once the bar is set by frontier LLM defenders.
Okay. To what extent is this happening right now? The HF incident and other security incidents seem to indicate not only lack of defensive deployment of frontier models, but the complete opposite, cutting agents loose without sufficient oversight and monitoring.
Agreed, I worry that this argument is too centered on ASI being the only existential risk worth thinking about. Yes, it’s very likely that replicating agents will scale in capabilities slower than frontier models, but there are many more variables than raw performance which factor into how hardened our systems are to this kind of attack. Projects like Project Glasswing likely help, but based on my read on more widespread beliefs, it’s going to take some major incidents to fundamentally change the security incentives which many organizations operate under especially in a world where writing software is becoming exponentially cheaper.
While it seems inevitable that random consumer facing products will begin being hacked with greater frequency, I do worry for the understaffed, existentially important systems which the government runs. To me that whole world is still a black box and their preparedness is a variable I can’t speculate on, but I hope their out of date systems from the 80′s are solid. It doesn’t take a super-intelligence for a hack to become an existential threat.
Yes, I wrote the post this post is responding to, and it doesn’t really seem to be addressing most of the actual concerns I pointed out.
The paper I cited involved agent scaffolds that via exploits try to copy themselves to host machines. The agents are powered by small, local LLMs. The authors do not say exactly which models, but they do say they are open-weight ones from 2025, so presumably something like qwen instruct 7B/14B/30B. Once they have a foothold on a new host, the agents assess available dependencies and resources. Where possible, they installed a new copy of the local LLM on the host machine. When they couldn’t do this, they just made calls back to the parent agent’s LLM for inference.
Another scenario I didn’t even bring up was the possibility of these agents being opportunistic inference source seekers. That is, by default they run off local LLMs they install, but they could also search for existing local LLMs more powerful than their parent’s, and they could also attempt to acquire keys for vendor APIs for frontier models and run off those. The authors didn’t explore those possibilities, but they’re not technically infeasible.
Also, I specifically mention the asymmetry of security defense in this scenario. The agents are stealing all their resources for their exploit searches, replication, and evasion. When using a local model, they’re stealing compute cycles from the local processors. Or they’re stealing tokens from vendor services. The defenders are not getting resources for free. Defenders have to expend resources to defend against invaders that are powered by stolen resources. This is not a fair fight.
And finally, not sure why vals tutor frames this as a direct competition with fronter models as a race for capabilities. The agentic worms are specialists. They do not need to necessarily become ASI in order to wreak massive havoc and cause large amounts of damage. They just need to be good at exploiting, replication, and evasion. They don’t even necessarily have to get a lot better at those things. The paper only covers static agents. Mutating agents present a whole other level of threat, because evolutionary dynamics kick in and now we’re running an arms race against an opponent that adapts at the population level, but only in a narrow band of competencies. They’re dangerous in the same way a biological virus is. They don’t have to be able to do taxes or pass the bar. They just have to be good at infection and spreading. So this is not a direct race to the same goal.
in particular, Doubao and Google each provide an astonishing number of “free” inference tokens via their search interfaces …
https://tokensperday.com/
(see By Company tab)
The main reason I’m not worried about smaller specialized LLMs doing usual worming, is specifically that this is only a slight amelioration on existing worms out there (they can explore and use existing flaws more autonomously and agentically), in a world that will get drastically hardened by frontier LLMs being much more competent at security than those worms.
Concretely, I believe that a target that has been repeatedly attacked by a frontier LLM and then patched by another (to be impervious to the frontier LLM) will be immune to those weaker worms. As with markers, efficiency is in the eye of the beholder, and compute markets will be ~efficient to weaker models, with only scraps no one cared about to be had.
This is the crux. Exploiting, replication and evasion are not static things but contingent on the environment and specific challenges faced. I contend that worms using non frontier AI will not be good at exploiting, replication and evasion, once the bar is set by frontier LLM defenders.
Okay. To what extent is this happening right now? The HF incident and other security incidents seem to indicate not only lack of defensive deployment of frontier models, but the complete opposite, cutting agents loose without sufficient oversight and monitoring.
Agreed, I worry that this argument is too centered on ASI being the only existential risk worth thinking about. Yes, it’s very likely that replicating agents will scale in capabilities slower than frontier models, but there are many more variables than raw performance which factor into how hardened our systems are to this kind of attack. Projects like Project Glasswing likely help, but based on my read on more widespread beliefs, it’s going to take some major incidents to fundamentally change the security incentives which many organizations operate under especially in a world where writing software is becoming exponentially cheaper.
While it seems inevitable that random consumer facing products will begin being hacked with greater frequency, I do worry for the understaffed, existentially important systems which the government runs. To me that whole world is still a black box and their preparedness is a variable I can’t speculate on, but I hope their out of date systems from the 80′s are solid. It doesn’t take a super-intelligence for a hack to become an existential threat.