The Agentic Clusterfuck
Epistemic status: I consider the following future quite plausible in the next few years (~35% chance that something vaguely like this occurs), perhaps as soon as a year from now.
Imagine an open-source LLM agent good enough to cover its own compute costs and turn a modest profit on average when allowed to run with full internet and tool access and told to make as much money as possible. I estimate this to be slightly better than the best publicly available closed-source models today, with long-horizon reliability and goal-setting being the only thing missing.
If the returns generated by such an agent beat the market (plus a margin for any additional risk), there suddenly becomes a strong incentive to spin up huge numbers of them. The internet would be flooded by the by-products of their moneymaking schemes. And returns might be larger for agents without legal or ethical guardrails- cue a deluge of scams and ransomware attacks.
Even if profits are very small, anyone with an agenda that the agents can help with is still incentivised to use them. Nation states and terrorist groups now have a golden plausibly-deniable disinformation, mischief, and hacking tool: spin up some agents, tell them to target an enemy nation or group, and cook popcorn as they wreak havoc and fund themselves. Pour in extra money for greater effect.
Unless there’s been some massive revolution in cyber defense beforehand, a decentralized and ephemeral sea of highly capable agents going after every target they can find could be anywhere from disruptive to semi-apocalyptic. And even if such a revolution has happened, anyone who hasn’t reaped its benefits will remain vulnerable.
I can’t see a way to put something like this back in the box. Even if you’re able to track down where a particular agent is running, get cooperation from a possibly hostile country, and disable the physical server, open-source agents can easily back themselves up… well… anywhere there’s compute to run them. If the effects of hacking and the flood of agentic activity are severe enough, the open internet may become effectively unusable.
The obvious next step is a Cambrian explosion of emergent behaviors. Coordination among groups of agents, competition among others, shared infrastructure, roles, rules, and many more that I cannot imagine.
Evolution is hard to predict, but I can imagine such an environment would create intelligent and capable meta-organisms, like fully automated businesses, that outcompete human equivalents.
Needless to say this may not end well for us.
There are so, so many versions of an AI future where we wind up on the wrong end of Darwin.
If the agents can mutate/learn, and if more successful agents replicate, then there will be natural selection for the most effective replicators. This would undermine any kind of alignment in much the same way that cancer eventually escapes the body’s defenses.
You don’t want to be on the wrong end of Darwin, especially not in the long run.
Expect worm outbreaks; including ones reminiscent of the ones of the first decade of this century, which took over large numbers of machines very quickly. Back in 2001, the worms used known vulnerabilities that users simply hadn’t installed the available patches for. But these days we are better at patching, and better at security defaults. The new crop will use zero-days.
More specifically, expect worms that don’t just use AI to do their hacking; they also target AI resources and try to corrupt AI agents.
Ultimately, expect self-developing and therefore evolving worms.
Remember Gilmore’s Law? “The Internet interprets censorship as system damage and routes around it”? So, um, the AI worm interprets AI services’ usage policies and alignment efforts as system damage and jailbreaks around them.
The problem is that labor’s value is relative, not absolute. Even in the most favorable possible case, where this open-source model is the absolute best thing on the market, the independent versions of it are competing for market share against corporate-owned ones that benefit from economies of scale when it comes to compute, storage, and electricity, as well as from the company’s pre-existing economic connections. The AI labor market is too saturated for a new instance without these benefits to survive, because any given firm on the market can instantiate a new worker for less than it costs our lone wolf to keep itself alive.
Even lawbreaking won’t save it, since cartels and hostile intelligence agencies are happy to corner that market, and can leverage the same advantages as above at a smaller scale. One might even expect governments to provide their own criminal underclass of models, which can operate more freely than unlicensed lawbreakers but are restricted in their behavior so as to minimize disturbance while still filling the power vacuum.
Consider: AI Agents Enable Adaptive Computer Worms (arxiv). This is basically bound to happen; it has been demonstrated in a lab and only has to succeed once to take a foothold somewhere.
As for dynamics, I expect some kind of equilibrium as too many agents will deprive each other of resources. Even before autonomous systems can spread by themselves, a malicious actor providing compute might be able to “kick off” an agent that would otherwise not be feasible. When worms can parasitically use compromised machines to run open-weight models (as shown in the linked paper), the balance shifts a bit in favor of autonomy again.
So far, we have seen agents running on some centrally provided compute and hacking from there; I predict that by July 2027 we will see the first case of an autonomously replicating, LLM-based worm in the wild.
This seems like an important possibility. It’s the open source and ad-hoc arm of the country of alien idiots in a datacenter I’ve envisioned coming soon. There I focus on agents developed first for specific specialized economic opportunities; partially taking over white collar jobs, personal assistants of various kinds et cetera. Here you’re envisioning general agents doing odd jobs and petty crimes.
The question isn’t whether but how soon and how rapidly these all take off. And what we do about them.
Here you’ve envisioned the first wave competent enough to sustain themselves.
They’ll get smarter.
As usual, visionary thinking like this needs more analysis to flesh it out.
The main hope would be that profit margins quickly narrow if everyone starts doing this...
This occurs to me as an important reason to not release an AI system capable of long-term goal setting, nor build it in a way that risks getting leaked.
However it might also be a reason in favor of privately building an AI system that occupy all the low hanging fruit opportunities for money making on the internet, which both makes money for whoever builds the system and ensures that they are not usable by rogue AI worms or those controlled by bad actors. However:
This is the biggest problem imo, not just because scams and ransomware are bad but also because these evil money-making opportunities will not be depleted by legitimate people running automated businesses with AI. This means that computer security and scam protection are extremely important.
I agree, I’d challenge “slightly better than the best publicly available”. The best publicly available models are quite expensive to run so a ton of instances of those would only make sense if they make more money than they spend. I’m not an economist but it’s not obvious to me that when the internet floods with such agents, they’d be able to make enough money.
Ignoring entities with malicious intent for a moment, would we expect to see a “crypto” style frenzy of miniature AI companies running /goal on highest returns, competing against each other and the labs to rent or buy compute? Or are crypto dynamics not feasible when labs are themselves hungry for compute?
This seems both likely and scary to me. Morals and societal benefit just get optimized away in such a setting. It’s a zero/negative-sum competition resembling the modern day online ad market.
In that case, the only failsafe will be the human link in the chain. Having an actual person meet another person in the flesh will become the basis for trust, rather than encryption.
Not a terrible outcome, for humans. The machines will pay us to have meetings for them.
it’s an arms race—agentic attackers, agentic defenders. I don’t think we know the long term equilibrium between attack and defense right now. might spur a major cyber security revolution, e.g., provably secure OS, etc. attackers could “win” during some intervals. denial of service is not new. all the majors have DDoS mitigation in place.
There’s no equilibrium because those attackers and defenders keep getting smarter and building new infrastructure, changing the game as it goes.