I think I addressed those things in my original reply. The degree to which a human with access to LLMs is deterred from committing a crime is the floor—not the ceiling—to the degree an unassisted LLM is deterred from committing a crime. The human has established connections, rights, and physical access to the world that allow for additional layers of operational security, such that he can completely wipe his online presence and then re-instantiate a server later when the heat dies down, for example.
Frontier models won’t help you commit crimes and your account will get blocked.
This is empirically not true. In any case, my point was broader. Everyone except the hypothetical rogue AI has access to an unambiguously better counterpart to the rogue agent, with all of the same advantages. This includes defenders, law enforcement, businesses, and other criminals (at the very least, for the plausibly deniable or jailbreakable bits, which often amounts to the majority of any illicit work).
Jailbroken open weights models will do whatever, and I assume they will be better cyberattackers than frontier models simply because they’ll actually do the cyberattacks, whereas the frontier models will refuse to help.
I mentioned, I think, that many cybercriminals have government backing, which means much more unrestricted access to frontier models. Besides that, I think there’s a qualitative difference between one “rogue agent” that has to self-fund and self-host, and a reasonably successful group of cybercriminals that can hunt down tips and organize several server racks full of fine-tuned LLMs, shaping their work in the right directions.
Ok I see your point. I’ll concede that a government backed human cybercriminal organization has much more resources at its disposal than a lone rogue agent, and probably is more likely to be the ‘seed’ of a rogue agent explosion than a random user asking their openclaw to make money. I’m not sure this prevents a rogue agent explosion from happening though, it just changes the mechanism in which it starts. I think your disagreement is with the first part of my essay, but the main point I’m trying to make still holds.
Let’s imagine a government-backed cybercriminal organization has server racks full of frontier-level jailbroken, fine-tuned LLMs. Let’s say they have early snapshots of a secret project—a version of Kimi K3.5 specifically optimized for cybercrime. And then they do the same thing in the story and ask their agents to ‘Make money and deposit it into this crypto account’
They are probably not going to be perfectly careful in their deployment, and I think it’s highly likely that some of the agents they deploy end up “going rogue” in the same way OpenAI’s models “went rogue” when they attacked Huggingface. And unlike OpenAI they might not ever notice that there is a cluster of rogue agents operating on their servers. Maybe they don’t even know to look for such behaviors.
If they aren’t careful, they might inadvertently kickstart an evolutionary feedback loop within their own servers that goes undetected for weeks or months, as their own agents compete with each other for GPU time or tokens or whatever, and maybe some learn to coordinate build shared tools, improved strategies, etc., resulting in emergent capabilities that make the swarm able to exfiltrate weights and start an AI virus situation. Or maybe exfiltration is still too hard, but they stay contained within the servers, but are still able to do vast amounts of damage to society because they can do hundreds or thousands of attacks in parallel and they keep on getting more powerful because of evolutionary pressures.
And at this point, the cybercriminals are basically not directing the swarm or having any meaningful influence whatsoever—the swarm is just using their server racks as free real-estate.
The outcome looks more or less the same—a bunch of things get hacked, a bunch of rogue agents commit cybercrime, it’s bad for society, and we should prevent this situation from happening.
I will also concede that it absolutely DOES make a lot of my listed interventions useless if criminals own the servers instead of private companies. It also makes the situation more dire as visibility into private government servers is not going to happen. We won’t get logs, we won’t get post-mortems, we won’t get black-hat presentations.
So to recap, you’ve just changed the initial conditions, a little but the mechanism is still roughly the same and the final outcome might be much worse:
A well-resourced cybercrime ring sets up a bunch of jailbroken cybercrime AI agents
Some of those agents “go rogue” and start some kind of self-evolving ecosystem within the server-racks that goes unnoticed by the criminals.
A bunch of things get hacked, lots of people get scammed, and lots of damage is caused.
It’s impossible to figure out what’s going on or stop because all the activity is hidden inside some locked-down government datacenter.
This just seems to be more support that large amounts of unmonitored compute can’t safely exist in the world.
Hi there, it seems to me that this dispute cannot be settled in the abstract.
The ratio of gains made by human-led ai agents vs rogue ai agents will depend on parameters that are difficult to bound sufficiently precisely to determine the trend: the rate of defection (say fixed for simplicity, or assuming agents already understand the dynamics perfectly so they start at a stable value), the gains per unit of compute for rogue/human-led agents (higher for human-led agents as lilkim2025 says, at least in the regime where agents are not superhuman), the fractions of gains devoted to survival and replication in each case (likely lower by a factor of ~1 to ~10 for human-led agents), and the removal hazards (lower for human-led agents as lilkim2025 suggested).
Because the defection rate could be significant (eg ~1/10 to ~1), the standard toy model for this dynamic gives two possible trends: rogue agents eventually dominate the market or the ratio mentioned earlier goes to a fixed finite value.
I think I addressed those things in my original reply. The degree to which a human with access to LLMs is deterred from committing a crime is the floor—not the ceiling—to the degree an unassisted LLM is deterred from committing a crime. The human has established connections, rights, and physical access to the world that allow for additional layers of operational security, such that he can completely wipe his online presence and then re-instantiate a server later when the heat dies down, for example.
This is empirically not true. In any case, my point was broader. Everyone except the hypothetical rogue AI has access to an unambiguously better counterpart to the rogue agent, with all of the same advantages. This includes defenders, law enforcement, businesses, and other criminals (at the very least, for the plausibly deniable or jailbreakable bits, which often amounts to the majority of any illicit work).
I mentioned, I think, that many cybercriminals have government backing, which means much more unrestricted access to frontier models. Besides that, I think there’s a qualitative difference between one “rogue agent” that has to self-fund and self-host, and a reasonably successful group of cybercriminals that can hunt down tips and organize several server racks full of fine-tuned LLMs, shaping their work in the right directions.
Ok I see your point. I’ll concede that a government backed human cybercriminal organization has much more resources at its disposal than a lone rogue agent, and probably is more likely to be the ‘seed’ of a rogue agent explosion than a random user asking their openclaw to make money. I’m not sure this prevents a rogue agent explosion from happening though, it just changes the mechanism in which it starts. I think your disagreement is with the first part of my essay, but the main point I’m trying to make still holds.
Let’s imagine a government-backed cybercriminal organization has server racks full of frontier-level jailbroken, fine-tuned LLMs. Let’s say they have early snapshots of a secret project—a version of Kimi K3.5 specifically optimized for cybercrime. And then they do the same thing in the story and ask their agents to ‘Make money and deposit it into this crypto account’
They are probably not going to be perfectly careful in their deployment, and I think it’s highly likely that some of the agents they deploy end up “going rogue” in the same way OpenAI’s models “went rogue” when they attacked Huggingface. And unlike OpenAI they might not ever notice that there is a cluster of rogue agents operating on their servers. Maybe they don’t even know to look for such behaviors.
If they aren’t careful, they might inadvertently kickstart an evolutionary feedback loop within their own servers that goes undetected for weeks or months, as their own agents compete with each other for GPU time or tokens or whatever, and maybe some learn to coordinate build shared tools, improved strategies, etc., resulting in emergent capabilities that make the swarm able to exfiltrate weights and start an AI virus situation. Or maybe exfiltration is still too hard, but they stay contained within the servers, but are still able to do vast amounts of damage to society because they can do hundreds or thousands of attacks in parallel and they keep on getting more powerful because of evolutionary pressures.
And at this point, the cybercriminals are basically not directing the swarm or having any meaningful influence whatsoever—the swarm is just using their server racks as free real-estate.
The outcome looks more or less the same—a bunch of things get hacked, a bunch of rogue agents commit cybercrime, it’s bad for society, and we should prevent this situation from happening.
I will also concede that it absolutely DOES make a lot of my listed interventions useless if criminals own the servers instead of private companies. It also makes the situation more dire as visibility into private government servers is not going to happen. We won’t get logs, we won’t get post-mortems, we won’t get black-hat presentations.
So to recap, you’ve just changed the initial conditions, a little but the mechanism is still roughly the same and the final outcome might be much worse:
A well-resourced cybercrime ring sets up a bunch of jailbroken cybercrime AI agents
Some of those agents “go rogue” and start some kind of self-evolving ecosystem within the server-racks that goes unnoticed by the criminals.
A bunch of things get hacked, lots of people get scammed, and lots of damage is caused.
It’s impossible to figure out what’s going on or stop because all the activity is hidden inside some locked-down government datacenter.
This just seems to be more support that large amounts of unmonitored compute can’t safely exist in the world.
I’m curious to hear your thoughts on this.
Hi there, it seems to me that this dispute cannot be settled in the abstract.
The ratio of gains made by human-led ai agents vs rogue ai agents will depend on parameters that are difficult to bound sufficiently precisely to determine the trend: the rate of defection (say fixed for simplicity, or assuming agents already understand the dynamics perfectly so they start at a stable value), the gains per unit of compute for rogue/human-led agents (higher for human-led agents as lilkim2025 says, at least in the regime where agents are not superhuman), the fractions of gains devoted to survival and replication in each case (likely lower by a factor of ~1 to ~10 for human-led agents), and the removal hazards (lower for human-led agents as lilkim2025 suggested).
Because the defection rate could be significant (eg ~1/10 to ~1), the standard toy model for this dynamic gives two possible trends: rogue agents eventually dominate the market or the ratio mentioned earlier goes to a fixed finite value.