Seconded as per my comment here (I gave July 2027 there and I think by EOY 2026 is a bit early). I predict further that it will take the shape of a kind of computer worm that opportunistically steals compute wherever it can. I know, for example, that plenty of universities got a bunch of capable GPUs scattered around that nobody’s really using, and that aren’t really secured or monitored either. I’m not even talking about “only” high-end consumer GPUs; I know of at least one case of an A100.
“Despite our best effort” I would soften a bit. Despite some reasonable effort, of the kind that you usually see with such things. Because I doubt we’ll see a best effort. It’s gonna be mostly a curiosity; most will laugh at it and move on.
I predict further that it will take the shape of a kind of computer worm that opportunistically steals compute wherever it can.
I predict that the first version will actually target inference API keys, and the inference will happen on hosted frozen models with maybe a LoRA thrown on top (there are a number of hosting services that are bring-your-own-LoRA). Perhaps that is the difference in timeline we expect—I don’t expect a self-replicator who spins up vllm on each new box, at least not as the first replicator. But I also don’t think the weights are the important bit for the replicator—I think the scaffold/prompts/context/memories are the genetic-code-analogue, rather than the weights. At some point I expect we see the vllm one but not until later and I really don’t think that’ll be a significant development at all from a practical pov.
“Despite our best effort” I would soften a bit
Ok fair. “Despite a lot of handwringing and some token efforts, as well as lots of legislation which could not possibly help with stopping the replicator but which does advance the proposing politicians’ pet causes or those of their donors”.
Such activity requires feeding lots of cyber-related prompts (including offensive) to powerful models (at least Kimi K3-level). Such a possibility would be very useful to human cybercriminals, why would API providers allow this without KYC?
Because some people who own GPUs are outside of the jurisdiction of the United States, and are willing to sell access to their GPUs in exchange for cryptocurrency.
The reason consumer GPUs and even a single A100 might be scattered around without use is that one can’t inference any useful coding agents on them, making the scenario you suggest impossible.
If they were to appear (very improbable by the end of this year and unlikely even next year), this hardware would become much more valuable both for its legitimate owners and for cybercriminals.
It’s quite obvious that professional (human) cybercriminals with advanced agents inferenced on large clusters (say, 8xH200s) will exploit such hardware much earlier and more effectively than AIs, making such compute basically unavailable to the latter
Recent work on how AI Agents Enable Adaptive Computer Worms has used a single A100 and demonstrated that a publicly available open-weight model running on such hardware is capable of taking over various machines inside a (constructed) network.
So the scenario I suggest isn’t impossible, it has been proven feasible in a lab setting. Granted, that setting serves only as a proof-of-concept, and the network only had hosts which were deliberately vulnerable in various ways and undefended (but vulnerabilities included e.g. copy fail and dirty frag, discovered after the model’s training cutoff date).
Your point that human cybercriminals are also interested and form some “healthy competition” stands, but someone of that group will also get the bright idea to build and release such a worm.
Seconded as per my comment here (I gave July 2027 there and I think by EOY 2026 is a bit early). I predict further that it will take the shape of a kind of computer worm that opportunistically steals compute wherever it can. I know, for example, that plenty of universities got a bunch of capable GPUs scattered around that nobody’s really using, and that aren’t really secured or monitored either. I’m not even talking about “only” high-end consumer GPUs; I know of at least one case of an A100.
“Despite our best effort” I would soften a bit. Despite some reasonable effort, of the kind that you usually see with such things. Because I doubt we’ll see a best effort. It’s gonna be mostly a curiosity; most will laugh at it and move on.
I predict that the first version will actually target inference API keys, and the inference will happen on hosted frozen models with maybe a LoRA thrown on top (there are a number of hosting services that are bring-your-own-LoRA). Perhaps that is the difference in timeline we expect—I don’t expect a self-replicator who spins up vllm on each new box, at least not as the first replicator. But I also don’t think the weights are the important bit for the replicator—I think the scaffold/prompts/context/memories are the genetic-code-analogue, rather than the weights. At some point I expect we see the vllm one but not until later and I really don’t think that’ll be a significant development at all from a practical pov.
Ok fair. “Despite a lot of handwringing and some token efforts, as well as lots of legislation which could not possibly help with stopping the replicator but which does advance the proposing politicians’ pet causes or those of their donors”.
Hosted where?
Such activity requires feeding lots of cyber-related prompts (including offensive) to powerful models (at least Kimi K3-level). Such a possibility would be very useful to human cybercriminals, why would API providers allow this without KYC?
Because some people who own GPUs are outside of the jurisdiction of the United States, and are willing to sell access to their GPUs in exchange for cryptocurrency.
If it’s something DSV4 Flash sized it could survive on 256 GB RAM regular CPU based servers / workstations.
The reason consumer GPUs and even a single A100 might be scattered around without use is that one can’t inference any useful coding agents on them, making the scenario you suggest impossible.
If they were to appear (very improbable by the end of this year and unlikely even next year), this hardware would become much more valuable both for its legitimate owners and for cybercriminals.
It’s quite obvious that professional (human) cybercriminals with advanced agents inferenced on large clusters (say, 8xH200s) will exploit such hardware much earlier and more effectively than AIs, making such compute basically unavailable to the latter
Recent work on how AI Agents Enable Adaptive Computer Worms has used a single A100 and demonstrated that a publicly available open-weight model running on such hardware is capable of taking over various machines inside a (constructed) network.
So the scenario I suggest isn’t impossible, it has been proven feasible in a lab setting. Granted, that setting serves only as a proof-of-concept, and the network only had hosts which were deliberately vulnerable in various ways and undefended (but vulnerabilities included e.g. copy fail and dirty frag, discovered after the model’s training cutoff date).
Your point that human cybercriminals are also interested and form some “healthy competition” stands, but someone of that group will also get the bright idea to build and release such a worm.