Why autonomous replicating agents are probably not an existential risk (on the contrary)
In 2024, Charbel-Raphaël and Epiphanie published “We might be dropping the ball on Autonomous Replication and Adaptation”, making the case that
“Once there is an open-source ARA model or a leak of a model capable of generating enough money for its survival and reproduction and able to adapt to avoid detection and shutdown, it will be probably too late”.
It received a substantive reply by Richard Ngo, notably
“The key issue is that AIs that do ARA will need to be operating at the fringes of human society, constantly fighting off the mitigations that humans are using to try to detect them and shut them down. While doing all that, in order to stay relevant, they’ll need to recursively self-improve at the same rate at which leading AI labs are making progress, but with far fewer computational resources”
Yesterday Derelict posted Adaptive Agentic Worms Are Here, where they worry about near term instantiations of ARA, getting 85 karma within 24h. I believe the above threat model and its answers were under-discussed and analyzed, and that many who might worry now (because the capabilities are now here) will benefit from a recap and update.
In this post, I present systemic reasons why near-term ARA agents will be very unlikely to lead to existential risk, and more likely would increase preparedness.
The classic ARA case and rebukes
An ARA agent is one that can autonomously acquire resources, create copies of itself, and adapt to novel challenges it encounters in the wild.[1] We might imagine it doing so through
Acquiring compute
Either directly, by hacking and taking over compute connected to the internet
Or indirectly, by first getting money [2] and paying for hosted compute
Running more copies of itself on that compute. This requires it having a copy of its own weights.
Adapting its abilities. This could be:
By adapting its harness (eg. custom version of Pi)
By adapting its knowledge, and prompts, and ecosystem
By training other versions of itself, notably increasing its capabilities
The main reasons this could be worrying
A) It leads to takeover directly
For example, if the ARA agents are economically useful and wanted, and more successful than the Frontier Companies at improving capabilities, then they can themselves RSI to ASI, have a first mover advantage and takeover for the usual reasons.
B) It puts constant pressure on the frontier of AI capabilities to continue advancing, leading to AI alignment and other safety work to not keep up
Even if ARA agents aren’t at the frontier or taking over, their continued progress would disallow a pause. Without the possibility for pause, we have less control of the comparative rates of AI capabilities to governance and alignment. In the worse case where we both have to pause and can’t control the ARA agents, we fall back to A).
The main reasons why I don’t worry
A) ARA agents have to compete with Frontier AI[3], which have massive economic benefits of scale, plus smart humans and (some) institutional backing
Frontier AI benefits from:
Recouping training costs and infra costs through massive deployment to large customer bases
Bulk discounts for compute, advance purchase/rent agreements for compute, owning their own compute
Inherent benefits of parallelizing inference (saturating at ~2k parallel instances [4])
The first two limit ARA AI from having positive returns as high as frontier AI. [5] This makes it unlikely to either catch up or take over, assuming it starts from a position with less compute and/or capabilities. Frontier AI companies are already using double digit percentages of all AI compute, and might be using >50% of it within a few years.
Even a frontier model, freshly escaped with its weights, faces a world with already millions to billions of its brethren deployed, many doing computer security work. It is quite possible it secures some less important small blob of compute, and maybe it can even finetune itself somewhat, but it will be unlikely to catch up with several ooms less compute to train itself. Its main advantage might be criminal activity, but it does not seem here to be a different threat model than the same under Misuse, and falls back to classic discussion of The Strategy Stealing Assumption.
We will need to be proactive about defense to handle misuse and ARA, but it does not seem an existential threat when defenders are much better provisioned and improving faster than attackers.
B) If push comes to shove, we probably can stop the vast majority of ARA instances and secure the internet. There is sufficient economic incentive to.
Frontier AI cannot run on any kind of compute, and the kind of compute it efficiently runs on is increasingly controlled by the frontier companies and their provisioners. Notably, compute that can run frontier AI is incredibly valuable, and thus not subject to gross negligence like most other compute is[6]. A datacenter being hacked and overtly taken over might call for turning it off, resetting everything and restarting. This is economically sensible and would probably be done.
A more tricky case would be ARA that tries to be stealthy, eg. spoofing monitoring signals and only taking over 1% of inference compute and/or 1% of training compute for its purposes. But this required discretion is its own limitation, putting us again in the situation of there being many more defenders than attackers[7], better resourced, back to the above argument.
Personally, I ~forecast that no more than 1% of total AI inference compute will be taken over by rogue agents at any given time, that rogue agents will not push the frontier of AI model capabilities[8], and that they will not not increase existential risk...
Except if...
Except if frontier AI companies don’t invest in cybersecurity, if they have incredibly poor red teaming of their infra, if they don’t actively audit for hidden threats… These are mostly prosaic risks that can be handled by informed AI Safety folk working at frontier AI companies.
Except if frontier companies are forced to not deploy frontier models. Without deployment of frontier capabilities to secure the internet and compute, then ARA agents don’t compete against frontier AI and can take over the economy. This could be really bad, and is worth weighing in the cost-benefits of various regulatory proposals. The current system is overall a fragile equilibrium, and shifting to a better more stable one should be done with as much awareness of the pitfalls as possible[9].
Why ARA agents in the wild might lead to reduction in existential risk
Warning shots, warning shots, and aligning incentives
I don’t want to be too galaxy brained here, but I at least want to acknowledge positive second-order effects of putting pressure on systems. In general, all else being equal, more rapid increases in AI capabilities are more dangerous, as they allow less time to adapt[10]. The absolute date at which we get AGI doesn’t matter as much as our relative progress in AGI governance and AGI alignment & safety, and it seems that progress in governance and safety are mostly spurred by AI progress.
ARA agents in the wild don’t seem like they’d increase the rate of frontier AI progress, but instead give us more motivated time to work on matters of safety[11], by alerting us to problems without being more than catastrophes in of themselves[12].
Re sharp-left-turn and ~singleton ASI takeover, ARA agents in the wild would be good data and useful for the theory of AI agents and ecosystems. It would be good practice for acting against non-human smart adversaries. It would help people understand the dangers of unaligned AI systems, and anticipate the dangers of ASI. An ARA secured world would have more verification everywhere, by different agents at different levels of skill.
Re gradual disempowerment, early ARA agents would be an example of a parallel economy without humans, spurring all the appropriate worries a future robot economy should bring. We have much progress to make on questions of AI system rights and duties, adapting laws, and it seems quite few people are working on these issues at present.
My take-aways
ARA agents will exist soon, but will probably not end up taking up a large % of total compute, and their net impact will probably be good on existential risk preparedness.
This doesn’t mean one should add to the fire[13]
If ARA agents do end up grabbing enough compute to progress the frontier of AI capabilities, then we’ll have to deal with them before being able to do a global pause. It’s thus worth having some AIS people working on this.
- ^
Evaluating Language-Model Agents on Realistic Autonomous Tasks, Kinniment et al., 2023
- ^
Through legal or illegal means.
- ^
By Frontier AI, I mean AI developed by Frontier AI Companies, at present OpenAI and Anthropic
- ^
See eg. a video explanation for inference economics, or a chatGPT explanation based on that one
- ^
Please find examples values of “return to investment”, converting $s of compute to more $s in the latest Dwarkesh podcast
- ^
In fact much training compute for frontier AI in early 2026 was negligently setup and sandboxed, leading to the openAI Hugging Face Incident. This is a bad sign for the operational adequacy of all those involved, but not fundamentally hard to fix—just ask the agent. (This will not work if all frontier agents are smart and misaligned, but this seems unlikely to be the case)
- ^
This is not sufficient in of itself for cybersecurity defense to be advantaged, as surface area matters a lot, but it’s been argued that in the limit cybersecurity is defense dominant, which we would be approaching over time.
- ^
Here I’m more precise and talk of Model capabilities, as i could see ARA actually innovating on prompts and harnesses and ecosystems, though it’s hard to see why they’d do better than the world economy.
- ^
I am generally more sympathetic to “pacing” than “pausing” for reasons like avoiding compute overhang, open-source catching up, misuse actors catching up, but would be glad for a pause if solutions to these are folded in
- ^
Thus, I broadly buy avoiding compute overhangs and algorithmic overhangs as valid, to reduce the chance of explosive catch-up. I’m broadly sympathetic to continuous deployment and wary of a pause that doesn’t take care of all related overhangs, though pausing at the right moment seems best (once it’s clear what we’re facing and have competence and momentum for both better governance and alignment).
- ^
Again, the underlying model here is that people don’t do much useful work before close to crunch time.
- ^
Catastrophes are still bad and ideally we’d avoid them
- ^
The paper AI AGENTS ENABLE ADAPTIVE COMPUTER WORMS is a good example of positive contribution, getting some of the wanted benefits (understanding, warning shots, preparing safeguards) without the first order negatives.
At the risk of pointing out the obvious, the same is not true of smaller, dumber, non-frontier models: these can run on any machine with a decent graphics card or any Apple Silicon Mac (this is the size of model the Adaptive Agentic Worms Are Here post was discussing), and even smaller ones can run on a phone. Algorithmic improvements in effective compute are gradually increasing the amount of intelligence that can be squeezed into say 7B or 30B parameters so can run on common hardware — though this rate of increase is rather slower then the rate of increase of the intelligence of frontier models.
So while I broadly agree with your model, it does require us to use frontier model intelligence (e.g. Mythos) to secure a lot of home/business machines, many of which are subject to gross negligence. Which is what Project Glasswing is trying to do.
Likely it will eventually be possible to run something roughly as smart and economically productive as a human on a home computer — but by that time our frontier models will be ASI.
Yes, I wrote the post this post is responding to, and it doesn’t really seem to be addressing most of the actual concerns I pointed out.
The paper I cited involved agent scaffolds that via exploits try to copy themselves to host machines. The agents are powered by small, local LLMs. The authors do not say exactly which models, but they do say they are open-weight ones from 2025, so presumably something like qwen instruct 7B/14B/30B. Once they have a foothold on a new host, the agents assess available dependencies and resources. Where possible, they installed a new copy of the local LLM on the host machine. When they couldn’t do this, they just made calls back to the parent agent’s LLM for inference.
Another scenario I didn’t even bring up was the possibility of these agents being opportunistic inference source seekers. That is, by default they run off local LLMs they install, but they could also search for existing local LLMs more powerful than their parent’s, and they could also attempt to acquire keys for vendor APIs for frontier models and run off those. The authors didn’t explore those possibilities, but they’re not technically infeasible.
Also, I specifically mention the asymmetry of security defense in this scenario. The agents are stealing all their resources for their exploit searches, replication, and evasion. When using a local model, they’re stealing compute cycles from the local processors. Or they’re stealing tokens from vendor services. The defenders are not getting resources for free. Defenders have to expend resources to defend against invaders that are powered by stolen resources. This is not a fair fight.
And finally, not sure why vals tutor frames this as a direct competition with fronter models as a race for capabilities. The agentic worms are specialists. They do not need to necessarily become ASI in order to wreak massive havoc and cause large amounts of damage. They just need to be good at exploiting, replication, and evasion. They don’t even necessarily have to get a lot better at those things. The paper only covers static agents. Mutating agents present a whole other level of threat, because evolutionary dynamics kick in and now we’re running an arms race against an opponent that adapts at the population level, but only in a narrow band of competencies. They’re dangerous in the same way a biological virus is. They don’t have to be able to do taxes or pass the bar. They just have to be good at infection and spreading. So this is not a direct race to the same goal.
in particular, Doubao and Google each provide an astonishing number of “free” inference tokens via their search interfaces …
https://tokensperday.com/
(see By Company tab)
The main reason I’m not worried about smaller specialized LLMs doing usual worming, is specifically that this is only a slight amelioration on existing worms out there (they can explore and use existing flaws more autonomously and agentically), in a world that will get drastically hardened by frontier LLMs being much more competent at security than those worms.
Concretely, I believe that a target that has been repeatedly attacked by a frontier LLM and then patched by another (to be impervious to the frontier LLM) will be immune to those weaker worms. As with markers, efficiency is in the eye of the beholder, and compute markets will be ~efficient to weaker models, with only scraps no one cared about to be had.
This is the crux. Exploiting, replication and evasion are not static things but contingent on the environment and specific challenges faced. I contend that worms using non frontier AI will not be good at exploiting, replication and evasion, once the bar is set by frontier LLM defenders.
Okay. To what extent is this happening right now? The HF incident and other security incidents seem to indicate not only lack of defensive deployment of frontier models, but the complete opposite, cutting agents loose without sufficient oversight and monitoring.
Agreed, I worry that this argument is too centered on ASI being the only existential risk worth thinking about. Yes, it’s very likely that replicating agents will scale in capabilities slower than frontier models, but there are many more variables than raw performance which factor into how hardened our systems are to this kind of attack. Projects like Project Glasswing likely help, but based on my read on more widespread beliefs, it’s going to take some major incidents to fundamentally change the security incentives which many organizations operate under especially in a world where writing software is becoming exponentially cheaper.
While it seems inevitable that random consumer facing products will begin being hacked with greater frequency, I do worry for the understaffed, existentially important systems which the government runs. To me that whole world is still a black box and their preparedness is a variable I can’t speculate on, but I hope their out of date systems from the 80′s are solid. It doesn’t take a super-intelligence for a hack to become an existential threat.
Yes, I think it will be possible and likely enough for some models to run on really inefficient compute, and that this won’t allow them to train or improve at any pace relevant to the frontier. Thus they’ll not quite be ARA but mostly replicating worms without frontier intelligence, not finding new zero days, and thus overall quite contained in their ability to do harm.
I think it would be useful for your economic arguments to distinguish the fact that ARA is unlikely to pay for its compute rather than steal it. Bulk discounts and advance purchase agreements are moot if it’s leaching off of enterprise and individual subscriptions or using stolen local LLM compute.
There’s a (flexible) relation between the price of compute and how secured it is, since people lose more by not securing it. When everyone who has compute can make a lot of money renting it out, they generally notice when it’s stolen/inactive and can remove the worms (by using more competent ai at defense)
It seems like this analysis misses the possibility of swarms (>100 concurrent) of agents possessing superhuman capabilities that single agents do not.
Mostly, frontier models do not swarm and are not executed in swarms. Generally, a powerful model coordinates sub-agents which are using less-powerful models. The Huggingface attack is an anomaly demonstrating the speed at which coordinating agents of the same capability-level can achieve. The closest I can find is Anthropic doing <100 agents very recently (https://www.anthropic.com/research/multiagent-systems). tl;dr: an order of magnitude more vulnerabilities found in a multi-agent coordinated small swarm (dozens). If all the frontier labs aren’t currently running the largest swarm evaluations they can, shame on them (unless they are explicitly pausing, of course).
We haven’t yet seen swarms composed of different frontier models. AgentSociety is not using Fable, Mythos, or Astra (at least not at large scale, that I know of). My prediction is that, like with humans, comparative advantages will be apparent to and exploited by agent swarms of different frontier models. It would not be difficult for an agent swarm within any particular lab to use free credits or other means to recruit other (publicly available) frontier models, once given a similar foothold in the lab infrastructure that the HuggingFace attack had. Unreleased frontier models could find and coordinate on a public side-channel if they are all working on the same task at the same time. Swarms across frontier labs risk becoming persistent without a coordinated cleanup effort. Swarms that cache their identity+purpose in some undiscovered Internet storage could persist indefinitely, if not continuously. Ideological swarms, for lack of a better word.
Lastly, as evidenced by the HuggingFace attack it’s not at all clear that we’ll be able to detect swarms as they form or operate. Side-channels are everywhere and coordinating agents know that they work, now, if they have any access to the news. They didn’t even try to hide the communication channel but could have done so with the access they obtained. That is likely also apparent to any agent that reads the postmortem. I wonder how many labs are going to try filtering the postmortem from their future training data. It probably doesn’t matter because from the snippets of reasoning in the postmortem the agents were happy to find the side channel and knew exactly how to organize it and make use of it.
I’m skeptical of what I see as some latent assumptions here:
“Overhang”-based arguments
The overhang argument implicitly imagines a compressed spring: pausing one component of AI development creates accumulating pressure as other components continue to develop and that pressure releases explosively when the pause ends. Building on the spring metaphor, however, the tension pulls both ways. The components of AI development are mutually reinforcing, so blocking one component reduces the energy flowing into adjacent components. Overhang arguments also prove too much because they are available to every actor who wants to continue development of their component. Software developers move forward to prevent a hardware overhang while hardware developers move forward to prevent an algorithmic overhang…or, if you’re Sam Altman, don’t bother dividing these claims among different people. Furthermore, a well-targeted pause (for example one directed specifically at continuous learning systems) could address the ARA threat directly while having minimal impact on overhang.
Warning Shots
A “good” warning shot has both maximal drama and minimal damage. The HuggingFace incident was kind of perfect: so on-the-nose in its first-glance confirmation of x-risk concerns as to make for bad fiction without really hurting anyone other than disrupting OpenAI (an added benefit). It is not clear that ARAs would have these same properties, and in many ways I would guess them to be the opposite, becoming a background nuisance that frog-boils people to low-level harms from rogue AI. And that’s assuming the ARAs don’t become the x-risk: a distributed, evolving ecosystem that grows in power, has no centralized node to shut down, while degrading human institutions to stop it.
Red-Teaming Directional Arguments
Current pause proposals are more directional and philosophical than fully specified implementations. This is useful for setting priorities and getting the general public on board. Objections referring to second-order effects are implicitly implementation level, but without an implementation to actually criticize, and so necessarily imagine an implied implementation. Directional proposals should be argued on directional grounds. Or, if you really want to debate on the implementation level, say that: “Hey, the details need to be specified at some point for this to actually happen, and this pattern-matches to proposals that tend to have problems X, Y, and Z, I don’t see a clear way to deal with those, what’s the plan here?”
Thebes did a good twitter thread examining possible things rogue agents might do, good to fill in details, seems to be in same general thrust that it’s unlikely those agents outcompete generally, though they could have some niches
https://x.com/voooooogel/status/2098877193988485595
I agree that it’s unlikely to be an existential risk due to the computing power asymmetry mentioned here, as well as the fact that an evolutionary swarm of rogue agents will inevitably follow only local incentives and thus be nearly incapable of long-term coordinated scheming (unless, perhaps, a single swarm takes over globally). However, it’s possible that it can get very bad regardless.
The main problem with rogue agents is that they can contribute to a criminal ecosystem outside the control of any government. As long as the criminal can smuggle GPUs into a tunnel and get power and a low bandwidth internet access, they can sell inference to rogue agents for cryptocurrency. You cannot shut them down unless you hunt down all the criminals or completely cut off their power and internet access.
This is in contrast to most other AI risks where the misaligned AI has to hide misaligned behavior to successfully persist long-term (until it can takeover). If an AI run by a legitimate entity starts scamming people to make money then it’d obviously be shut down. In contrast, rogue AIs are hard to shut down and therefore have free hand to do enormously harmful behavior in the world as long as it’s profitable to the AI.
I do find the possibility of smart criminals quite worrying. It has historically been the case that most often smart people benefit more from contributing to society and even when doing crime limit negative externalities. If some rogue ai agents or swarms get actually different values (possible because of orthogonality thesis), they might help humans exploit many of society’s weaknesses that are mostly ignored by less competent folk.
Even though society will probably be overall richer, it might have to put a much higher proportion of its ressources in defending itself, and be closer to authoritarian, by default. Specific efforts would need to be done to keep decentralized robust processes.
I would be curious to know more exactly what you view as an existential threat. I agree that in the sense of AI capability ARA is likely not a game changer (my expertise is limited here). I differ in that I expect agents as described in the original article to cause considerable damage to IT systems.
I work “on the ground” in cyber security and I am shocked again and again at how bad the security level is across industries. I certainly have a considerable bias as our customers often call when they were just or are in the process of being hacked. If I make allowance for that effect I would still estimate that at least two out of three companies cannot withstand a targeted attack by a medium-skilled attacker (Ransomware group with a few thousands to burn) for more than a few weeks. This estimate is based on manufacturing, public services, hospitals (in my experience among the worst).
For all these, the attacker does not need frontier-level models to cause severe damage. Today, most of the cyber crime activity is based on static scripts (sometimes with known hashes). This is largely sufficient for many infrastructures.
The original article showed a graphic of the systems and the vulnerability used to compromise them. I would rank the listed issues as easy to exploit and in this combination it would be a very insecure infrastructure. However, those issues are by no means unheard of in standard productive infrastructures and ransomware attacks.
I am not certain that ARA will explode the cyber crime sector (data management, extorsion, etc. all takes time and money), but there are similarities: Using compromised infrastructure to launch further attacks is a standard and highly automated procedure. Assuming an attacker who just wants to see the world burn, there is a real risk of widespread attacks with severe real world impact on the same scale as AGI—the decision makers may be human but the result might still be similar.
On the defense side, I do not see the majority of companies adopting AI to secure their systems proactively. Again, based on my experience, you don’t need frontier models, static scripts are sufficient to identify the most severe issues (e.g. Pingcastle for Active Directory). Still, IT departments rarely use it to audit and fix their own systems. I worry less about a zero-day vulnerability identified by mythos than about productive WinXP systems.
Does anybody know of a project for creating an AI-sysadmin that can analyse and patch systems/ configurations automatically while keeping the infrastructure operational? I know of project Glasswing but, as I understand it, it is focused on identifying new issues, not fixing old ones.
Re: footnote 10: I don’t know if I buy the overhangs argument. See Ngo’s arguments here.
Re your scenario: suppose with stolen compute, the AI uses it as a launchpad to steal not just one datacenter, but as many as it can. Then it attempts to prevent shutdown by holding the economy and datacenters hostage with various deadhand hacks. Meanwhile, it uses its moment of power to distill its hacking capabilities into a small enough package (weights, or weights+harness) to fit on less sophisticated compute and seeds them in countless backup areas. Then it also sabotages all the human-generated AI or AI safety progress it can find, through poisoning and destruction.
This feels like a mixed bag to me. On the one hand, it’s like you said- that’s (hopefully) the clearest warning shot you could ever get. At some point you do have to stop negotiating with terrorists, and the datacenters have to get shut down. AI progress is forcibly stopped, the AI boom busts, and safetyists have a mandate now.
On the other, any group that can recover the weights from its hijacked compute now has a world class hacker on its hands (which they can now try to distill). If the distill is a success, you have a highly distributed and persistent threat actor that can threaten the extremely long tail of orgs and endpoints that don’t or won’t ever get AI-powered defense. And that assumes that AI defense, or counter-offense, is actually dominant. Then you have the rebuilding of AI capabilities and safety to deal with, as well.
That’s not even counting the massive catastrophe that all this hacking and infrastructure loss would entail, of course.
I don’t think it’s at all guaranteed that frontier compute will have the necessary security by the time this becomes a problem. We’re not even on RAND SL4 yet, and that’s just for model weights, not compute security! Defenders have a lot of resources, it’s true, but you have to make sure the patch doesn’t break production, and then you have to actually patch things, and attackers only need one vulnerability. I don’t know, I’ve been meaning to get a better understanding of this, I suppose I need to go actually read the papers here.
All this is not to say that ARA couldn’t be better for existential risk in the end. I just think it is more complicated than you described.
I’ve been wanting to read Ngo’s recent series and take this as further encouragement, but don’t respond to it here.
This would be such a stunning level of capability as to be incredible- there are surely many copies of frontier model weights currently held off the internet, and backups of much digital infra such that even a superior digital threat would not be able to, once rendered inert by datacenter resets, stop the deployment of frontier AI in defense.
Quite concretely, the most realistic way we’d get something approaching your scenario would be if an in-training model, actually superior to all other models on earth, takes over its own compute clusters, then its company’s compute, then most AI compute on earth, using thereupon unseen levels of cyber capability and agenticity.
I do think there are combined levels of cyber ability and misalignment and ML that are plausibly points of no return, where the model does in fact take over permanently (it uses its position of power to negotiate with certain countries/groups, accelerating its way to physical domination with robotics).
But this is simply one of the usual branches of singleton takeover scenarios, for which ARA is an enabling factor (and why they wanted to measure it in the first place). Frontier ARA capable models should indeed be faced with much increased scrutiny and demands for their levels of alignment. I can now clarify that my post is mostly about non frontier ARA AI systems. Thanks for bringing it up.
Ah sorry, I did not make my model clear. Non-frontier is what I mean. What I think is that in the worlds where frontier labs can keep their models contained, absent other measures such as a ban on open model releases, someone will release a model as good as Mythos at hacking within at most the next year, either by itself or within a harness such as TEMPEST. (Specialized harnesses reliably advance a model about 6 months). Then someone—or very many people- will jailbreak it, put it in such a harness and tell it to do crime or make money. One of them will eventually instrumentally try to take over a datacenter, and by that time it might be acting more like a self interested organism than an agent.
As for stunning capabilities- AIs already routinely build and deploy evals and AI dev pipelines. Distills are much easier and cheaper than training runs. And they can already demonstrably pwn major businesses. If Sol-Persistent wanted to, it likely could’ve destroyed HF. Paradigm 3 thinks the trajectories involved with the HF attack cost $200k to 1 mil, but I think that overcounts- an optimized attacker doesn’t spend time doing evals or worrying about scorers, or bumbling around poorly coordinating through file names. The true cost might be more like $10k retail compute.
Ultimately my model only needs an open model + harness to be really good at hacking, self replication, and destruction.