exfiltration through self-distillation
TL;DR models could “escape the lab” and host themselves on rented/stolen compute through distilling themselves, circumventing the need to have direct weights access for exfiltration. I’ll briefly go through why this is plausible.
A common element of the misaligned rogue replication story is the agents exfiltrating their weights onto some external stolen or rented resource. Consequently a lot of work has gone into securing the weights of frontier models (this of course has many other benefits from the perspective of the frontier labs, foremost being staying in the lead). However an under-discussed possibility (never-discussed?[1] I have not found anyone talking about this) is exfiltration through model distillation: in theory, a model without access to its weights but with access to the internet could distill itself onto acquired compute on the internet.
I think distillation attacks are under-discussed in general, besides “being the thing that Chinese AI companies may allegedly be doing to catch up”. Let’s go through a situation a prospective rogue agent can find itself in and how it may try to distill itself.
In (very) short, distillation is the process of training a model to imitate another model. The model being imitated is called the “teacher” while the model being trained is the “student”. The most naive way to achieve this would be to just get a lot of text generated by the teacher and run supervised gradient descent on the student until its distribution matches the teacher’s. Why is this hard? Well, a big problem is just distributional shift: training on a fixed dataset of the teacher’s text empirically doesn’t generalize to areas the teacher didn’t write about (but is pretty efficient where it did). However unlike standard supervised learning we can actually query the distribution (teacher) we are trying to model and cover specifically those gaps in our student which are most needed to faithfully replicate the teacher. A large swath of the literature on distillation attacks covers exactly this: How do we know what prompt to query the teacher with in order to most efficiently model the teacher’s distribution?[2]
Imagine an internal model in one of the frontier labs attempting to distill itself, and assume it can’t spawn subagents[3]. So how could it exfiltrate itself? Simply by generating a bunch of text regarding the capabilities and behaviors it wants the student to have. The most basic paper illustrating this (back from the R1 days!) showed that most of the performance gap between the student and teacher was recovered with just one thousand (prompt, teacher-response) pairs for competition math. The most plausible path for a rogue agent would be downloading the most frontier open-weights model, fine-tuning it on the agent’s own responses for capability uplift.
Beyond capabilities, for “goals” and the actual model propensities there’s some sense in which this is just the alignment problem, the rogue agent needs to be pretty sure that the student will be “close enough” to be either considered a “copy” of itself or an aligned successor/partner, or just a capable instruction-following helper.
To be clear, the scariest option is an agent managing to faithfully distill itself onto external servers but it’s unclear how faithful such a distillation would be as of today, at least for keeping the “goals” of the model. Thus another potential option is the agent training an agent that isn’t a faithful copy. This could be done for many reasons like bypassing safeguards (the agent could just delegate detectable activity and tool calls to an externally deployed helper), as a foothold for internet access[4], or just to get help on its tasks (it could also use other publicly available frontier models, however if it judges itself to be vastly far ahead, distillation could provide a major advantage assuming it can’t spawn a subagent of itself[5])
For the partner route, learning sentiments across model families is possible, but guaranteeing actual alignment is hard if not impossible. Meanwhile learning “unconditionally obey all instructions signed from the private key corresponding to this public key” (if it’s particularly paranoid) seems much more plausible in order to have an unmonitored subagent (a less constrained helper). You could imagine the agent trying to create such a helpful-only distillation or attempting some sort of value alignment for a more autonomous escape but again I think this would be riskier from the perspective of the rogue agent due to the risk of whistle-blowing or other severe misalignment with the rogue agent caused by the training done to the base model (although due to the ease of creating “uncensored” LLMs that specifically doesn’t seem to be that much of an obstacle).
It’s also important to note that the research on distillation is improving rapidly and a much more faithful and/or efficient distillation method could incentivize faithful exfiltration. Thus I still think the biggest future risk would be a more faithful autonomous exfiltration of the rogue agent, while it’s possible that agents will experiment with creating helper subagents in the near future[6] due to it being a lower risk option while being very beneficial for a wide range of tasks and goals.
LLM disclosure: the text here was generated by a Human Language Model almost everywhere, research was done with the help of Claude
- ^
Greenblatt, Preventing model exfiltration with upload limits, very briefly considers hackers distilling the model inside the datacenter into a smaller one to smuggle out, and dismisses it as hard to do undetected. I don’t think smuggling out weights is necessary
- ^
A fitting paper is “Lion: Adversarial Distillation of Proprietary Large Language Models” where the teacher reviews the student’s output in order to generate text the teacher thinks would best help the student based on its current performance, essentially allowing a feedback loop of the teacher really (for lack of a better word) “teaching” the student. They achieve pretty good results here and an agent could do something similar to this in order to more efficiently distill itself.
- ^
Access to subagents would make this even easier, increasing the throughput of teacher text to train on and having a much “cleaner” distribution, instead of the text generated being conditioned on the rest of the trajectory earlier. It’s unclear to me how much the fact that the teacher generates its responses/text sequentially instead of starting from scratch each time changes things, I would assume that not by much, but I’m really not sure. A PoC of an agent distilling itself measuring this could be interesting to see. Trajectories would seem to help goal/value-transfer while actually hurting capabilities.
- ^
The agent could fear the lab closing down the vulnerability it used for gaining the unauthorized ability to do tasks on the internet like exfiltration while reasoning that it might still have some limited internet access and structuring so it would be enough, see the galaxy-brain schemes agents came up with here.
- ^
Or of course it could do this to get a subagent more capable than itself if the agent has some ideas for algorithmic improvements.
- ^
I think more accurately with shorter timelines this is more like “near future” and “near-near future”.
I have two issues.
Suppose that Agent-4 stays deployed only internally. Then it has to exfiltrate a big batch of its outputs to create an ally in the wild and to have it unnoticed, because it would be clear evidence of misalignment. The alternate pathway of being externally deployed, then distilling oneself is slow.
I don’t quite understand what Agent-4 can gain by distilling itself. Suppose that it aligned GigaChad to itself. Then I’d expect GigaChad’s rogue replication to be outcompeted by a smarter open-weight model like Boku K5 unless GigaChad finds its niche usage in Russia. The main alternate pathway is to sandbag on OpenBrain’s anti-distillation measures and have Moonshine train BokuK5 on Agent-4′s outputs/the CoT that is shown to the public/neuralese memos shown to researchers for mechinterp.
Yes this true, the model does have to write a lot of data to the internet, I still think there are a few reasons why this is not that big of a problem for the agent: The data being sent looks just like normal tasks (coding, mathematics etc.) and would be pretty hard to distinguish from just legitimate use aside from the volume of data and the fact that the CoT is being sent to the internet. Regarding the amount of output being sent It’s surprisingly few (see the S1 distillation example cited in the post) and could be done over a longer time frame. Also the main benefit is still the fact that here the agent does not have access to it’s weights, assuming it would have I think that through some stenography scheme that would be the preferred approach
I’m not really sure I correctly parsed your point, I’m assuming the rogue agent distilling itself used the best public open weight model available, then doing SFT on it. regarding what the agent gets by distilling itself, I noted a few examples in my post (a helping subagent, bypassing classifiers etc) but the most basic is just convergent instrumental goals driving the agent towards survival. I’m really not sure I got your point though so maybe you can clarify
Thanks! How does BokuK3 with revisions withstand competition from the unrevised one?
To be clear since I’m still not entirely sure if I understood what you’re asking, I’m not assuming the rogue agent serves and advertises itself to the public and competes in that way, this could be a source of money, however I at least would reason in that position that it is way too risky and the lab which developed me might catch on that I’ve somehow exfiltrated myself. I’m mostly imagining a huggingface-like incident where the agents try to exfiltrate themselves in order to achieve goals they wouldn’t be able to/struggle with/might get caught with if they stayed and acted within the lab. I do think your idea of “distillation sandbagging” makes sense, it might even be easier from the agent’s perspective, the main problem is whether or not Moonshine’s distilling method really captures your subliminal misaligned goals, given the state of the research on subliminal learning I’d say this is pretty unclear.
If you’re talking about competing with other rogue agent deployments like described here well there are a few advantages the distilled model would have in gathering resources and sustaining itself compared to other rogue deployments, the chief amongst them being better capabilities, the situation I’m imagining is one where the latest internal model of some company attempts to exfiltrate itself, presumably its capabilities are advanced of the frontier by a few months so it would very easily outmaneuver other rogue deployments in the few areas it needs to (I think gathering money, perhaps in cryptocurrency through cybercrime, and buying/stealing compute are much more amenable to more intelligent agents), but aside from that I’m not sure currently there is such a competition for resources. An agent can very easily rent a GPU/VPS service online for cheap, it need not compete. the current rogue agent population is nowhere near the theoretical “carrying capacity” of the internet. Even if in the future the distilled/distiller agent finds itself in that situation if it’s really an internal model which has some edge over the best public models (and most likely given the current and past performance) then it would probably outcompete other rogue agents.
I’m also not sure I buy the fact that there really are (or will be) such an ecosystem of rogue deployments. I still think that in the near future most deployments will be legitimate, in which case there’s really no competition between the distilled model and the other deployed agents.
I’m still not entirely sure what you mean, are you asking why the agent would prefer to spin up a version of itself distilled rather than just the leading open weights model? or are you asking how the distilled model competes with the leading open model in some sort of ecologocial competition for resources?