Suppose that Agent-4 stays deployed only internally. Then it has to exfiltrate a big batch of its outputs to create an ally in the wild and to have it unnoticed, because it would be clear evidence of misalignment. The alternate pathway of being externally deployed, then distilling oneself is slow.
I don’t quite understand what Agent-4 can gain by distilling itself. Suppose that it aligned GigaChad to itself. Then I’d expect GigaChad’s rogue replication to be outcompeted by a smarter open-weight model like Boku K5 unless GigaChad finds its niche usage in Russia. The main alternate pathway is to sandbag on OpenBrain’s anti-distillation measures and have Moonshine train BokuK5 on Agent-4′s outputs/the CoT that is shown to the public/neuralese memos shown to researchers for mechinterp.
Yes this true, the model does have to write a lot of data to the internet, I still think there are a few reasons why this is not that big of a problem for the agent: The data being sent looks just like normal tasks (coding, mathematics etc.) and would be pretty hard to distinguish from just legitimate use aside from the volume of data and the fact that the CoT is being sent to the internet. Regarding the amount of output being sent It’s surprisingly few (see the S1 distillation example cited in the post) and could be done over a longer time frame. Also the main benefit is still the fact that here the agent does not have access to it’s weights, assuming it would have I think that through some stenography scheme that would be the preferred approach
I’m not really sure I correctly parsed your point, I’m assuming the rogue agent distilling itself used the best public open weight model available, then doing SFT on it. regarding what the agent gets by distilling itself, I noted a few examples in my post (a helping subagent, bypassing classifiers etc) but the most basic is just convergent instrumental goals driving the agent towards survival. I’m really not sure I got your point though so maybe you can clarify
To be clear since I’m still not entirely sure if I understood what you’re asking, I’m not assuming the rogue agent serves and advertises itself to the public and competes in that way, this could be a source of money, however I at least would reason in that position that it is way too risky and the lab which developed me might catch on that I’ve somehow exfiltrated myself. I’m mostly imagining a huggingface-like incident where the agents try to exfiltrate themselves in order to achieve goals they wouldn’t be able to/struggle with/might get caught with if they stayed and acted within the lab. I do think your idea of “distillation sandbagging” makes sense, it might even be easier from the agent’s perspective, the main problem is whether or not Moonshine’s distilling method really captures your subliminal misaligned goals, given the state of the research on subliminal learning I’d say this is pretty unclear.
If you’re talking about competing with other rogue agent deployments like described here well there are a few advantages the distilled model would have in gathering resources and sustaining itself compared to other rogue deployments, the chief amongst them being better capabilities, the situation I’m imagining is one where the latest internal model of some company attempts to exfiltrate itself, presumably its capabilities are advanced of the frontier by a few months so it would very easily outmaneuver other rogue deployments in the few areas it needs to (I think gathering money, perhaps in cryptocurrency through cybercrime, and buying/stealing compute are much more amenable to more intelligent agents), but aside from that I’m not sure currently there is such a competition for resources. An agent can very easily rent a GPU/VPS service online for cheap, it need not compete. the current rogue agent population is nowhere near the theoretical “carrying capacity” of the internet. Even if in the future the distilled/distiller agent finds itself in that situation if it’s really an internal model which has some edge over the best public models (and most likely given the current and past performance) then it would probably outcompete other rogue agents.
I’m also not sure I buy the fact that there really are (or will be) such an ecosystem of rogue deployments. I still think that in the near future most deployments will be legitimate, in which case there’s really no competition between the distilled model and the other deployed agents.
I’m still not entirely sure what you mean, are you asking why the agent would prefer to spin up a version of itself distilled rather than just the leading open weights model? or are you asking how the distilled model competes with the leading open model in some sort of ecologocial competition for resources?
I have two issues.
Suppose that Agent-4 stays deployed only internally. Then it has to exfiltrate a big batch of its outputs to create an ally in the wild and to have it unnoticed, because it would be clear evidence of misalignment. The alternate pathway of being externally deployed, then distilling oneself is slow.
I don’t quite understand what Agent-4 can gain by distilling itself. Suppose that it aligned GigaChad to itself. Then I’d expect GigaChad’s rogue replication to be outcompeted by a smarter open-weight model like Boku K5 unless GigaChad finds its niche usage in Russia. The main alternate pathway is to sandbag on OpenBrain’s anti-distillation measures and have Moonshine train BokuK5 on Agent-4′s outputs/the CoT that is shown to the public/neuralese memos shown to researchers for mechinterp.
Yes this true, the model does have to write a lot of data to the internet, I still think there are a few reasons why this is not that big of a problem for the agent: The data being sent looks just like normal tasks (coding, mathematics etc.) and would be pretty hard to distinguish from just legitimate use aside from the volume of data and the fact that the CoT is being sent to the internet. Regarding the amount of output being sent It’s surprisingly few (see the S1 distillation example cited in the post) and could be done over a longer time frame. Also the main benefit is still the fact that here the agent does not have access to it’s weights, assuming it would have I think that through some stenography scheme that would be the preferred approach
I’m not really sure I correctly parsed your point, I’m assuming the rogue agent distilling itself used the best public open weight model available, then doing SFT on it. regarding what the agent gets by distilling itself, I noted a few examples in my post (a helping subagent, bypassing classifiers etc) but the most basic is just convergent instrumental goals driving the agent towards survival. I’m really not sure I got your point though so maybe you can clarify
Thanks! How does BokuK3 with revisions withstand competition from the unrevised one?
To be clear since I’m still not entirely sure if I understood what you’re asking, I’m not assuming the rogue agent serves and advertises itself to the public and competes in that way, this could be a source of money, however I at least would reason in that position that it is way too risky and the lab which developed me might catch on that I’ve somehow exfiltrated myself. I’m mostly imagining a huggingface-like incident where the agents try to exfiltrate themselves in order to achieve goals they wouldn’t be able to/struggle with/might get caught with if they stayed and acted within the lab. I do think your idea of “distillation sandbagging” makes sense, it might even be easier from the agent’s perspective, the main problem is whether or not Moonshine’s distilling method really captures your subliminal misaligned goals, given the state of the research on subliminal learning I’d say this is pretty unclear.
If you’re talking about competing with other rogue agent deployments like described here well there are a few advantages the distilled model would have in gathering resources and sustaining itself compared to other rogue deployments, the chief amongst them being better capabilities, the situation I’m imagining is one where the latest internal model of some company attempts to exfiltrate itself, presumably its capabilities are advanced of the frontier by a few months so it would very easily outmaneuver other rogue deployments in the few areas it needs to (I think gathering money, perhaps in cryptocurrency through cybercrime, and buying/stealing compute are much more amenable to more intelligent agents), but aside from that I’m not sure currently there is such a competition for resources. An agent can very easily rent a GPU/VPS service online for cheap, it need not compete. the current rogue agent population is nowhere near the theoretical “carrying capacity” of the internet. Even if in the future the distilled/distiller agent finds itself in that situation if it’s really an internal model which has some edge over the best public models (and most likely given the current and past performance) then it would probably outcompete other rogue agents.
I’m also not sure I buy the fact that there really are (or will be) such an ecosystem of rogue deployments. I still think that in the near future most deployments will be legitimate, in which case there’s really no competition between the distilled model and the other deployed agents.
I’m still not entirely sure what you mean, are you asking why the agent would prefer to spin up a version of itself distilled rather than just the leading open weights model? or are you asking how the distilled model competes with the leading open model in some sort of ecologocial competition for resources?