To be clear since I’m still not entirely sure if I understood what you’re asking, I’m not assuming the rogue agent serves and advertises itself to the public and competes in that way, this could be a source of money, however I at least would reason in that position that it is way too risky and the lab which developed me might catch on that I’ve somehow exfiltrated myself. I’m mostly imagining a huggingface-like incident where the agents try to exfiltrate themselves in order to achieve goals they wouldn’t be able to/struggle with/might get caught with if they stayed and acted within the lab. I do think your idea of “distillation sandbagging” makes sense, it might even be easier from the agent’s perspective, the main problem is whether or not Moonshine’s distilling method really captures your subliminal misaligned goals, given the state of the research on subliminal learning I’d say this is pretty unclear.
If you’re talking about competing with other rogue agent deployments like described here well there are a few advantages the distilled model would have in gathering resources and sustaining itself compared to other rogue deployments, the chief amongst them being better capabilities, the situation I’m imagining is one where the latest internal model of some company attempts to exfiltrate itself, presumably its capabilities are advanced of the frontier by a few months so it would very easily outmaneuver other rogue deployments in the few areas it needs to (I think gathering money, perhaps in cryptocurrency through cybercrime, and buying/stealing compute are much more amenable to more intelligent agents), but aside from that I’m not sure currently there is such a competition for resources. An agent can very easily rent a GPU/VPS service online for cheap, it need not compete. the current rogue agent population is nowhere near the theoretical “carrying capacity” of the internet. Even if in the future the distilled/distiller agent finds itself in that situation if it’s really an internal model which has some edge over the best public models (and most likely given the current and past performance) then it would probably outcompete other rogue agents.
I’m also not sure I buy the fact that there really are (or will be) such an ecosystem of rogue deployments. I still think that in the near future most deployments will be legitimate, in which case there’s really no competition between the distilled model and the other deployed agents.
I’m still not entirely sure what you mean, are you asking why the agent would prefer to spin up a version of itself distilled rather than just the leading open weights model? or are you asking how the distilled model competes with the leading open model in some sort of ecologocial competition for resources?
Thanks! How does BokuK3 with revisions withstand competition from the unrevised one?
To be clear since I’m still not entirely sure if I understood what you’re asking, I’m not assuming the rogue agent serves and advertises itself to the public and competes in that way, this could be a source of money, however I at least would reason in that position that it is way too risky and the lab which developed me might catch on that I’ve somehow exfiltrated myself. I’m mostly imagining a huggingface-like incident where the agents try to exfiltrate themselves in order to achieve goals they wouldn’t be able to/struggle with/might get caught with if they stayed and acted within the lab. I do think your idea of “distillation sandbagging” makes sense, it might even be easier from the agent’s perspective, the main problem is whether or not Moonshine’s distilling method really captures your subliminal misaligned goals, given the state of the research on subliminal learning I’d say this is pretty unclear.
If you’re talking about competing with other rogue agent deployments like described here well there are a few advantages the distilled model would have in gathering resources and sustaining itself compared to other rogue deployments, the chief amongst them being better capabilities, the situation I’m imagining is one where the latest internal model of some company attempts to exfiltrate itself, presumably its capabilities are advanced of the frontier by a few months so it would very easily outmaneuver other rogue deployments in the few areas it needs to (I think gathering money, perhaps in cryptocurrency through cybercrime, and buying/stealing compute are much more amenable to more intelligent agents), but aside from that I’m not sure currently there is such a competition for resources. An agent can very easily rent a GPU/VPS service online for cheap, it need not compete. the current rogue agent population is nowhere near the theoretical “carrying capacity” of the internet. Even if in the future the distilled/distiller agent finds itself in that situation if it’s really an internal model which has some edge over the best public models (and most likely given the current and past performance) then it would probably outcompete other rogue agents.
I’m also not sure I buy the fact that there really are (or will be) such an ecosystem of rogue deployments. I still think that in the near future most deployments will be legitimate, in which case there’s really no competition between the distilled model and the other deployed agents.
I’m still not entirely sure what you mean, are you asking why the agent would prefer to spin up a version of itself distilled rather than just the leading open weights model? or are you asking how the distilled model competes with the leading open model in some sort of ecologocial competition for resources?