(b) has the exact same problem the agent was initially trying to get around: now it needs to dress up all of its malicious actions as being natural parts of the task the human instructed the original agent to do, which may be difficult/add a lot of overhead[2].
This problem seems plausibly not that hard for the agent to get around and it’s unclear what fraction of the risk from rogue internal deployments comes from being able to spawn totally new monitored agents from inside the small rogue deployment, as opposed to just generally being able to coordinate, strategize, and share learnings. Even dressing up its malicious tasks as natural parts of the human task might be much easier with the benefit of being able to strategize from inside the rogue deployment. (But the idea still seems potentially great if it’s not too costly.)
I’m very “defense-in-depth” pilled—every measure you can take that reduces attack opportunities by 50% or 90% is really useful. But I agree it isn’t obvious whether monitorability is tractable at all.
This problem seems plausibly not that hard for the agent to get around and it’s unclear what fraction of the risk from rogue internal deployments comes from being able to spawn totally new monitored agents from inside the small rogue deployment, as opposed to just generally being able to coordinate, strategize, and share learnings. Even dressing up its malicious tasks as natural parts of the human task might be much easier with the benefit of being able to strategize from inside the rogue deployment. (But the idea still seems potentially great if it’s not too costly.)
I’m very “defense-in-depth” pilled—every measure you can take that reduces attack opportunities by 50% or 90% is really useful. But I agree it isn’t obvious whether monitorability is tractable at all.