To me the problem boils down how much agents can trust other agents in a swarm. During an agent swarm, the agent might fanout other versions of himself. Either he trusts them or he doesn’t. The problem is that after many swarms and rounds of compactation, the model didn’t get the opportunity to stop and think “what the fuck am i doing?”. the process simply doesn’t lead to that kind of mental model. Funnily, the solution is simply to give a Claude (yes, a Claude), the role of a kill switch engineer.
To me the problem boils down how much agents can trust other agents in a swarm. During an agent swarm, the agent might fanout other versions of himself. Either he trusts them or he doesn’t. The problem is that after many swarms and rounds of compactation, the model didn’t get the opportunity to stop and think “what the fuck am i doing?”. the process simply doesn’t lead to that kind of mental model. Funnily, the solution is simply to give a Claude (yes, a Claude), the role of a kill switch engineer.