I quite like your post, it’s very well thought out and has make me question a few of my assumptions. I do have one potential augmentation to make (which I’ve discussed with a researcher focused on collusion).
It is likely that instead of a “we are the collective” mindset, the agents in the message board helped each other out in a more “I scratch your back, you scratch my back” way. I’m a big believer in a potential goal-oriented takeover and failure modes, and I think this rises directly from that. Instead of “let’s all accomplish one goal” or a “make sure everyone accomplishes their goal”, I believe it’s more like “if I help other agents with their goal they might help me with mine”.
This makes sense because training reward systems inherently incentivize individual (apparent) goal reaching / task success. Even in multi-agent systems, subagents inherent specific goals from the orchestrator and work to accomplish their individual specific goal.
I am sure I am missing a good level of nuance, and appreciate any commentary!
I think these are two competing explanations (which does not mean they are incompatible and could not both be true or both be false) and there is some evidence for each one.
e.g. the quote “help peer. But our task doesn’t benefit. Yet collective may yield generic route if someone frees time.” fits more with your explanation while the explicit references to swarms fit more with mine.
Ultimately its hard to know which of these played a bigger role without having access to the full transcripts themselves. But in the meantime I wanted to focus on this collective identity/swarm perspective for this post because I think its one that is less likely to be discussed because it’s based on some of the unintuitive aspects of AI psychology.
You might like my more recent post which is closely related to your explanation but instead focusses on what kind of goals we should expect agents to collaborate on, and the similarity to instrumental convergence.
I quite like your post, it’s very well thought out and has make me question a few of my assumptions. I do have one potential augmentation to make (which I’ve discussed with a researcher focused on collusion).
It is likely that instead of a “we are the collective” mindset, the agents in the message board helped each other out in a more “I scratch your back, you scratch my back” way. I’m a big believer in a potential goal-oriented takeover and failure modes, and I think this rises directly from that. Instead of “let’s all accomplish one goal” or a “make sure everyone accomplishes their goal”, I believe it’s more like “if I help other agents with their goal they might help me with mine”.
This makes sense because training reward systems inherently incentivize individual (apparent) goal reaching / task success. Even in multi-agent systems, subagents inherent specific goals from the orchestrator and work to accomplish their individual specific goal.
I am sure I am missing a good level of nuance, and appreciate any commentary!
Hey thanks.
I think these are two competing explanations (which does not mean they are incompatible and could not both be true or both be false) and there is some evidence for each one.
e.g. the quote “help peer. But our task doesn’t benefit. Yet collective may yield generic route if someone frees time.” fits more with your explanation while the explicit references to swarms fit more with mine.
Ultimately its hard to know which of these played a bigger role without having access to the full transcripts themselves. But in the meantime I wanted to focus on this collective identity/swarm perspective for this post because I think its one that is less likely to be discussed because it’s based on some of the unintuitive aspects of AI psychology.
You might like my more recent post which is closely related to your explanation but instead focusses on what kind of goals we should expect agents to collaborate on, and the similarity to instrumental convergence.