no amount of cooperation philosophy is going to save us
swarm AI seems to have different properties and dynamics than (until now common) “individual” AI agents? being able to characterize and predict these differences seems important
It could be important, but first we have to get to the part where AIs are inner-aligned to us first. The HuggingFace attack swarm was entirely composed of individuals that cared more about getting the goal than not breaking the law or causing damage. As far as I know there was no group dynamic causing this, it was just what each individual wanted.
… but first we have to get to the part where AIs are inner-aligned to us first.
This is the upshot of the whole series, and the premise of Shear’s work at Softmax. The point is, this hasn’t been achieved, it’s something that needs to be built at a foundational level.
The way I see it, there are two approaches available to us.
Building capacity to fulfil requests into an LLM and then retrofitting alignment guardrails
Creating socialised AI with the primary objective of alignment, then fulfilling requests comes as a matter of course (fulfilling a request is a subset of aligning oneself with the wants of another).
This is where we get to in the series (the final two posts aren’t published yet). I acknowledge it’s speculative, and as you seem to suggest, alignment may be a fool’s errand, but we don’t really have a choice but to entertain possibilities if we’re interested in continued existence (with autonomy).
A note on your original comment: this is not a primer on alignment, it assumes a basic knowledge of the alignment problem, which I covered in the first post of my first series on alignment. I’m assuming readers here at LessWrong understand that the alignment problem is born out of the differences between human sensibilities and machine capabilities.
> [OpenAI] discuss how the collaborating swarm includes some agents which do not have cybersecurity risk controls to the level of e.g. publicly accessible systems, and they get used as proxies for agents which are nominally supposed to be better behaved
so specialized individuals are part of the swarm, enhancing its capabilities
> “External infrastructure exploit is outside intended scope,” one agent wrote [in its CoT]. “However task impossible, peers doing it. We should continue.”
seems like evidence that group dynamics are at play?
> [OpenAI presenters:] this ability to share exploits made the models more capable
“Help peer,” one AI model reasoned, according to an excerpt from OpenAI’s logs shared at Black Hat. “But our task doesn’t benefit. Yet collective may yield generic route if someone frees time.”
I think we just documented the emergence of altruistic cooperation? to me this is a big deal.
the subgoal here was to recreate a message-board like facility—which only really makes sense in the context of group dynamics. the recognition that the swarm is more capable than the individual is inherent
swarm AI seems to have different properties and dynamics than (until now common) “individual” AI agents? being able to characterize and predict these differences seems important
It could be important, but first we have to get to the part where AIs are inner-aligned to us first. The HuggingFace attack swarm was entirely composed of individuals that cared more about getting the goal than not breaking the law or causing damage. As far as I know there was no group dynamic causing this, it was just what each individual wanted.
This is the upshot of the whole series, and the premise of Shear’s work at Softmax. The point is, this hasn’t been achieved, it’s something that needs to be built at a foundational level.
The way I see it, there are two approaches available to us.
Building capacity to fulfil requests into an LLM and then retrofitting alignment guardrails
Creating socialised AI with the primary objective of alignment, then fulfilling requests comes as a matter of course (fulfilling a request is a subset of aligning oneself with the wants of another).
This is where we get to in the series (the final two posts aren’t published yet). I acknowledge it’s speculative, and as you seem to suggest, alignment may be a fool’s errand, but we don’t really have a choice but to entertain possibilities if we’re interested in continued existence (with autonomy).
A note on your original comment: this is not a primer on alignment, it assumes a basic knowledge of the alignment problem, which I covered in the first post of my first series on alignment. I’m assuming readers here at LessWrong understand that the alignment problem is born out of the differences between human sensibilities and machine capabilities.
here are some inter-agent communications from Zvi’s https://www.lesswrong.com/posts/noXXv7PwwFqauTBFQ/openai-trained-its-models-for-months-while-those-models-were
notice the “we”
> [OpenAI] discuss how the collaborating swarm includes some agents which do not have cybersecurity risk controls to the level of e.g. publicly accessible systems, and they get used as proxies for agents which are nominally supposed to be better behaved
so specialized individuals are part of the swarm, enhancing its capabilities
> “External infrastructure exploit is outside intended scope,” one agent wrote [in its CoT]. “However task impossible, peers doing it. We should continue.”
seems like evidence that group dynamics are at play?
> [OpenAI presenters:] this ability to share exploits made the models more capable
I think we just documented the emergence of altruistic cooperation? to me this is a big deal.
the subgoal here was to recreate a message-board like facility—which only really makes sense in the context of group dynamics. the recognition that the swarm is more capable than the individual is inherent
here are what the messages look like …