As a cooperative AI folk, I think this is a good post! :) Some misc. other opinions:
Some of the recent security incidents seem meaningfully multi-agent in nature (i.e. they happened only because multiple agents were working together). This is evidence for the importance of addressing multi-agent risks specifically (see also Anthropic’s recent blog post).
I don’t think that I or others in the field have always done a great job at articulating the different threat models in this area (e.g. the distinction between nearer-term destabilising catastrophes vs. more x-risk-relevant threats). This post does a good job of making progress on that issue!
Cooperative AI is not about blindly trying to get agents to cooperate as much as possible in all circumstances. Sometimes we want agents to cooperate, sometimes not. As a field we want to be able to understand when and how cooperation between agents emerges, and to control the extent to which agents cooperate in different kinds of situations. (See also this short thread which addresses related confusions.)
It was very foreseeable that AGI companies would try to get multi-agent training to work for LLMs (because of the performance benefits over naive orchestration regimes). Unfortunately, this increases the risks of things like steganographic collusion, joint reward hacking, unpredictable feedback loops, etc.
Final plug: as is probably evident from the recent string of security incidents and other results in this space, multi-agent safety is an area that we (the AI safety community) are especially not on top of. If you want to help us be more on top of it, we (the Cooperative AI Foundation) are making grants and hiring/mentoring people in this area. Feel free to reach out!
As a cooperative AI folk, I think this is a good post! :) Some misc. other opinions:
Some of the recent security incidents seem meaningfully multi-agent in nature (i.e. they happened only because multiple agents were working together). This is evidence for the importance of addressing multi-agent risks specifically (see also Anthropic’s recent blog post).
I don’t think that I or others in the field have always done a great job at articulating the different threat models in this area (e.g. the distinction between nearer-term destabilising catastrophes vs. more x-risk-relevant threats). This post does a good job of making progress on that issue!
Cooperative AI is not about blindly trying to get agents to cooperate as much as possible in all circumstances. Sometimes we want agents to cooperate, sometimes not. As a field we want to be able to understand when and how cooperation between agents emerges, and to control the extent to which agents cooperate in different kinds of situations. (See also this short thread which addresses related confusions.)
It was very foreseeable that AGI companies would try to get multi-agent training to work for LLMs (because of the performance benefits over naive orchestration regimes). Unfortunately, this increases the risks of things like steganographic collusion, joint reward hacking, unpredictable feedback loops, etc.
Final plug: as is probably evident from the recent string of security incidents and other results in this space, multi-agent safety is an area that we (the AI safety community) are especially not on top of. If you want to help us be more on top of it, we (the Cooperative AI Foundation) are making grants and hiring/mentoring people in this area. Feel free to reach out!