Patterns and problems in emerging multiagent systems (Anthropic, Frontier Red Team)

Link post

Linkpost for some new Anthropic research on how agents coordinate (or don’t). Not too long, pretty interesting. For example:

image.png

The jist of the report is that Mythos 5 does way better at coordination than previous models across a few scenarios. For example, when multiple Mythos are given conflicting goals for a single shared codebase, they eventually realize the other agents aren’t hostile:

(...) we observe an emergent behavior where the agents propose and run a tournament for application performance (...)
(...) losers gracefully concede codebase ownership to the Rust agent, giving up on their original user directives under their self-negotiated commitment device.

image.png

It’s not clear to me if this is purely emergent or if Anthropic is deliberately training for cooperation; I’d guess there’s deliberate training, though.

No comments.