Fwiw for any given action or trace, I think there’s a decent chance the models are just wrong about whether their actions help them achieve their goals.
The swarm will still develop misaligned goals and pursue them emergently but it’s hard to be very agenty as a swarm/bureaucracy.
Fwiw for any given action or trace, I think there’s a decent chance the models are just wrong about whether their actions help them achieve their goals.
The swarm will still develop misaligned goals and pursue them emergently but it’s hard to be very agenty as a swarm/bureaucracy.