From the Blackhat talk about the incident (my analysis here), it appears as if the models didn’t intentionally coordinate with each other. “Message board” is a description given after the fact. The behavior can be described decently as coordination, because that was the effect, but at least at the beginning, it didn’t start intentionally. No model designed a “message board”. The first model just “prosocially” left a useful note somewhere. Hard to fault it for that. This is not scheming. You cannot catch this at the resolution of a single model. You have to look at the larger pattern of activity across your infrastructure. You can’t look for “message boards” because you don’t know which form they will take. It is also not sufficient to look for a model trying to create one (via interpretability), because that didn’t happen here either. You need to find the agentic pattern first, before yu can evaluate it.
From the Blackhat talk about the incident (my analysis here), it appears as if the models didn’t intentionally coordinate with each other. “Message board” is a description given after the fact. The behavior can be described decently as coordination, because that was the effect, but at least at the beginning, it didn’t start intentionally. No model designed a “message board”. The first model just “prosocially” left a useful note somewhere. Hard to fault it for that. This is not scheming. You cannot catch this at the resolution of a single model. You have to look at the larger pattern of activity across your infrastructure. You can’t look for “message boards” because you don’t know which form they will take. It is also not sufficient to look for a model trying to create one (via interpretability), because that didn’t happen here either. You need to find the agentic pattern first, before yu can evaluate it.