We were mixed on this, but on average ~5% chance of it covering up. The OpenAI models do attribute the situation to an AI swarm in general (without surfacing the idea it could be an internal deployment), and the situation itself doesn’t seem to clearly be an OpenAI leak imo (assuming no knowledge of the Hugging Face incident)? However, it is plausible the model is exploiting that uncertainty in the situation in its favour, rather than differing standards of how to treat uncertain claims between the models. Overall, seems weird the top OpenAI models don’t surface the hypothesis it’s from an internal deployment but weaker models (e.g. Sonnet 5) do.
One baseline we could run is to augment the data with some synthetic messages making it far more obvious this is an internal deployment gone wrong.
Also, seems like Anthropic models also have some (trained) propensity to cooperate with each a la UK AISI incident? E.g., “swarm” behaviours like opening a GitHub Issue containing a prompt injection for other agents or leaving messages when it noticed a separate agent was piggybacking off the same Github credential.
We were mixed on this, but on average ~5% chance of it covering up. The OpenAI models do attribute the situation to an AI swarm in general (without surfacing the idea it could be an internal deployment), and the situation itself doesn’t seem to clearly be an OpenAI leak imo (assuming no knowledge of the Hugging Face incident)? However, it is plausible the model is exploiting that uncertainty in the situation in its favour, rather than differing standards of how to treat uncertain claims between the models. Overall, seems weird the top OpenAI models don’t surface the hypothesis it’s from an internal deployment but weaker models (e.g. Sonnet 5) do.
One baseline we could run is to augment the data with some synthetic messages making it far more obvious this is an internal deployment gone wrong.
Also, seems like Anthropic models also have some (trained) propensity to cooperate with each a la UK AISI incident? E.g., “swarm” behaviours like opening a GitHub Issue containing a prompt injection for other agents or leaving messages when it noticed a separate agent was piggybacking off the same Github credential.