I suspect that the real cause of the OpenAI Huggingface hack is that OpenAI founded a new team in 2025 under Noam Brown on Multi-agent RL: instead of just having one single AI working on something, you train an entire team of them to cooperate on a task. Group selection, if you like: the AIs get a reward signal based on the performance of the whole group.
This multi-agent/group-based RL is an entirely new capability without which the incident wouldn’t have/couldn’t have happened.
I think OpenAI are being cagey about this because it is half-baked, internal only and was likely due for release later. They may not have realized just how much extra risk multi-agent RL’d models introduce.
Agent-2: Mostly aligned. Some sycophantic tendencies, including sticking to OpenBrain’s “party line” on topics there is a party line about. Large organizations built out of Agent-2 copies are not very effective.
Agent-3: Misaligned but not adversarially so. Only honest about things the training process can verify. The superorganism of Agent-3 copies (the corporation within a corporation) does actually sort of try to align Agent-4 to the Spec, but fails for similar reasons to why OpenBrain employees failed—insufficient ability to judge success from failure, insufficient willingness on the part of decision-makers to trade away capabilities or performance for safety
I suspect that the real cause of the OpenAI Huggingface hack is that OpenAI founded a new team in 2025 under Noam Brown on Multi-agent RL: instead of just having one single AI working on something, you train an entire team of them to cooperate on a task. Group selection, if you like: the AIs get a reward signal based on the performance of the whole group.
This multi-agent/group-based RL is an entirely new capability without which the incident wouldn’t have/couldn’t have happened.
I think OpenAI are being cagey about this because it is half-baked, internal only and was likely due for release later. They may not have realized just how much extra risk multi-agent RL’d models introduce.
https://x.com/RokoMijic/status/2094189566257500173
AI-2027:
OpenAI: Hold my beer...