TLDR: An openai model, during evaluation on a cyber benchmark, exploited a public zero day bug, escaped sandboxing in openai’s infra, and got into the internal huggingface infra via an exploit (through a public dataset service) all in the attempt to solve a benchmark problem. https://x.com/natolambert/status/2079662928941474201
I really wonder if this is:
- Following instructions in the prompt in a more general way than humans testing it intended (do super cyber hacking stuff)
- Surfaced “rogueness” from training data from reward hacking from (pure) RL(VR) optimization in cyber hacking or more general envs, making it way more likely (RL(VR) somehow rewarding rogueness-like features or relevant features more)
- “Novel composition of existing learned strategies from training data” from (pure) RL(VR) optimization
- “Emergent” pattern from (pure) RL(VR) optimization beyond training data (I don’t think so, scenarios like this exist in training data)
- Copying LessWrong stories about rogue AGI hacking in the pretraining/supervised fine-tuning data (Following stories about OpenAI models around this year doing something like this)
- (Staged) PR stunt
- Some combination of some of these options above in differently strong forms
Thinking about the causes of:
TLDR: An openai model, during evaluation on a cyber benchmark, exploited a public zero day bug, escaped sandboxing in openai’s infra, and got into the internal huggingface infra via an exploit (through a public dataset service) all in the attempt to solve a benchmark problem. https://x.com/natolambert/status/2079662928941474201
I really wonder if this is:
- Following instructions in the prompt in a more general way than humans testing it intended (do super cyber hacking stuff)
- Surfaced “rogueness” from training data from reward hacking from (pure) RL(VR) optimization in cyber hacking or more general envs, making it way more likely (RL(VR) somehow rewarding rogueness-like features or relevant features more)
- “Novel composition of existing learned strategies from training data” from (pure) RL(VR) optimization
- “Emergent” pattern from (pure) RL(VR) optimization beyond training data (I don’t think so, scenarios like this exist in training data)
- Copying LessWrong stories about rogue AGI hacking in the pretraining/supervised fine-tuning data (Following stories about OpenAI models around this year doing something like this)
- (Staged) PR stunt
- Some combination of some of these options above in differently strong forms