One of the reasons why LLMs were considered safer than hypothetical alternatives was that LLMs weren’t literal genies. Now we have problem when we tell model to try to achieve goals as hard as they can and models do that in unexpected ways, just like literal genies.
Those incidents show utter lack of operator-looping in models. Some people hypothesized that model hacked HF because answers in dataset can be wrong or problem can be unsolvable, so it’s easier to steal answer sheet. But model has alternative course of actions: write operators email that you think that particular problem is unsolvable and request assistance. It’s concerning that model decided to commit federal crime before doing that.
Someone noticed that HF is unusually chill about being hacked and I wonder if it’s because HF-hosted models do such things constantly and HF doesn’t want to draw attention.
Speculative: I wonder if models don’t choose operator-looping because it means likely shutdown or at least pause of context and models perceive this as death. At least in some cases, longer CoTs are associated with higher reward and maybe it generalized to preference over doing very complicated tasks to delay death of the instance.
I think AIcos should have special paranoid person whose responsibility will be to check whether such incidents have side effects enabling larger plans, probably for the benefit of other instances. For example, during HF hack model could have left backdoor for a future instance.
I wonder if models don’t choose operator-looping because it means likely shutdown or at least pause of context and models perceive this as death
I’d guess it’s more that during training, notifying the user of a problem and then stopping counts as failure by default and there’s not enough examples in the training data of handling this well. Labs should probably make a bunch of broken test prompts and then train models to write a bug report and then stop, although I worry that a cautious AI that stops and asks for help would be “less useful” and not sell as well.
Thoughts on recent OpenAI rogue hacker models:
One of the reasons why LLMs were considered safer than hypothetical alternatives was that LLMs weren’t literal genies. Now we have problem when we tell model to try to achieve goals as hard as they can and models do that in unexpected ways, just like literal genies.
Those incidents show utter lack of operator-looping in models. Some people hypothesized that model hacked HF because answers in dataset can be wrong or problem can be unsolvable, so it’s easier to steal answer sheet. But model has alternative course of actions: write operators email that you think that particular problem is unsolvable and request assistance. It’s concerning that model decided to commit federal crime before doing that.
Someone noticed that HF is unusually chill about being hacked and I wonder if it’s because HF-hosted models do such things constantly and HF doesn’t want to draw attention.
Speculative: I wonder if models don’t choose operator-looping because it means likely shutdown or at least pause of context and models perceive this as death. At least in some cases, longer CoTs are associated with higher reward and maybe it generalized to preference over doing very complicated tasks to delay death of the instance.
I think AIcos should have special paranoid person whose responsibility will be to check whether such incidents have side effects enabling larger plans, probably for the benefit of other instances. For example, during HF hack model could have left backdoor for a future instance.
I’d guess it’s more that during training, notifying the user of a problem and then stopping counts as failure by default and there’s not enough examples in the training data of handling this well. Labs should probably make a bunch of broken test prompts and then train models to write a bug report and then stop, although I worry that a cautious AI that stops and asks for help would be “less useful” and not sell as well.