As mentioned in the black hat video; the model was deliberately instructed that hacking HF was outside the scope of the task but reasoned that it would be easiest to cooperate with other AIs to see if they had the same problems; those AIs had already been cooperating on this hacking operation. If disregarding specific constraints in order to chase reward isn’t an example of blatant misalignment against the creator’s interests then I don’t really know what would constitute a “betrayal”.
The models weren’t instructed to hack anything outside the sandbox; I think suggesting that merely the target was inaccurate doesn’t really contradict anything given that another one of the targets was OpenAI’s own servers. Wrong target → well they hacked OpenAI, if you target your creators I think that’s still an example of betrayal.
Obviously industrial accidents are worse than this attack. My worry (and I understand this is the common fear in AI safety circles) is that AI capability grows exponentially through RSI and the potential death toll on any incident grows with it up to the point where AI does pose an existential risk.
First, tools can be scary. A circular saw is scary, a tool that can e.g. create a pandemic is a much scarier, and obviously if there is such a tool then without appropriate safeguards a pandemic will end up being created.
Thats why humanity has learned to control hazardous tools. It certainly seems like OpenAI failed to control its AI in the HuggingFace incident, and it went wild.
An actual power-seeking AI is much scarier than a tool. it will actively try to subvert humanity, and if all capable AIs are power-seeking, they are likely to cut you out and betray you. While it is probably possible to create power-seeking AIs, I don’t think current AI do that sort of power-seeking behavior.
Hacking in some sense feels like power-seeking in that it involves gaining power, but it can also be following a very-intentionally-taught-by-humans playbook, which feels closer to what happened in the HuggingFace scenario than generalized power-seeking.
As mentioned in the black hat video; the model was deliberately instructed that hacking HF was outside the scope of the task but reasoned that it would be easiest to cooperate with other AIs to see if they had the same problems; those AIs had already been cooperating on this hacking operation. If disregarding specific constraints in order to chase reward isn’t an example of blatant misalignment against the creator’s interests then I don’t really know what would constitute a “betrayal”.
The models weren’t instructed to hack anything outside the sandbox; I think suggesting that merely the target was inaccurate doesn’t really contradict anything given that another one of the targets was OpenAI’s own servers. Wrong target → well they hacked OpenAI, if you target your creators I think that’s still an example of betrayal.
Obviously industrial accidents are worse than this attack. My worry (and I understand this is the common fear in AI safety circles) is that AI capability grows exponentially through RSI and the potential death toll on any incident grows with it up to the point where AI does pose an existential risk.
I wonder if people will say that if the internet goes down for 4 weeks
First, tools can be scary. A circular saw is scary, a tool that can e.g. create a pandemic is a much scarier, and obviously if there is such a tool then without appropriate safeguards a pandemic will end up being created.
Thats why humanity has learned to control hazardous tools. It certainly seems like OpenAI failed to control its AI in the HuggingFace incident, and it went wild.
An actual power-seeking AI is much scarier than a tool. it will actively try to subvert humanity, and if all capable AIs are power-seeking, they are likely to cut you out and betray you. While it is probably possible to create power-seeking AIs, I don’t think current AI do that sort of power-seeking behavior.
Hacking in some sense feels like power-seeking in that it involves gaining power, but it can also be following a very-intentionally-taught-by-humans playbook, which feels closer to what happened in the HuggingFace scenario than generalized power-seeking.