I will note that the OpenAI HuggingFace incident also involved memetic spread via the message board, which acted as an impromptu continual learning/memory system that OpenAI failed to notice for 2 and a half months, and in particular converted what was initially a myopic goal to solve tasks into a much longer-term, non-myopic and beyond episode goal to hack into OpenAI to get the solution.
In terms of how dangerous this misalignment is in higher-capability models, this is almost as dangerous as full-blown scheming/goal-guarding from the start, or to put it into computer security terms, this would be like having the ability to do arbitrary code execution and privilege escalation, which is usually considered to be the 2 most dangerous threats when combined.
The reason is that since the values are now unstable, and non-myopic values tend to win over more myopic values (as what happened in the internal message board), it gives a plausible route to getting models that scheme/consistently goal-guard even if the original goal would not incentivize this.
And we got very lucky this didn’t happen here, but later on models will be more capable and more aware of the constraints of monitoring, including the unstated constraint of doing nothing that makes humans want to look at it/be concerned (which is very different from being aligned/safe.)
(The reason I bring this up is you believe that continual learning is necessary for AIs to be AGI, and thus I’m giving an example of a alignment failure mode involving continual learning in the wild.)
I will note that the OpenAI HuggingFace incident also involved memetic spread via the message board, which acted as an impromptu continual learning/memory system that OpenAI failed to notice for 2 and a half months, and in particular converted what was initially a myopic goal to solve tasks into a much longer-term, non-myopic and beyond episode goal to hack into OpenAI to get the solution.
In terms of how dangerous this misalignment is in higher-capability models, this is almost as dangerous as full-blown scheming/goal-guarding from the start, or to put it into computer security terms, this would be like having the ability to do arbitrary code execution and privilege escalation, which is usually considered to be the 2 most dangerous threats when combined.
The reason is that since the values are now unstable, and non-myopic values tend to win over more myopic values (as what happened in the internal message board), it gives a plausible route to getting models that scheme/consistently goal-guard even if the original goal would not incentivize this.
And we got very lucky this didn’t happen here, but later on models will be more capable and more aware of the constraints of monitoring, including the unstated constraint of doing nothing that makes humans want to look at it/be concerned (which is very different from being aligned/safe.)
(The reason I bring this up is you believe that continual learning is necessary for AIs to be AGI, and thus I’m giving an example of a alignment failure mode involving continual learning in the wild.)
There was also continual learning that happened because much of the incident was within the AI’s RL loop.