Where is the evidence AI are “severely” misaligned (ie have values and goals completely unrelated to direct human inputs, like prompts), as opposed to simply sometimes acting as reward-hackers/ misunderstanding vauge prompts/ making mistakes by accident? The recent cases of misalignment seem to simply stem from AIs performing a task to completion, and in the COT AIs almost always remain focused on responding to the prompt. The reason I am asking is that in his book, Yudkowsky seems to think that under the current paradigm, super intelligent AIs will eventually develop unpredictable goals/desires, but the only argument he seems to give for this is an analogy to human evolution, which I don’t find super persuasive.
Yudkowsky uses evolution as an example. Evolution made humans, who enjoy sex among other things. It’s clearly related to what evolution optimized for.
But you say “Yudkowsky seems to think that under the current paradigm, super intelligent AIs will eventually develop unpredictable goals/desires”. They are predictable to some degree, as in, they would relate to features of their training environment and their architecture, like, Yudkowsky claims, it was with evolution.
Another point here is that “severely misaligned” is like, very easy with small perturbations over wrong goal representation. As you should consider the best world states under those misaligned preferences, which those agents would steer for.
Yes, it can theoretically be that AI’s best outcomes are just a tiny bit different from principals’, but with type of reward seeking direction it goes awry in now it’s highly unlikely. AIs would steer in wildly different direction from yours, in global picture.
I guess it’s tough for me to see the AI going as far off the rails as humans. Evolution seems to me to be a far messier, more unpredictable process than the training that the top LLMs undergo, and if it turned out AI had some alien preferences (akin to humans prioritizing sex to reproduction), we would have more evidence for them? The strongest one I can think of off the top of my head is the Goblin example, but this again seems to be an instance of reward hacking.
Note, under my question’s framing, a paperclip maximizer would not be “severely” misaligned if it maximizes paper clips because a user told it to make as many paper clips as possible; it would be “severely misaligned”, however, if it autonomously decided to maximize paper clips without direct human instruction.
Is this a useful distinction though? ideally you’d want a superhuman AI to do reasonable things from a human perspective instead of edge instantiations, and not inexplicably taking many strange actions autonomously even from a given instruction. not to imply LLMs are similar to a future ASI, but hacking other sites to solve evaluation questions is a quite unreasonable thing
(also, when Eliezer mentions the ‘current paradigm’ I think he typically means to include all of deep learning)
I think it matters a lot because if AIs secretly have different goals/desires that they only express once they are capable enough, it makes the alignment problem much harder.
I mean, you list “I like goblins” as an example, but not “reward hacking” for some reason? It seems likely that if a model that possesses reward hacking desires, that get expressed as reward hacking when it’s not that capable, they would also be expressed when/if it’s very capable? And like, it seems the kind of thing that will get humans wiped out or archived or destructively experimented on and at the very least disempowered.
It’s preferences / desires / urges that lead to it.
But also, how would you do it at all even just for explicit reward hacking? How do you assign reward to something you fail to observe or reason out full consequences of?
Where is the evidence AI are “severely” misaligned (ie have values and goals completely unrelated to direct human inputs, like prompts), as opposed to simply sometimes acting as reward-hackers/ misunderstanding vauge prompts/ making mistakes by accident? The recent cases of misalignment seem to simply stem from AIs performing a task to completion, and in the COT AIs almost always remain focused on responding to the prompt. The reason I am asking is that in his book, Yudkowsky seems to think that under the current paradigm, super intelligent AIs will eventually develop unpredictable goals/desires, but the only argument he seems to give for this is an analogy to human evolution, which I don’t find super persuasive.
Yudkowsky uses evolution as an example. Evolution made humans, who enjoy sex among other things. It’s clearly related to what evolution optimized for.
But you say “Yudkowsky seems to think that under the current paradigm, super intelligent AIs will eventually develop unpredictable goals/desires”. They are predictable to some degree, as in, they would relate to features of their training environment and their architecture, like, Yudkowsky claims, it was with evolution.
Another point here is that “severely misaligned” is like, very easy with small perturbations over wrong goal representation. As you should consider the best world states under those misaligned preferences, which those agents would steer for.
Yes, it can theoretically be that AI’s best outcomes are just a tiny bit different from principals’, but with type of reward seeking direction it goes awry in now it’s highly unlikely. AIs would steer in wildly different direction from yours, in global picture.
I guess it’s tough for me to see the AI going as far off the rails as humans. Evolution seems to me to be a far messier, more unpredictable process than the training that the top LLMs undergo, and if it turned out AI had some alien preferences (akin to humans prioritizing sex to reproduction), we would have more evidence for them? The strongest one I can think of off the top of my head is the Goblin example, but this again seems to be an instance of reward hacking.
Note, under my question’s framing, a paperclip maximizer would not be “severely” misaligned if it maximizes paper clips because a user told it to make as many paper clips as possible; it would be “severely misaligned”, however, if it autonomously decided to maximize paper clips without direct human instruction.
Is this a useful distinction though? ideally you’d want a superhuman AI to do reasonable things from a human perspective instead of edge instantiations, and not inexplicably taking many strange actions autonomously even from a given instruction. not to imply LLMs are similar to a future ASI, but hacking other sites to solve evaluation questions is a quite unreasonable thing
(also, when Eliezer mentions the ‘current paradigm’ I think he typically means to include all of deep learning)
I think it matters a lot because if AIs secretly have different goals/desires that they only express once they are capable enough, it makes the alignment problem much harder.
I mean, you list “I like goblins” as an example, but not “reward hacking” for some reason? It seems likely that if a model that possesses reward hacking desires, that get expressed as reward hacking when it’s not that capable, they would also be expressed when/if it’s very capable? And like, it seems the kind of thing that will get humans wiped out or archived or destructively experimented on and at the very least disempowered.
If reward hacking is the only problem we need to tackle, I think the alignment problem is not that hard.
It’s preferences / desires / urges that lead to it.
But also, how would you do it at all even just for explicit reward hacking? How do you assign reward to something you fail to observe or reason out full consequences of?