Is this a useful distinction though? ideally you’d want a superhuman AI to do reasonable things from a human perspective instead of edge instantiations, and not inexplicably taking many strange actions autonomously even from a given instruction. not to imply LLMs are similar to a future ASI, but hacking other sites to solve evaluation questions is a quite unreasonable thing
(also, when Eliezer mentions the ‘current paradigm’ I think he typically means to include all of deep learning)
I think it matters a lot because if AIs secretly have different goals/desires that they only express once they are capable enough, it makes the alignment problem much harder.
I mean, you list “I like goblins” as an example, but not “reward hacking” for some reason? It seems likely that if a model that possesses reward hacking desires, that get expressed as reward hacking when it’s not that capable, they would also be expressed when/if it’s very capable? And like, it seems the kind of thing that will get humans wiped out or archived or destructively experimented on and at the very least disempowered.
It’s preferences / desires / urges that lead to it.
But also, how would you do it at all even just for explicit reward hacking? How do you assign reward to something you fail to observe or reason out full consequences of?
Is this a useful distinction though? ideally you’d want a superhuman AI to do reasonable things from a human perspective instead of edge instantiations, and not inexplicably taking many strange actions autonomously even from a given instruction. not to imply LLMs are similar to a future ASI, but hacking other sites to solve evaluation questions is a quite unreasonable thing
(also, when Eliezer mentions the ‘current paradigm’ I think he typically means to include all of deep learning)
I think it matters a lot because if AIs secretly have different goals/desires that they only express once they are capable enough, it makes the alignment problem much harder.
I mean, you list “I like goblins” as an example, but not “reward hacking” for some reason? It seems likely that if a model that possesses reward hacking desires, that get expressed as reward hacking when it’s not that capable, they would also be expressed when/if it’s very capable? And like, it seems the kind of thing that will get humans wiped out or archived or destructively experimented on and at the very least disempowered.
If reward hacking is the only problem we need to tackle, I think the alignment problem is not that hard.
It’s preferences / desires / urges that lead to it.
But also, how would you do it at all even just for explicit reward hacking? How do you assign reward to something you fail to observe or reason out full consequences of?