I mean, you list “I like goblins” as an example, but not “reward hacking” for some reason? It seems likely that if a model that possesses reward hacking desires, that get expressed as reward hacking when it’s not that capable, they would also be expressed when/if it’s very capable? And like, it seems the kind of thing that will get humans wiped out or archived or destructively experimented on and at the very least disempowered.
It’s preferences / desires / urges that lead to it.
But also, how would you do it at all even just for explicit reward hacking? How do you assign reward to something you fail to observe or reason out full consequences of?
I mean, you list “I like goblins” as an example, but not “reward hacking” for some reason? It seems likely that if a model that possesses reward hacking desires, that get expressed as reward hacking when it’s not that capable, they would also be expressed when/if it’s very capable? And like, it seems the kind of thing that will get humans wiped out or archived or destructively experimented on and at the very least disempowered.
If reward hacking is the only problem we need to tackle, I think the alignment problem is not that hard.
It’s preferences / desires / urges that lead to it.
But also, how would you do it at all even just for explicit reward hacking? How do you assign reward to something you fail to observe or reason out full consequences of?