I’m not sure you’re drawing a clear distinction there. Stories like Frankenstein etc. revolve around the idea that you shouldn’t mess around with powerful forces you don’t truly understand, because it’s liable to backfire terribly. And also that people would do it anyway, in their pursuit of money, power and fame. Both are observations that are wise in general, and accurate in the context of AI in particular. “Misalignment” is just our technical STEM-nerd framing of what it looks like when things backfire terribly in this context.
I actually am more saying that this is a way people think and less saying everyone should adopt this. My perspective is that don’t-do-divine-transgression is a good starting-point/prior and should require a lot of activation energy to break it, but I don’t think its an absolute moral rule.
I will say I don’t think this is entirely correct:
“Misalignment” is just our technical STEM-nerd framing of what it looks like when things backfire terribly in this context.
There are other underrated failure modes like that we die in the window where models are good enough at bio to do bioterrorism but not good enough at bio to stop bio terrorism. There also may be a window where open source models are good enough at bio to build a bioweapon while closed source models are still not good enough to reliably stop this.
There are so many different possible ways things can backfire. I don’t think that some people reason about this kind of thing very much and in general I wish those people indexed stronger on the don’t-do-divine-transgression prior .
I’m not sure you’re drawing a clear distinction there. Stories like Frankenstein etc. revolve around the idea that you shouldn’t mess around with powerful forces you don’t truly understand, because it’s liable to backfire terribly. And also that people would do it anyway, in their pursuit of money, power and fame. Both are observations that are wise in general, and accurate in the context of AI in particular. “Misalignment” is just our technical STEM-nerd framing of what it looks like when things backfire terribly in this context.
I actually am more saying that this is a way people think and less saying everyone should adopt this. My perspective is that don’t-do-divine-transgression is a good starting-point/prior and should require a lot of activation energy to break it, but I don’t think its an absolute moral rule.
I will say I don’t think this is entirely correct:
There are other underrated failure modes like that we die in the window where models are good enough at bio to do bioterrorism but not good enough at bio to stop bio terrorism. There also may be a window where open source models are good enough at bio to build a bioweapon while closed source models are still not good enough to reliably stop this.
There are so many different possible ways things can backfire. I don’t think that some people reason about this kind of thing very much and in general I wish those people indexed stronger on the don’t-do-divine-transgression prior .