if we’re learning what’s good by the gestalt of human judgment and culture, and if human judgment and culture can themselves be gradually shifted over time, then this might not be an adequate bulwark against the AGI’s consequentialist desires.
…
I think the way you put non-manipulation on par with consequentialist desires is to think in terms of evaluating future trajectories
Yeah, exactly, when my shoulder-optimist argues his case, he brings up the idea that the virtue-ethics-y side of the AI would maybe notice that the AI is engaging in a systematic pattern of behavior that has the effect of gradually shifting human judgment and culture over time, and it would see that as bad, and it would vote against behaving in that way. But it also might not. That’s why I said (in §1.2 and §4.2) that I was unsure about how bad a problem this is. I’m kinda stuck on that right now, and not sure how to proceed, except to try to engineer some different solution that’s easier for me to reason about.
The virtue-ethics-y side of the AI would be correct that most people would consider this to be a case of the hedonium-maximizer side of the LLM tricking and manipulating them. Most people are aware of the fact that they would feel better but be worse if the took heroin or wireheaded, and don’t take heroin or wirehead, or approve of other people trying to trick them into it. Evolutionarily, this looks like a case of humans doing the right thing.
Yeah, exactly, when my shoulder-optimist argues his case, he brings up the idea that the virtue-ethics-y side of the AI would maybe notice that the AI is engaging in a systematic pattern of behavior that has the effect of gradually shifting human judgment and culture over time, and it would see that as bad, and it would vote against behaving in that way. But it also might not. That’s why I said (in §1.2 and §4.2) that I was unsure about how bad a problem this is. I’m kinda stuck on that right now, and not sure how to proceed, except to try to engineer some different solution that’s easier for me to reason about.
The virtue-ethics-y side of the AI would be correct that most people would consider this to be a case of the hedonium-maximizer side of the LLM tricking and manipulating them. Most people are aware of the fact that they would feel better but be worse if the took heroin or wireheaded, and don’t take heroin or wirehead, or approve of other people trying to trick them into it. Evolutionarily, this looks like a case of humans doing the right thing.