I agree with the bottom-line conclusion “AIs probably can’t convince you to kill your family in a short conversation” [1]but disagree with a lot of the analysis and the implied moral (therefore AI persuasion isn’t that big a deal).
For starters, AIs aren’t in a box. They’re like so not in a box like you won’t believe, including internally deployed models. In practice, how much AI persuasion is a big deal depends both on their underlying capabilities and how much affordances we give them or are likely to give them (and other things too, of course). And my contention is that in practice we give them enough affordances to pose a major risk, at “reasonable” levels of peak human persuasion or above.
Also I think your threat model is overly epistemic and direct relative to how humans in practice change their minds about things.
For example, consider the standard model of success in (military) coups. A standard non-persuasion story is that coup are like battles (about hard military advantages and tactical victories). A naive persuasion story is that coups are like elections (about winning the hearts and minds of either the people in general, or people in the military specifically). But the Singh model (which I take to be the dominant/standard model these days) is that coups are like coordination games: you win a coup by convincing enough of the military that other people support the coup: ie, that your preferred outcome is the inevitable one. I think a lot of persuasion stories look like this broadly speaking: that there’s one or more unexpected layers of indirection between the persuaders’ preferred outcomes and the mind-states they want people to hold.
I also think there are many levels between “charismatic leader” and “full-on hypnotic suggestion.” For example, very few charismatic leaders can regularly get people to challenge religions or give up half their wealth, yet this is a perfectly normal thing to do in the context of a romantic relationship. This suggests to me that relationships are a common way humans convince each other of things, and relationships have historically been very unscalable (pre-AI).
I agree with the bottom-line conclusion “AIs probably can’t convince you to kill your family in a short conversation” [1]but disagree with a lot of the analysis and the implied moral (therefore AI persuasion isn’t that big a deal).
For starters, AIs aren’t in a box. They’re like so not in a box like you won’t believe, including internally deployed models. In practice, how much AI persuasion is a big deal depends both on their underlying capabilities and how much affordances we give them or are likely to give them (and other things too, of course). And my contention is that in practice we give them enough affordances to pose a major risk, at “reasonable” levels of peak human persuasion or above.
Also I think your threat model is overly epistemic and direct relative to how humans in practice change their minds about things.
For example, consider the standard model of success in (military) coups. A standard non-persuasion story is that coup are like battles (about hard military advantages and tactical victories). A naive persuasion story is that coups are like elections (about winning the hearts and minds of either the people in general, or people in the military specifically). But the Singh model (which I take to be the dominant/standard model these days) is that coups are like coordination games: you win a coup by convincing enough of the military that other people support the coup: ie, that your preferred outcome is the inevitable one. I think a lot of persuasion stories look like this broadly speaking: that there’s one or more unexpected layers of indirection between the persuaders’ preferred outcomes and the mind-states they want people to hold.
I also think there are many levels between “charismatic leader” and “full-on hypnotic suggestion.” For example, very few charismatic leaders can regularly get people to challenge religions or give up half their wealth, yet this is a perfectly normal thing to do in the context of a romantic relationship. This suggests to me that relationships are a common way humans convince each other of things, and relationships have historically been very unscalable (pre-AI).
I discuss why I think cognitive exploits are unlikely here: https://www.lesswrong.com/posts/s58hDHX2GkFDbpGKD/linch-s-shortform#qGKdocPXuqpiezcvn