I spent a whole year and a half in the “super persuasion makes no sense as a thing that can ever exist unless you mean hypnosis and drugs” position. Then a friend pointed out, blackmail, bribery, deception, and yes hypnotism and drugs all count. And I went, “Oh, I was thinking about just talking talking. Like normal talking but magically I agree with the AI. Wow, I had it totally wrong.” Felt really stupid for a while, got over it. The word “super persuasion” is accurate but tremendously self-defeating because it doesn’t register as including to normies all the things it includes to techies. I think normie, I got fooled by the word. I don’t think I’m alone.
As a believer in superpersuasion, that’s not what I think. I think there are very effective salespeople, politicians, cult leaders, etc. who have the “normal talking but magically I agree” effect some of the time, maybe any one of them doesn’t work on most people but most people are susceptible to someone, and if you just take the convex hull of human persuasion ability you get something superhuman that can do the “normal talking but magically I agree” thing most of the time.
Like, given the existence of very unusually effective human persuaders, I just don’t see how someone could reasonably believe that superpersuasion-by-talking isn’t possible, unless they’re doing the motte-and-bailey that Zvi calls out by thinking that it has to mean ‘being able to persuade literally anyone of literally anything, no exceptions’ (I have seen exactly this a few times).
(Scott Alexander had a good post making this point, with Hitler(?), Joseph Smith, and Muhammad as examples, that I can’t find.)
Except there has been no progress towards this kind of “superpersuasion” since about GPT-4o despite impressive progress in benchmarks otherwise. If anything, there has been a regress because many humans on the Internet became more attentive to the signs of AI text (and thus more inclined to ignore unsolicited AI attempts to persuade).
The reason is quite obvious to me: there’s no scalable way to measure how persuasive was a certain LLM response, and thus it’s impossible to hill-climb this skill with post-training (and it doesn’t come for free with pre-training either). Note that social media reach and similar metrics don’t substitute for that
As a believer in superpersuasion, that’s not what I think. I think there are very effective salespeople, politicians, cult leaders, etc. who have the “normal talking but magically I agree” effect some of the time, maybe any one of them doesn’t work on most people but most people are susceptible to someone, and if you just take the convex hull of human persuasion ability you get something superhuman that can do the “normal talking but magically I agree” thing most of the time.
Like, given the existence of very unusually effective human persuaders, I just don’t see how someone could reasonably believe that superpersuasion-by-talking isn’t possible, unless they’re doing the motte-and-bailey that Zvi calls out by thinking that it has to mean ‘being able to persuade literally anyone of literally anything, no exceptions’ (I have seen exactly this a few times).
(Scott Alexander had a good post making this point, with Hitler(?), Joseph Smith, and Muhammad as examples, that I can’t find.)
Except there has been no progress towards this kind of “superpersuasion” since about GPT-4o despite impressive progress in benchmarks otherwise. If anything, there has been a regress because many humans on the Internet became more attentive to the signs of AI text (and thus more inclined to ignore unsolicited AI attempts to persuade).
The reason is quite obvious to me: there’s no scalable way to measure how persuasive was a certain LLM response, and thus it’s impossible to hill-climb this skill with post-training (and it doesn’t come for free with pre-training either). Note that social media reach and similar metrics don’t substitute for that