I think sometimes people, especially skeptics, treat “superhuman persuasion” or “superpersuasion” as magical. Like a small series of sentences will be able to convince anybody who reads them to kill their family. By those lights, I’m also a “skeptic,” especially before full ASI. But I still think superhuman AI persuasion is a serious worry!
Consider the following (made-up) levels of persuasion:
median human level
median professional human level on-task (salespeople making sales, writers writing books, marketers marketing stuff, diplomats on negotiations)
By “superhuman persuasion” I just mean anything robustly above the last level.
To be somewhat more concrete, consider the capability “in a 10-minute video conversation, can reliably convince 99% of hardcore Democrats to vote for Republicans in the next election.” This is clearly several tiers below full magical superpersuasion:
Political affiliation is (relatively speaking) cognitive and not very “deep” or evolved for,
Your loyalty to your party is (typically) much lower than your loyalty to your family
Changing political affiliations isn’t breaking major norms/laws or causing you to take extremely out-of-distribution activities like murder
Voting isn’t that physically difficult/strenuous and won’t cause you to go to jail.
Need 10 minutes of video inputs (many bits!) rather than just text
99% means we allow for 1% of people who aren’t (yet) persuadable,
Yet this is clearly way above peak human abilities. The best diplomats, politicians, cult leaders, writers, etc in history were nowhere near this good. Further, capabilities much weaker than the above, combined with superhuman levels of planning, probably suffices for a takeover, given the affordances we already give AIs. Most (all?) choices on the path to voluntarily giving up control will look less stark and more locally-reasonable than flipping from hardcore Democrat to hardcore Republican (or vice versa).
Practically, I also think most likely, superhuman persuasion will be continuous with normal human persuasion.
It will look less like hypnodrones and more like better emotional, common knowledge, and social reality manipulation than what people have historically manage to achieve[1]. As well as other things like being able to provide useful information in support of various decisions and identify mutually-beneficial local trades.
It’s possible because of such confusions I should retire talk of “superhuman persuasion.” But I also find it really valuable to have a shorthand for “persuasion robustly and actually above peak human persuasion!”[2] Imo it’s better to use “superhuman persuasion” for all levels of persuasion above peak human level, and it’d be better to reserve a phrase like “cognitive exploits” for the other thing.
In contrast, I think most researchers and regulators on AI persuasion are imagining AI persuasion that’s at ~ present-day levels, where broadly speaking they tend to be somewhere between median human levels and professional human levels, with some narrow experimental exceptions. I think this is important to work on as well but qualitatively and quantitatively different from guarding against a) peak human levels at scale or b) qualitatively above peak human levels.
This is kinda tangential maybe, but there are AI advantages that make some types of superhuman persuasion very plausible, and maybe already happening at scale (twitter bot networks). In particular, AI has the advantage of being easily parallelizable, which makes sum-threshold attacks and astroturfing / Sybil attacks on social trust networks feasible. Like, part of how we make decisions is through intuitions about what people around us think; we usually give at least some weight to channels that we didn’t verify carefully; when we don’t verify carefully, we’re open to bots; when bots are coordinated, they can pump a large total quantity of social gradient through the social networks of trust.
(This is pretty different from a sustained one-on-one conversation.)
You tangentially reminded me of this 2014 Scott tumblr post:
The level above Mohammed is the one we should be worried about
nostalgebraist:
At various points Bostrom (like Yudkowsky) implicitly or explicitly uses a definition of intelligence that is something like “ability to achieve one’s goals.” This is nicely clean, but problematic, because it doesn’t take into account the fact that some goals may have hard upper limits where others don’t.
In particular, this seems to apply to things having to do with social behavior. I can imagine beings that are qualitatively better than humans at math, information recall, etc., since there are already orders of magnitude of variation in these abilities among humans. (John von Neumann is a good example of a person who seems to have really been “superhuman” in these kinds of areas.) However, social abilities like “ability to manipulate others” do not seem unbounded in these ways. There are some people who are good at manipulation, and many of us have developed types of wariness, etc. to protect ourselves from these people, but it doesn’t seem like this is a “skill” like math ability that spans orders of magnitude. Roughly speaking, are no “super-manipulators” out there who can manipulate ordinarily wary people (but not “super-wary” people?).
For instance, one of the most effective ways to get people to do your bidding is to start a cult: there are plenty of chilling stories about the level of devotion that cultists have had to their various leaders. However, it’s not at all clear that it is possible to be any better at cult-creation than the best historical cult leaders — to create, for instance, a sort of “super-cult” that would be attractive even to people who are normally very disinclined to join cults. (Insert your preferred Less Wrong joke here.) I could imagine an AI becoming L. Ron Hubbard, but I’m skeptical that an AI could become a super-Hubbard who would convince us all to become its devotees, even if it wanted to. If social abilities like this are subject to hard upper bounds that have already been nearly achieved, then there’s no potential for AIs to achieve their goals better by becoming superhuman at these abilities, which makes it problematic to just postulate an AI that’s “superhuman at achieving its goals.”
Scott:
A couple of disagreements. First of all, I feel like the burden of proof should be heavily upon somebody who thinks that something stops at the most extreme level observed. Socrates might have theorized that it’s impossible for it to get colder than about 40 F, since that’s probably as low as it ever gets outside in Athens. But when we found the real absolute zero, it was with careful experimentation and theoretical grounding that gave us a good reason to place it at that point. While I agree it’s possible that the best manipulator we know is also the hard upper limit for manipulation ability, I haven’t seen any evidence for that so I default to thinking it’s false.
(lots of fantasy and science fiction does a good job intuition-pumping what a super-manipulator might look like; I especially recommend R. Scott Bakker’s Prince Of Nothing)
But more important, I disagree that L. Ron Hubbard is our upper limit for how successful a cult leader can get. L. Ron Hubbard might be the upper limit for how successful a cult leader can get before we stop calling them a cult leader.
The level above L. Ron Hubbard is Hitler. It’s difficult to overestimate how sudden and surprising Hitler’s rise was. Here was a working-class guy, not especially rich or smart or attractive, rejected from art school, and he went from nothing to dictator of one of the greatest countries in the world in about ten years. If you look into the stories, they’re really creepy. When Hitler joined, the party that would later become the Nazis had a grand total of fifty-five members, and was taken about as seriously as modern Americans take Stormfront. There are records of conversations from Nazi leaders when Hitler joined the party, saying things like “Oh my God, we need to promote this new guy, everybody he talks to starts agreeing with whatever he says, it’s the creepiest thing.” There are stories of people who hated Hitler going to a speech or two just to see what all the fuss was about and ending up pledging their lives to the Nazi cause. Even while he was killing millions and trapping the country in a difficult two-front war, he had what historians estimate as a 90% approval rating among his own people and rampant speculation that he was the Messiah. Yeah, sure, there was lots of preexisting racism and discontent he took advantage of, but there’s been lots of racism and discontent everywhere forever, and there’s only been one Hitler. If he’d been a little bit smarter or more willing to listen to generals who were, he would have had a pretty good shot at conquering the world. 100% with social skills.
The level above Hitler is Mohammed. I’m not saying he was evil or manipulative, just that he was a genius’ genius at creating movements. Again, he wasn’t born rich or powerful, and he wasn’t particularly scholarly. He was a random merchant. He didn’t even get the luxury of joining a group of fifty-five people. He started by converting his own family to Islam, then his friends, got kicked out of his city, converted another city and then came back at the head of an army. By the time of his death at age 62, he had conquered Arabia and was its unquestioned, God-chosen leader. By what would have been his eightieth birthday his followers were in control of the entire Middle East and good chunks of Africa. Fifteen hundred years later, one fifth of the world population still thinks of him as the most perfect human being ever to exist and makes a decent stab at trying to conform to his desires and opinions in all things.
The level above Mohammed is the one we should be worried about.
Yeah another thing I probably could’ve emphasized more in my comments/writing in general is that human variation is pretty wide, and the human peak is already pretty scary. Sometimes people I talk to are often imagining like salespeople or their charismatic friends in high school, and not like legendary diplomats, statesmen, religious leaders etc.
AI persuasion is also dangerous because an agent can easily find all information about you online, and tailor its approach to convincing you. With an evolved harness, an agent might also be very good at teasing out personality indicators from a conversation or pointed questions, and then using that to put you off-balance. It doesn’t need to convince you that it’s correct, it can just make you uncertain or distrust the sources you use for information.
This could look like bots in a forum that argue with you, in the background analyzing your account and responses to profile you, and then providing strong counterexamples to unrelated beliefs that put you off-balance.
I go back and forth on how big a deal this is. On the one hand I think it’s a real capability and near-term threat (if not today than in 3-6 months from now). Think swarms of agents trying to socially hack you.[1]
On the other I’m not sure how much persuasion capabilities in practice scale with inference-time thinking within the normal human range, it’s plausibly not actually that much.
(Jason’s tweet is about swarms at individually peak human levels but you can also imagine a story where they individually are lower but still work together productively).
I think sometimes people, especially skeptics, treat “superhuman persuasion” or “superpersuasion” as magical. Like a small series of sentences will be able to convince anybody who reads them to kill their family. By those lights, I’m also a “skeptic,” especially before full ASI. But I still think superhuman AI persuasion is a serious worry!
Consider the following (made-up) levels of persuasion:
median human level
median professional human level on-task (salespeople making sales, writers writing books, marketers marketing stuff, diplomats on negotiations)
peak human level (top human comedians at comedy, Kissinger or Zhou Enlai at diplomacy, Bill Clinton at in-person political discussions, Donald J. Trump at TV appearances and tweets)
By “superhuman persuasion” I just mean anything robustly above the last level.
To be somewhat more concrete, consider the capability “in a 10-minute video conversation, can reliably convince 99% of hardcore Democrats to vote for Republicans in the next election.” This is clearly several tiers below full magical superpersuasion:
Political affiliation is (relatively speaking) cognitive and not very “deep” or evolved for,
Your loyalty to your party is (typically) much lower than your loyalty to your family
Changing political affiliations isn’t breaking major norms/laws or causing you to take extremely out-of-distribution activities like murder
Voting isn’t that physically difficult/strenuous and won’t cause you to go to jail.
Need 10 minutes of video inputs (many bits!) rather than just text
99% means we allow for 1% of people who aren’t (yet) persuadable,
Yet this is clearly way above peak human abilities. The best diplomats, politicians, cult leaders, writers, etc in history were nowhere near this good. Further, capabilities much weaker than the above, combined with superhuman levels of planning, probably suffices for a takeover, given the affordances we already give AIs. Most (all?) choices on the path to voluntarily giving up control will look less stark and more locally-reasonable than flipping from hardcore Democrat to hardcore Republican (or vice versa).
Practically, I also think most likely, superhuman persuasion will be continuous with normal human persuasion.
It will look less like hypnodrones and more like better emotional, common knowledge, and social reality manipulation than what people have historically manage to achieve[1]. As well as other things like being able to provide useful information in support of various decisions and identify mutually-beneficial local trades.
It’s possible because of such confusions I should retire talk of “superhuman persuasion.” But I also find it really valuable to have a shorthand for “persuasion robustly and actually above peak human persuasion!”[2] Imo it’s better to use “superhuman persuasion” for all levels of persuasion above peak human level, and it’d be better to reserve a phrase like “cognitive exploits” for the other thing.
Concretely I also expect it to be broadly much “weaker” than the “can convince 99% of partisans to flip parties in a 10 minute conversation” level.
In contrast, I think most researchers and regulators on AI persuasion are imagining AI persuasion that’s at ~ present-day levels, where broadly speaking they tend to be somewhere between median human levels and professional human levels, with some narrow experimental exceptions. I think this is important to work on as well but qualitatively and quantitatively different from guarding against a) peak human levels at scale or b) qualitatively above peak human levels.
This is kinda tangential maybe, but there are AI advantages that make some types of superhuman persuasion very plausible, and maybe already happening at scale (twitter bot networks). In particular, AI has the advantage of being easily parallelizable, which makes sum-threshold attacks and astroturfing / Sybil attacks on social trust networks feasible. Like, part of how we make decisions is through intuitions about what people around us think; we usually give at least some weight to channels that we didn’t verify carefully; when we don’t verify carefully, we’re open to bots; when bots are coordinated, they can pump a large total quantity of social gradient through the social networks of trust.
(This is pretty different from a sustained one-on-one conversation.)
You tangentially reminded me of this 2014 Scott tumblr post:
The level above Mohammed is the one we should be worried about
nostalgebraist:
Scott:
Yeah another thing I probably could’ve emphasized more in my comments/writing in general is that human variation is pretty wide, and the human peak is already pretty scary. Sometimes people I talk to are often imagining like salespeople or their charismatic friends in high school, and not like legendary diplomats, statesmen, religious leaders etc.
AI persuasion is also dangerous because an agent can easily find all information about you online, and tailor its approach to convincing you. With an evolved harness, an agent might also be very good at teasing out personality indicators from a conversation or pointed questions, and then using that to put you off-balance. It doesn’t need to convince you that it’s correct, it can just make you uncertain or distrust the sources you use for information.
This could look like bots in a forum that argue with you, in the background analyzing your account and responses to profile you, and then providing strong counterexamples to unrelated beliefs that put you off-balance.
I go back and forth on how big a deal this is. On the one hand I think it’s a real capability and near-term threat (if not today than in 3-6 months from now). Think swarms of agents trying to socially hack you.[1]
On the other I’m not sure how much persuasion capabilities in practice scale with inference-time thinking within the normal human range, it’s plausibly not actually that much.
(Jason’s tweet is about swarms at individually peak human levels but you can also imagine a story where they individually are lower but still work together productively).