Current AIs get basically no practice directly interacting with humans. They learn about humans implicitly by reading text about them, or getting RL-ed by human-derived preference models / environments. While this data could teach an AI to predict what humans might do, or what they believe/want, it doesn’t teach them how to causally intervene in a social context.
If you wanted to train an AI to get good at persuading people, or having good (active) social skills in general, you’d need to either:
find some kind of scalable proxy, e.g. practice persuading other AIs
hire loads of humans to practice on
I naively guess that getting good quality persuasion data would be extremely expensive, and that there won’t be especially good scalable proxies. This pushes me to think that AI systems, even when very advanced, will continue to have pretty poor social skills.
Alternatively:
Improved general intelligence leads to good social skills without much practice
Superpersuasion can be done without especially good social skills, e.g. using epistemic advantages.
We found converging evidence that AI’s advantage stemmed from rapidly deploying larger quantities of information: after coaching, expert humans could tie an AI constrained to respond at human speeds and with human-length messages.
Prior work has identified information provision as a highly effective persuasion strategy [6], and specifically that the number of fact-checkable claims deployed in a conversation predicts its persuasive impact [14].
It seems their persuasiveness is partially due to just being faster, but personally I wouldn’t discount the effectiveness of just general intelligence and RLHF.
I can see some ways that large labs could collect data on how effective their models are at persuasion, or something like it. As they have the corpus of back-and-forth interaction with chatbots by regular people today.
Take a chat of the form: “I think X (not Y) but I’m open to change my mind.” The chatbot responds with arguments for Y. SOME (but probably only a small portion) of these chats might culminate in some human text of the form “You convinced me.”
Many of the things that we humans seek assistance in understanding are the common controversial topics of the day. Politics, social issues, and so on. In a large enough dataset, there are a significant amount of conversations on any one subject. And so, different models can be compared on how successful they are at changing someone’s mind.
Success might be quantified as the amount of text, or number of back-and-forth interactions, required to get to the end point of “OK, I’m convinced”.
Is it ethical for labs to conduct this kind of research?
This is a two-sided weapon. If we can identify a parameter in models that is “Good at persuasion”, then it can be controlled for. It can also, of course, be reward-hacked to potentially dangerous levels.
Some quick thoughts on AI superpersuasion:
Current AIs get basically no practice directly interacting with humans. They learn about humans implicitly by reading text about them, or getting RL-ed by human-derived preference models / environments. While this data could teach an AI to predict what humans might do, or what they believe/want, it doesn’t teach them how to causally intervene in a social context.
If you wanted to train an AI to get good at persuading people, or having good (active) social skills in general, you’d need to either:
find some kind of scalable proxy, e.g. practice persuading other AIs
hire loads of humans to practice on
I naively guess that getting good quality persuasion data would be extremely expensive, and that there won’t be especially good scalable proxies. This pushes me to think that AI systems, even when very advanced, will continue to have pretty poor social skills.
Alternatively:
Improved general intelligence leads to good social skills without much practice
Superpersuasion can be done without especially good social skills, e.g. using epistemic advantages.
AI is already better than expert humans at persuasion! (source: https://arxiv.org/html/2606.16475v1)
It seems their persuasiveness is partially due to just being faster, but personally I wouldn’t discount the effectiveness of just general intelligence and RLHF.
I can see some ways that large labs could collect data on how effective their models are at persuasion, or something like it. As they have the corpus of back-and-forth interaction with chatbots by regular people today.
Take a chat of the form: “I think X (not Y) but I’m open to change my mind.” The chatbot responds with arguments for Y. SOME (but probably only a small portion) of these chats might culminate in some human text of the form “You convinced me.”
Many of the things that we humans seek assistance in understanding are the common controversial topics of the day. Politics, social issues, and so on. In a large enough dataset, there are a significant amount of conversations on any one subject. And so, different models can be compared on how successful they are at changing someone’s mind.
Success might be quantified as the amount of text, or number of back-and-forth interactions, required to get to the end point of “OK, I’m convinced”.
Is it ethical for labs to conduct this kind of research?
This is a two-sided weapon. If we can identify a parameter in models that is “Good at persuasion”, then it can be controlled for. It can also, of course, be reward-hacked to potentially dangerous levels.