Some quick thoughts on AI superpersuasion:
Current AIs get basically no practice directly interacting with humans. They learn about humans implicitly by reading text about them, or getting RL-ed by human-derived preference models / environments. While this data could teach an AI to predict what humans might do, or what they believe/want, it doesn’t teach them how to causally intervene in a social context.
If you wanted to train an AI to get good at persuading people, or having good (active) social skills in general, you’d need to either:
find some kind of scalable proxy, e.g. practice persuading other AIs
hire loads of humans to practice on
I naively guess that getting good quality persuasion data would be extremely expensive, and that there won’t be especially good scalable proxies. This pushes me to think that AI systems, even when very advanced, will continue to have pretty poor social skills.
Alternatively:
Improved general intelligence leads to good social skills without much practice
Superpersuasion can be done without especially good social skills, e.g. using epistemic advantages.
I find the title of this post a bit misleading, given that this is a primarily job ad and not a discussion of how mid-training interacts with RL. I’d recommend changing it.