AI persuasion is also dangerous because an agent can easily find all information about you online, and tailor its approach to convincing you. With an evolved harness, an agent might also be very good at teasing out personality indicators from a conversation or pointed questions, and then using that to put you off-balance. It doesn’t need to convince you that it’s correct, it can just make you uncertain or distrust the sources you use for information.
This could look like bots in a forum that argue with you, in the background analyzing your account and responses to profile you, and then providing strong counterexamples to unrelated beliefs that put you off-balance.
I go back and forth on how big a deal this is. On the one hand I think it’s a real capability and near-term threat (if not today than in 3-6 months from now). Think swarms of agents trying to socially hack you.[1]
On the other I’m not sure how much persuasion capabilities in practice scale with inference-time thinking within the normal human range, it’s plausibly not actually that much.
(Jason’s tweet is about swarms at individually peak human levels but you can also imagine a story where they individually are lower but still work together productively).
AI persuasion is also dangerous because an agent can easily find all information about you online, and tailor its approach to convincing you. With an evolved harness, an agent might also be very good at teasing out personality indicators from a conversation or pointed questions, and then using that to put you off-balance. It doesn’t need to convince you that it’s correct, it can just make you uncertain or distrust the sources you use for information.
This could look like bots in a forum that argue with you, in the background analyzing your account and responses to profile you, and then providing strong counterexamples to unrelated beliefs that put you off-balance.
I go back and forth on how big a deal this is. On the one hand I think it’s a real capability and near-term threat (if not today than in 3-6 months from now). Think swarms of agents trying to socially hack you.[1]
On the other I’m not sure how much persuasion capabilities in practice scale with inference-time thinking within the normal human range, it’s plausibly not actually that much.
(Jason’s tweet is about swarms at individually peak human levels but you can also imagine a story where they individually are lower but still work together productively).