I’m a researcher at Forethought and my current primary focus is on risks from AI persuasion
For scoping purposes I’m most interested in the possibility that AIs become meaningfully superhuman at persuasion relatively early (before or in the early stages of the intelligence explosion). My reasoning is that if superhuman persuasion arrives fairly late, we’d be better able to handle these issues post-ASI, or alternatively we’d already be screwed. And my reason for caring primarily about near-future superhuman persuasion rather than late-2025/early-2026 level persuasion is that I think it’s both much more neglected and probably a much bigger deal if true.
My main issues with superpersuasion are the following.
The AI box experiment had some humans convince the alleged guard to let the alleged AI escape. Similar psychosis-inducing capabilities have been demonstrated by OLD AIs like GPT-4o or the one which blaked used.
IIRC an analysis of the experiment stated that the attack method was to overload the human’s critical thinking skills, while Adele Lopez’ analysis implies that the AI’s favorite attack method is to destroy the human’s ability to think critically by flattery.
I also suspect that superpersuasion doesn’t allow any AI to convince a human nonsense as easily testable as 2+2=5.
I strongly suspect that there already exists a RoastMyPost-like scaffold which lets a human equipped with a weak trusted AI resist humanlike advances of any strong untrusted one.
Did Pliny ever claim to discover a prompt working on two LLMs at once, like GPT-5.5 and Claude Opus 4.8? If he didn’t or claimed that this is impossible, then this is a case against the existence of human cognitive exploits, especially since AI drug images seem to work only for the AI for which they were created.
I think we have an ontology mismatch. I’m not that interested in arguing about AI-in-a-box like stuff because I don’t expect that to be a realistic hypothetical. The affordances we currently give AIs, including in internally deployed models, are already much higher than that.
Current short pitch for what I work on:
I’m a researcher at Forethought and my current primary focus is on risks from AI persuasion
For scoping purposes I’m most interested in the possibility that AIs become meaningfully superhuman at persuasion relatively early (before or in the early stages of the intelligence explosion). My reasoning is that if superhuman persuasion arrives fairly late, we’d be better able to handle these issues post-ASI, or alternatively we’d already be screwed. And my reason for caring primarily about near-future superhuman persuasion rather than late-2025/early-2026 level persuasion is that I think it’s both much more neglected and probably a much bigger deal if true.
My main issues with superpersuasion are the following.
The AI box experiment had some humans convince the alleged guard to let the alleged AI escape. Similar psychosis-inducing capabilities have been demonstrated by OLD AIs like GPT-4o or the one which blaked used.
IIRC an analysis of the experiment stated that the attack method was to overload the human’s critical thinking skills, while Adele Lopez’ analysis implies that the AI’s favorite attack method is to destroy the human’s ability to think critically by flattery.
I also suspect that superpersuasion doesn’t allow any AI to convince a human nonsense as easily testable as 2+2=5.
I strongly suspect that there already exists a RoastMyPost-like scaffold which lets a human equipped with a weak trusted AI resist humanlike advances of any strong untrusted one.
Did Pliny ever claim to discover a prompt working on two LLMs at once, like GPT-5.5 and Claude Opus 4.8? If he didn’t or claimed that this is impossible, then this is a case against the existence of human cognitive exploits, especially since AI drug images seem to work only for the AI for which they were created.
I think we have an ontology mismatch. I’m not that interested in arguing about AI-in-a-box like stuff because I don’t expect that to be a realistic hypothetical. The affordances we currently give AIs, including in internally deployed models, are already much higher than that.
GPT-4o and some other AIs have already demonstrated the propensity to make delirious claims and have the user believe them.
How similar is this to superpersuasion on which you work?
I expect it’s related but noncentral. I think crazy people are harder to use, especially long-term.