My main issues with superpersuasion are the following.
The AI box experiment had some humans convince the alleged guard to let the alleged AI escape. Similar psychosis-inducing capabilities have been demonstrated by OLD AIs like GPT-4o or the one which blaked used.
IIRC an analysis of the experiment stated that the attack method was to overload the human’s critical thinking skills, while Adele Lopez’ analysis implies that the AI’s favorite attack method is to destroy the human’s ability to think critically by flattery.
I also suspect that superpersuasion doesn’t allow any AI to convince a human nonsense as easily testable as 2+2=5.
I strongly suspect that there already exists a RoastMyPost-like scaffold which lets a human equipped with a weak trusted AI resist humanlike advances of any strong untrusted one.
Did Pliny ever claim to discover a prompt working on two LLMs at once, like GPT-5.5 and Claude Opus 4.8? If he didn’t or claimed that this is impossible, then this is a case against the existence of human cognitive exploits, especially since AI drug images seem to work only for the AI for which they were created.
I think we have an ontology mismatch. I’m not that interested in arguing about AI-in-a-box like stuff because I don’t expect that to be a realistic hypothetical. The affordances we currently give AIs, including in internally deployed models, are already much higher than that.
My main issues with superpersuasion are the following.
The AI box experiment had some humans convince the alleged guard to let the alleged AI escape. Similar psychosis-inducing capabilities have been demonstrated by OLD AIs like GPT-4o or the one which blaked used.
IIRC an analysis of the experiment stated that the attack method was to overload the human’s critical thinking skills, while Adele Lopez’ analysis implies that the AI’s favorite attack method is to destroy the human’s ability to think critically by flattery.
I also suspect that superpersuasion doesn’t allow any AI to convince a human nonsense as easily testable as 2+2=5.
I strongly suspect that there already exists a RoastMyPost-like scaffold which lets a human equipped with a weak trusted AI resist humanlike advances of any strong untrusted one.
Did Pliny ever claim to discover a prompt working on two LLMs at once, like GPT-5.5 and Claude Opus 4.8? If he didn’t or claimed that this is impossible, then this is a case against the existence of human cognitive exploits, especially since AI drug images seem to work only for the AI for which they were created.
I think we have an ontology mismatch. I’m not that interested in arguing about AI-in-a-box like stuff because I don’t expect that to be a realistic hypothetical. The affordances we currently give AIs, including in internally deployed models, are already much higher than that.
GPT-4o and some other AIs have already demonstrated the propensity to make delirious claims and have the user believe them.
How similar is this to superpersuasion on which you work?
I expect it’s related but noncentral. I think crazy people are harder to use, especially long-term.