A lot of rationality-philosophy talks about types of arguments in terms of asymmetry: a good argument is one that’s much better at convincing you of something that’s true, than it is at convincing you of something that’s false. Imo discussion of “AI superpersuasion” would benefit from incorporating a bit more of this perspective; I want AIs to be able to convince me to do anything if that thing is in my best interest (and not be able to convince me if it isn’t), and defense feels mostly like a matter of enabling people to figure out which is which.
A lot of rationality-philosophy talks about types of arguments in terms of asymmetry: a good argument is one that’s much better at convincing you of something that’s true, than it is at convincing you of something that’s false. Imo discussion of “AI superpersuasion” would benefit from incorporating a bit more of this perspective; I want AIs to be able to convince me to do anything if that thing is in my best interest (and not be able to convince me if it isn’t), and defense feels mostly like a matter of enabling people to figure out which is which.
Yeah definitely tracking this! At some other points I use the phrase “illegitimately persuasive” as shorthand for what I don’t want the AIs to do.