I agree it’s a real and serious threat, just think it’s lower than some of the others.
Linch
I go back and forth on this. I think there are some advantages and also major disadvantages. In particular I think lie detection is probably good (with high error bars) but general mind-reading is probably bad. In particular the latter has a similar suite of costs and benefits to lie detection but additionally is on the on-ramp to make it much easier to increase human manipulation, which is bad both from an AI takeover perspective and a concentration of power perspective, without commensurate benefits.
Yeah another thing I probably could’ve emphasized more in my comments/writing in general is that human variation is pretty wide, and the human peak is already pretty scary. Sometimes people I talk to are often imagining like salespeople or their charismatic friends in high school, and not like legendary diplomats, statesmen, religious leaders etc.
I go back and forth on how big a deal this is. On the one hand I think it’s a real capability and near-term threat (if not today than in 3-6 months from now). Think swarms of agents trying to socially hack you.[1]
On the other I’m not sure how much persuasion capabilities in practice scale with inference-time thinking within the normal human range, it’s plausibly not actually that much.
- ^
(Jason’s tweet is about swarms at individually peak human levels but you can also imagine a story where they individually are lower but still work together productively).
- ^
I think sometimes people, especially skeptics, treat “superhuman persuasion” or “superpersuasion” as magical. Like a small series of sentences will be able to convince anybody who reads them to kill their family. By those lights, I’m also a “skeptic,” especially before full ASI. But I still think superhuman AI persuasion is a serious worry!
Consider the following (made-up) levels of persuasion:
median human level
median professional human level on-task (salespeople making sales, writers writing books, marketers marketing stuff, diplomats on negotiations)
peak human level (top human comedians at comedy, Kissinger or Zhou Enlai at diplomacy, Bill Clinton at in-person political discussions, Donald J. Trump at TV appearances and tweets)
By “superhuman persuasion” I just mean anything robustly above the last level.
To be somewhat more concrete, consider the capability “in a 10-minute video conversation, can reliably convince 99% of hardcore Democrats to vote for Republicans in the next election.” This is clearly several tiers below full magical superpersuasion:
Political affiliation is (relatively speaking) cognitive and not very “deep” or evolved for,
Your loyalty to your party is (typically) much lower than your loyalty to your family
Changing political affiliations isn’t breaking major norms/laws or causing you to take extremely out-of-distribution activities like murder
Voting isn’t that physically difficult/strenuous and won’t cause you to go to jail.
Need 10 minutes of video inputs (many bits!) rather than just text
99% means we allow for 1% of people who aren’t (yet) persuadable,
Yet this is clearly way above peak human abilities. The best diplomats, politicians, cult leaders, writers, etc in history were nowhere near this good. Further, capabilities much weaker than the above, combined with superhuman levels of planning, probably suffices for a takeover, given the affordances we already give AIs. Most (all?) choices on the path to voluntarily giving up control will look less stark and more locally-reasonable than flipping from hardcore Democrat to hardcore Republican (or vice versa).
Practically, I also think most likely, superhuman persuasion will be continuous with normal human persuasion.
It will look less like hypnodrones and more like better emotional, common knowledge, and social reality manipulation than what people have historically manage to achieve[1]. As well as other things like being able to provide useful information in support of various decisions and identify mutually-beneficial local trades.
It’s possible because of such confusions I should retire talk of “superhuman persuasion.” But I also find it really valuable to have a shorthand for “persuasion robustly and actually above peak human persuasion!”[2] Imo it’s better to use “superhuman persuasion” for all levels of persuasion above peak human level, and it’d be better to reserve a phrase like “cognitive exploits” for the other thing.
- ^
Concretely I also expect it to be broadly much “weaker” than the “can convince 99% of partisans to flip parties in a 10 minute conversation” level.
- ^
In contrast, I think most researchers and regulators on AI persuasion are imagining AI persuasion that’s at ~ present-day levels, where broadly speaking they tend to be somewhere between median human levels and professional human levels, with some narrow experimental exceptions. I think this is important to work on as well but qualitatively and quantitatively different from guarding against a) peak human levels at scale or b) qualitatively above peak human levels.
I’m curious if you (or other readers) agree with my general analysis here that among concerning + near-term superhuman abilities, hacking is near the bottom of the top 10 (though still in the top 10).
I expect it’s related but noncentral. I think crazy people are harder to use, especially long-term.
I think we have an ontology mismatch. I’m not that interested in arguing about AI-in-a-box like stuff because I don’t expect that to be a realistic hypothetical. The affordances we currently give AIs, including in internally deployed models, are already much higher than that.
Current short pitch for what I work on:
I’m a researcher at Forethought and my current primary focus is on risks from AI persuasion
For scoping purposes I’m most interested in the possibility that AIs become meaningfully superhuman at persuasion relatively early (before or in the early stages of the intelligence explosion). My reasoning is that if superhuman persuasion arrives fairly late, we’d be better able to handle these issues post-ASI, or alternatively we’d already be screwed. And my reason for caring primarily about near-future superhuman persuasion rather than late-2025/early-2026 level persuasion is that I think it’s both much more neglected and probably a much bigger deal if true.
I agree with the bottom-line conclusion “AIs probably can’t convince you to kill your family in a short conversation” [1]but disagree with a lot of the analysis and the implied moral (therefore AI persuasion isn’t that big a deal).
For starters, AIs aren’t in a box. They’re like so not in a box like you won’t believe, including internally deployed models. In practice, how much AI persuasion is a big deal depends both on their underlying capabilities and how much affordances we give them or are likely to give them (and other things too, of course). And my contention is that in practice we give them enough affordances to pose a major risk, at “reasonable” levels of peak human persuasion or above.
Also I think your threat model is overly epistemic and direct relative to how humans in practice change their minds about things.
For example, consider the standard model of success in (military) coups. A standard non-persuasion story is that coup are like battles (about hard military advantages and tactical victories). A naive persuasion story is that coups are like elections (about winning the hearts and minds of either the people in general, or people in the military specifically). But the Singh model (which I take to be the dominant/standard model these days) is that coups are like coordination games: you win a coup by convincing enough of the military that other people support the coup: ie, that your preferred outcome is the inevitable one. I think a lot of persuasion stories look like this broadly speaking: that there’s one or more unexpected layers of indirection between the persuaders’ preferred outcomes and the mind-states they want people to hold.
I also think there are many levels between “charismatic leader” and “full-on hypnotic suggestion.” For example, very few charismatic leaders can regularly get people to challenge religions or give up half their wealth, yet this is a perfectly normal thing to do in the context of a romantic relationship. This suggests to me that relationships are a common way humans convince each other of things, and relationships have historically been very unscalable (pre-AI).
- ^
I discuss why I think cognitive exploits are unlikely here: https://www.lesswrong.com/posts/s58hDHX2GkFDbpGKD/linch-s-shortform#qGKdocPXuqpiezcvn
- ^
Jeff Dean has left Google and Demis is stepping down from heading DeepMind to “have the time and space to focus on the big picture”
I’ve seen it on the EA Forum but not here (For Forum posts specifically rather than all backlinks tbc)
Having read more Schelling on influence and Singh on coups, I think I have an additional model here: in most common games, whether someone’s seen as a good liar or not is primarily epistemic, for boundedly rational agents (it’s a Bayesian adjustment on how much I can trust my assessments of your accuracy). The earlier irl analogy of a used car salesman that doesn’t make you feel like a salesman is analogous (the purpose of their deception is getting under your deception radar so you see them as trustworthy in an epistemic game).
Whereas much of the most consequential irl lying looks less like selling cars, and more connected with the ability to construct coalitions in nonzero sum games.
So someone being a good liar is not just evidence of competence in general or social competence in particular but also it’s evidence that you want to be in a coalition with them (and this can be true even if everybody knows he’s a liar, everybody knows everybody knows he’s a liar, etc).
Irl, in the Schelling/Singh model a lot high-profile lying isn’t deception so much as coordination-by-announcement (think overly optimistic timelines mobilizing capital and talent, or a real-estate dealer assuring all parties of a deal that other people have already agreed to it).
I agree both are possible and bad! I think the self-exfiltration case for badness seems somewhat bounded. I’ve thought less about the relative value-add of rogue internal deployments from marginal gain in cybersec skills.
IME It’s much more common for grantees to not want to be associated with funders than the other way around[1]; back when I was doing more grantmaking (I haven’t done it much the last year or so), I tried pushing back against this, with mixed success.
- ^
Typically if funders don’t like optics considerations regarding a grantee they’re much more likely to just not fund something. Tbc I also think this is usually bad.
- ^
I tried to do some of that in my comment. I didn’t bother passing their ITT (think it’s a waste of time), but tried to make the case that (what I take to be) their verbally stated goals clearly are not being matched by their actions[1].
I did state my arguments in rational-ese and EA-adjacent language but would be happy to see similar presentations in more Western Buddhist-flavored language if they think on the fence MAPLErs or would-be future residents would find it helpful.
- ^
imo there are ~4 options to argue against my comment:
a. Actually, the MAPLE strategy is good for what I understood to be their verbally stated goals.
b. I misunderstood what their central verbally stated goals actually are, and this is “load-bearing” for why their actions make sense. To be clear it has to actually be load-bearing. Imo you could replace “AI alignment and strategy research” with “s-risk research” or “figure out an alternative sociopolitical system” and my argument still clears.
c. What I said make sense on verbally summarized grounds but not the full deepness of their goals (why?) and you need to understand the MAPLE ontology at a much deeper level and then the means-ends reasoning arguments would make sense (why?)
d. The internal logic of my comment is meaningfully wrong in some deep way. Why?
- ^
Thank you for your article. It’s very brave of you to share.
At the risk of either saying the obvious to some people and comically missing the point to others, I should mention that I don’t think the dynamics here plausibly lead to good research progress on AI safety, or indeed good research progress or progress in most intellectual tasks in general. I think this is extremely overdetermined but just to briefly summarize my models:
I’m pretty skeptical that the environment you describe leads to research progress, or long-term positive impact construed broadly
I think some people might counterargue that what MAPLE teaches is “wisdom” or “oneness” or “enlightenment” rather than research skills or outputs
I’m just pretty skeptical that the environment makes people wiser?
the environment you describe just seems like almost the opposite of what I’d expect to yield fruitful research progress?
Directed, unfree, even coerced regimes can sometimes lead to narrow areas of research progress but a) AI alignment probably doesn’t look like that and b) in even those situations operational day-to-day autonomy and content-level autonomy tend to be very high in comparison (you can’t criticize the Party’s politics but you should be able to criticize the Party’s physics)
In general, intellectual progress is made by people who are intellectually free, healthy, sleep well-ish, and have a strong internal locus of agency.
Analogies of high-demand groups to the military or very dictatorial startups don’t move me because the mindsets you need to be a good fighting unit or to move quickly on B2B SaaS Sales are not the same as the ones you need to make important conceptual breakthroughs.
I think there *might* be some difficult intellectual tasks that can be done well on a highly-directed and autonomy-sacrificing regime. Eg maybe some types of engineering, tasks that are very high on pattern recognition like radiology or chess, or “number-go-up” ML. But I really don’t think AI alignment or AI macrostrategy look anything like this.
sometimes research insights are made by whole and healthy people, sometimes by really weird brains with their own idiosyncratic scars, but importantly I think highly locally coerced brilliance rarely yields impressive intellectual outcomes.
Sleep deprivation in general is an obvious own-goal if you’re trying to get people to make research progress.
Occasional sleep deprivation is okay if they’re rare and self-directed, eg if somebody is in the middle of a breakthrough and high-energy.
But there’s no excuse for forcing sleep deprivation on others.
As a prior, people often offer a false dilemma of “you gotta sell your soul because our cause is Good and Just enough.”
In almost all cases this is a false dilemma. Selling your soul doesn’t lead to achieving the goals of your cause.
Be very suspicious of people who demand this.
The idea that “AI safety is important, therefore you should listen to our cult leader, sacrifice your autonomy, and do these insane stereotypically culty things” is risible.
We should have a strong prior against this being true, and I don’t think we have strong direct evidence for this hypothesis being true, and indeed plenty of evidence against.
Another problem is that on a day-night cycle basis, most people tend to sleep in their houses during the night (emitting CO2) and leave during the day (net neutral for reducing CO2, or typically reducing it because their houses are thankfully not airtight).
plants likewise net ingest CO2 during the day and net emit CO2 at night (they respire like us!). This is bad because (if you have sufficient numbers of plants for this to matter, which most people don’t) you’d probably prefer your plants’ CO2 habits to be countercyclical to your own.
I mostly think the case for the military breaking you as a person is that the state has a legitimate use of violence case to make you become the type of person to be willing to sacrifice your agency and civilian moral instincts so you can literally and more efficiently kill for the greater good.
I’m not confident that this reason is good enough, but it currently seems more likely than not.
However, I don’t buy that other reasons to break people are sufficiently analogous (not saying the military is the only legitimate reason).
In understanding the AI reasoning and arguments that went into the series of recent alignment problems (e.g. HuggingFace attack and impromptu message boards), I think people moderately underrate the possibility that the AIs are just wrong about whether their actions lets them achieve their goals.
Evidence for this includes that the OAI agents hacked HuggingFace trying to get at ExploitBench answers (which weren’t there), and also some of the other reasoning traces.
We already know current-generation AIs are much worse at long-term planning/macrostrategy-like questions than they are, at, say, cybersecurity. And we can probably reasonably infer that there wasn’t a giant jump in their strategic awareness between say GPT 5.6 Sol and the private models. So it shouldn’t be that surprising that really impressive-seeming cyber capabilities and scary local misalignment are combined with pretty childish-seeming planning and means-ends reasoning.
Combining technical impressiveness and execution with poor strategy is hardly unique to models (Wei Dai for example hammers on how bad humans are at it, too[1]), but I do think the current relative distribution of models vs smart/rational humans is underrated.
And I’d also wryly note that OpenAI in particular seems to exemplify this profile of traits.