Epistemics and Coordination: It’s complicated!
Money is pouring in, people are looking for new areas to fund, and the invisible hand is starting to grab a bit at AI for epistemics and coordination (AIFEC hereafter). It also got invoked in AI2040 as a potential part of the winning strategy.
My feelings here are mixed — I think the best version of AIFEC is great, but also the existing public writeups are only a few cycles deep on tracing out the different ways that the obvious plan backfires. And regrettably I think some people have correctly written off AIFEC because what they have read appears a bit naive to them. In my heart I always planned to do a proper writeup of my thoughts when things were a bit less busy, but, well, now we’re in the 100x funding era. So here’s my scrappy, hopefully-better-than-nothing attempt.
The six big claims:
AIFEC could be great! It’s a tractable way to do good on the margin, and the best version is a legitimate theory of victory
It could also easily backfire, especially for implementations that depend on scaling with inference
“Better epistemics” is often more hostile than it seems, and people have good reasons to be wary of things which profess to help them understand what’s going on
Vagueness about what AIFEC is doing lets you ignore tradeoffs: Yes you can just go make a bunch of AIFEC startups, and yes AIFEC could help with international coordination, but those scrappy startups aren’t what solve US-China tensions.
Sometimes people just don’t get along
But still, AIFEC could be great!
The basic case is pretty good!
Here is my favourite case for AIFEC:
If humanity ends up blowing itself up, it’s going to be a combination of three factors. Firstly, maybe we just rationally incurred some risk, the same way that getting on an airplane might kill you. Secondly, maybe we underestimated the scale of the risk. Thirdly, maybe some people took risks that personally benefit them at the expense of others, generating negative externalities.
I find this story pretty compelling, and surprisingly different in form to the standard arguments about AI x-risk. It’s actually much more general. And when you squint at the three factors, well, the first isn’t even really a problem, the second is basically poor epistemics, and the third is basically poor coordination. So maybe if we just get enough epistemics and coordination, we’re in the clear?
And it sure seems like AI is about to make this way easier — so much more cognition on tap, plus all kinds of neat new form factors like arbitrating minds that genuinely vanish after the fact.
Add to that: it’s not really all-or-nothing in the same way that alignment is. Decent AIFEC lift should help with all our other problems. In fact, AIFEC should help a fair bit with AIFEC — the better our epistemics and coordination, the more effectively we can coordinate to appropriately invest in further work.
One can imagine a kind of dizzying spiral to heaven, where we pass the Coasean friction event horizon, slay Moloch, and fuse into a liberal-libertarian omnimind in which every individual has their own authentic preferences while the group deftly dances along the pareto frontier. X-risk drops to the socially optimal level, the socialist calculation problem is finally solved, and we can all finally just get along.
Hyperbole aside, it’s worth emphasising that something in the realm of AIFEC is a legitimate full-blown wincon up there with aligned AGI. Indeed, a lot of people’s tacit plan seems to be “align the AI and then let it solve democracy and all the rest”, but one could just as well go in the other direction: “solve politics and ignorance, and then just be reasonable about advanced AI”
Alas, it is not so simple.
Problem 1: Naive AIFEC could easily backfire
The biggest splash of cold water is that AIFEC won’t be free. Especially in this era of inference scaling and an increasingly closed frontier, the fruits of AIFEC will not be equally distributed. There are certainly some bits of technological progress that are remarkably egalitarian — even billionaires use facebook, gmail, and iphones. But if the selling point is that you get to throw lots of intelligence at solving problems, then you’re going to differentially favour the people who can actually call up that intelligence at scale.
Indeed, there are ways this could backfire. It’s certainly possible that AI will enable everybody to seamlessly coordinate, but first it will enable small groups to, ahem, collude. If you’re worried about things like coups and permanent underclasses, it’s not clear that more coordination makes your life any easier. Similarly, if you build a machine that turns compute into better decisions and more situational awareness, those benefits will mainly accrue to the people who actually have the compute.
Problem 2: “Better AIFEC” can be a somewhat hostile move
It’s nice to think that everyone who believes false things is basically just misguided, and that the only reason they don’t take up arms for the truth is that they haven’t seen it yet. I think the real picture is unfortunately a bit more complicated.
The problem is, historically there have been many groups that have weaponised very convincing arguments to get their way. One reasonable response to a seemingly faultless argument for a seemingly crazy conclusion is to throw up your hands and assume you’re being swindled — in other words, epistemic learned helplessness. Even if team AIFEC is actually the good guys, people will be correct to be suspicious. Governments in particular depend a lot on restricting what kind of information is admissible, as does the judicial system.
On top of that, I do think that before you start swinging the club of truth it’s worth taking a beat to ask how pure your motivations really are. The EA/rationality community has a bit of a history of swinging the club of truth in a more hostile way — “save the drowning child” etc. Moreover, in the realm of politics, often the real sleight of hand is just changing what things are salient — elections are won over people’s sense of what the election is really about. Framings and deliberative processes are rarely as neutral as they seem. So even if your tools don’t privilege a specific answer directly, the choice of deliberative structure is, well, a choice.
The same is true of coordination. To give an obvious example, democracies are meant to represent the will of the people, but the structure of representation pretty directly modulates that will. The choice between proportional representation and first past the post, bicameral houses, election cycles, district boundaries and so on are superficially choices about how to structure the coordination, but they also sometimes obviously favour certain conclusions. It’s really hard to build neutral coordination structures, and people are right to be sceptical of anyone who pretends otherwise.
One slightly thornier point: believing false things is actually a pretty powerful coordination mechanism. Many groups cohere around a mixture of surprising truths and blatant falsehoods. When you try to bring the truth to such people, they will actually fight back. As a modest example, deconverting a child from the religion held by their entire family is actually pretty unpleasant and arguably not very nice. This is true to a lesser extent for e.g. adults and the political tribe of their social milieu. Now, maybe it’s worth it if you’re actually right, but you should expect them to fight back!
So: plenty of good to be done, but plenty of pitfalls along the way.
Problem 3: Vagueness lets you ignore tradeoffs
I think one of the major appeals of AIFEC is that you can just let a thousand flowers bloom by sending a lot of bright young things off to found startups and seeing which ones succeed. This is certainly scalable and likely to produce some hits. Unfortunately I’m not sure it’s enough, and the other parts are a lot harder.
What coordination and epistemics do we actually need? One way I like to think about this is in terms of what the channel is that leads from the so-called better angels of our nature to actual global action. Probably most of that comes down to robust democratic oversight and international coordination. These are very difficult! And I do not expect that even the top percentile AIFEC startup moves the needle that much, because these processes are by design pretty resistant to hopping on the latest technological fad.
Similarly, plenty of AIFEC wins don’t seem to me to really be on the critical path. Deliberative processes, for example, seem like somewhere that AI could be enormously helpful, in a way that might help overcome a major traditional limitation of democracy, but I don’t see that being super crucial for dealing with worlds where people lose all their leverage. It seems helpful, certainly, but not obviously necessary and definitely not sufficient.
Getting lots of products off the ground seems great, both in case some are great and to build institutional capacity and expertise, but we can’t neglect the other steps, and it’s important to think at least a bit about what specifically moves the needle.
This problem isn’t at all unique to AIFEC — the general AI risk movement has a bit of a problem with coming up with some new cause celebre and then unleashing a torrent of work that nominally fits the category without being useful. And I don’t want to overstate it: the mass of AIFEC startups probably will produce some hits, along with some important general lessons and greater capacity. But the best version of AIFEC does need to grapple with it.
Problem 4: Sometimes people just don’t get along
At the risk of psychologising, I think part of the appeal of AIFEC is that it lines up with a kind of technocratic mistake theory instinct that basically the current race towards a cliff-edge is one huge misunderstanding, and if everyone could be provided with the right information and the right structure, things would all work out.
I think this is pretty true! And I yearn for it to be more true. But it is not entirely true. Sometimes people just don’t get along.
Some people would rather take AI soon so as to increase the odds of their own immortality or the immortality of their family, even at the risk of destroying all of earth. Some people really do want to tile the universe with hedonium. Some people’s order of preferences is basically “my country wins” > “annihilation” > “enemy country wins”. Some people are currently reaping the benefits of not internalising the risk externalities they create. Some people are causal decision theorists. Global coordination needs to either include these people or, well, conspicuously not include them.
And they are not dumb![1] They are not merely passive processes that will fail to notice your attempts to re-engineer the environment around them. Sometimes the ruling party decides to block political change even if it’s obviously reasonable behind some veil of ignorance, because they correctly notice that in the moment it is to their detriment.
Some of this stuff you can bargain about, but, well, the structure of the bargaining is not neutral. And you can only bargain so much with someone whose parents are slowly getting older and sicker.
I really don’t know what to do about that. It makes me pretty sad. And we are going to run into it more and more, so the sooner we can deal with it, the better. But damn, it’s tricky.
But overall, I am still pro
I reel off these problems not because I don’t believe in AIFEC, but because I do believe in what the better version of AIFEC could be. I’m baffled by how overinvested we are in strategies that, from my perspective, seem specifically geared towards helping frontier labs navigate the acute risk period. I think in an adequate world we’d be pushing simultaneously on every plausible path to victory, at least until we hit the thresholds where they started to trade hard against each other.
One obvious reason to be long on AIFEC is that it automatically scales with AI capabilities. Some version of it is clearly coming, and will clearly be helpful. And right now in particular there’s some opportunity to steer things, to get certain balls rolling sooner. Even if you’re extremely bitter lesson-pilled, a three month lead could be worth a lot.
There are a few other big topics I didn’t cover here that do feel relevant to me. In no particular order:
More so than I’d like, AIFEC tends to privilege individual-level agency, whereas I am pretty bought in on group agency being real and very important for risks.
AIFEC doesn’t grapple as much as I’d like with power and leverage, and indeed the frame seems to slightly bounce off them. This is related to problem 4 — AIFEC hasn’t grappled as much as I’d like with conflict theory.
I feel we lack institutional knowledge about how you serve both God and money.
Finally, I continue to feel that one of the most underrated pieces of work I contributed to was The Choice Transition, which is basically an attempt to articulate how AIFEC might get humanity into a stable basin from which we can reliably avoid bad outcomes and slowly build up to the good ones under our own volition. Compared to all the other paths to victory, AIFEC seems like the only one with this property — that humanity can correctly recognise itself as having the power to work towards the best outcomes — and my liberal instincts feel that this is worth holding onto.
(crossposted from my new-ish blog)
- ^
Except the causal decision theorists
This I label ‘immortality or die trying’ vibe and yes it’s not unheard of.
This is the imposition of unacceptable externalities you mentioned earlier! i.e. something that (at least in vague optimistic idealism land) AIFEC counters quite specifically. Those people, where they’re gambling with everyone else’s lives and the lives of their children, basically have to lose. That’s the game.
I do expect that it’s just always going to be very difficult to include fundamentalists and fanatics of various types in most kinds of wider coalition, and that won’t easily change even with the greatest of (AIFEC or otherwise) tech powers. At least, not without changing their (stated and often deeply held) preferences fairly substantially, which is both difficult and arguably immoral.
Even in those cases, my hope is that most partisans are, at least largely, partisan out of concern, fear, and loyalty, which under sufficiently healthy conditions (difficult!) aren’t incompatible with some bargains, especially of the gentle-liberal-order kind (live and let live sort of thing).
Some specific matters may be difficult to reconcile (who gets to control Jerusalem? should we launch the probably-kill-everybody-forever-but-maybe-immortalise-some-people machine today or next year?) but even there, compromises are plentiful in history and today.
Point of tentative disagreement. (I can see you’ve marginally toned this down since private copy, and I marginally agree more now.)
Sufficient? No way. But protective and helpful and multiplying with other efforts, not necessarily forever-sustainably but with real impact? Seems quite promising to me! (cf your ‘dizzying spiral to heaven’...)
Deliberative processes (if we mean the same thing) can ideally do all of
making people justifiably feel engaged (reducing fuel for populist and fantasist reaction)
revealing (creating common knowledge of) unacceptable risks or attractive opportunities for collective action
serve the information-passing and -aggregating function(s) of society
As for how this interacts with ‘real leverage’, well… one way I imagine this is that both the residual ‘real hard leverage’ and the ‘thick veneer’ (Leicht) combine with deliberative processes (and surfacing of societal priorities) by delaying (again, not necessarily forever-sustainably!) the further erosion of leverage.
My vision of broad-based collective epistemic uplift mainly goes via improving salience of reasonable nuance and important context, about e.g. what lines of argument and evidence support a given case, who has which interests and has done what shady stuff, and what alternative positions there are. On the whole I don’t want (or need?) there to be too many crazy-seeming conclusions being argued by long-winded chains of reasoning.
For more targeted institutional foresight, say (e.g. tech-tree and AI-progress futures work), sometimes you do just need a bunch of detailed evidence. I think that’s ultimately only going to work on people who have the time and willingness to spend that effort, but we can make tools and systems which help them leverage that effort, and which serve up the most important info.
(I previously commented this privately and your response was interesting, I didn’t get round to/know how to respond, but I’m interested in the thread and in others’ thoughts.)
point taken! I think I might be more sceptical than you about context provision, for the same general reason that argument mapping is hard.
I wonder about crazy-seeming conclusions. I think we maybe will need them? maybe a useful prompt is ‘how could AIFEC have helped in the pandemic?’
(aforementioned reply from aforementioned private discussion)
Do you have some examples of things that you would be against? What do you think people should not build? What are you the most suspicious about?
Honestly right now we’re far enough from the frontier that I think it makes sense for people to just do whatever and see what works, in the same way that I think most junior safety researchers shouldn’t worry too much about infohazards or capability externalities. For example, field-wide, I want to push back against AIFEC being just the for-profit startup-y area, or being purely about strategies that scale with inference compute, but it seems premature to tell individuals to avoid whole areas.
universal basic cryonics[1] now!
+ sufficiently good thinking/epistemics to see the imo-true claim that the technical problem of reviving one will likely be solved
(Here’s smth I wrote a month ago.)
Plan A might be a kinda rough situation, given you’d have millions of people lobbying for the project to go faster. Like, 30% of the pop (60-89) is taking an additional 30% chance of death if the singularity happens after 20 years compared with 10 years.
One solution is to cryo everyone 60+. I’m not entirely dissuaded by “cryo is too sci fi for them” because the original worry is that the cohort is lobbying for an early singularity! Maybe I should invest in some cryo startups. I bet even Leopold hasn’t backchained this far.
A similar solution is a “soft cryo”, i.e. medically induced comas. It seems pretty plausible that during an intelligence explosion, doctors will say “Keep that patient on life support for another year, the medicine will be better.”
Another solution is simplying paying risk premiums to the 60-89 age range, to compensate them / stop them lobbying for early foom.
(some messages i sent to a friend at coefficient giving 6 months ago, with a few edits)
maybe openphil global health and wellbeing should fund some cryonics evangelists and scalers
assuming developing advanced AI goes well for humanity (which openphil effectively seems to have a decently high probability on) or we ban AI and also don’t destroy humanity/capitalism/institutions, you are likely buying > billions of QALYs for each actual current human who would have died that gets cryopreserved. (edit: hmm i guess this is assuming recovery from the outputs of a human is fucked. if you can recover someone from their writing as well then maybe there’s not much to gain from being cryofrozen. not sure what i think about this.) i think currently the usual price per person is like 100k usd but i think it could be like 10k with worse methods that are probably just fine, and at massive scale it could be like 1k — if very widespread, cryonics could plausibly even be cheaper than other usual things people do with a body
i mean it’s unclear if these cost numbers are really directly relevant in the naive way. like, you wouldn’t compare these to bednet usd/QALY numbers, because openphil wouldn’t be paying for these individuals getting cryonics probably. tho idk, i guess openphil could be initially paying for some fixed costs of setting things up which could be well estimated per person by these numbers maybe. well, openphil could just also be paying for individual people lol. anyway, my guess is that the correct calculus will make the bet look much better than just paying for each cryopreservation
unironically yes, I think some people have tried to differentially accelerate AI for healthcare with exactly this in mind
To a large extent, a big part of the reason this isn’t tried is that it’s far, far more difficult, fora variety of reasons to make the general public be reasonable and that leading to us take the optimal level of risk, at minimum, compared to aligning AGI, and I’d still say this is correct (politics is more tractable than people on LW thought, but this is mostly downstream of warning shots like Mythos that woke up the executive branch, and to a lesser extent congress without needing to wake up the public.)
I agree that EA/rationality (especially EA) has a bad habit of relying on claiming morally-laden facts as true, and more generally one of the single most important constraints for anyone working in the field of AIFEC is to avoid marking moral/value claims as true/correct.
And I also agree that rationalists tend to be blunter and less persuasive than they should, even when it is true.
The other issue is that true claims, especially in AI will tend to be very, very complicated relative to simple but wrong claims, and one of the larger updates I’ve made about the AI revolution is that models have to be very, very complicated, by design, and this means you will always be plagued with bias and very difficult to falsify models.
So yeah, epistemic learned helplessness is common.
Maybe worth saying here is I mostly don’t buy this theory of change, precisely because I mostly do think robust democratic oversight and international coordination is far too intractable, and I think it’s more useful to focus on smaller sets of actors like the executive branch or individual members of congress.
I still don’t think the products are good, but the standard you are asking for is both extremely unrealistic and unnecessary.
Some other useful examples here are people who genuinely believe war is a positive good, because that allows them to exterminate whichever disfavored ethnic group, and you don’t care as much about the resources destroyed as you care about extermination.
More prosaically, people can genuinely disagree with liberalism/democracy being good, or believe that centralization into a dictator, junta or oligarchy is actually just good (populist movements are generally this).
More generally, people faking/genuinely believing that their ethnic group is being exterminated, so we need to go to war against the outgroup, no matter the cost is unfortunately pretty common, and critically the belief does not need to be true in order for it to cause wars.
Agree that frontier lab epistemics are likely over-represented in AIFEC, but I’d have a different distribution of work that is valuable to work on than you do.
It’s worth saying though that this will require people to trust the AIs pretty largely, and frontier labs have obvious and large incentives to do this, but most others don’t and indeed have anti-incentives to trust AIs, and even if the AIs are in fact jusitified in being trustworthy, it’s plausible that people won’t use AI because they incorrectly trust the AIs.
To be fair, that’s because group-level agency without incentives is rare, and most claims of groups acting agentically are describable in terms of their incentives.
Put another way, AIFEC already grants that group-level agency will exist assuming the incentives are right, and it’s correct to limit it to that.
I can plausibly agree with this, though I also think power/leverage is probably overrated as a problem, and this is mostly due to my view that a lot of the more negative results of absolute power/leverage only really hold when you have don’t have a small bit of caring (assuming AIs make the cost of loyalty go to essentially 0 through alignment), and while this does exist (given psychopaths/sociopaths), this is actually pretty rare, and more generally I think the amount of people that have non-zero altruism is large enough such that we can actually have leaders equivalent to the wise/benevolent ones we often see in fiction (and because AI enables very long lifespans, this means that leadership turnover stops, meaning a potential degradation of leadership skill is avoided).
(The main mechanism here is basically the tails coming apart, but in a positive direction).
And finally, my main conclusion is that if you want to work on AIFEC, you should probably focus on smaller groups as your end goal than trying to get large groups of the general public, and I give some reasons for hope on why epistemic efforts focused on small groups could make outcomes much better, even if they don’t involve democracy or liberalism.
Thanks for the detailed comments!
Honestly, I’m not sure. Like, yeah, getting everyone to be reasonable seems hard, but aligning AI also seems really hard, and maybe ‘reasonable enough’ is easier than ‘aligned enough’? I mean, I think this hasn’t been tried partly because of severe capacity constraints; I think my optimal portfolio has a bit more in this bucket.
Right, but part of what I’m trying to get across here is that the framing counts for a lot, see eg yudkowsky on the drowning child. Like, even among the space of true things there’s a lot of degrees of freedom to exploit people’s intuitions.
Yeah, my take is: democratic oversight is a larger and more complex challenge, but also I really really don’t want to give up on it. I think people vary in how much they feel like “brief period of pseudo-dictatorship” is an acceptable part of the plan; definitely I’d take that over doom but I would also trade some doom-odds to avoid it.
Yeah, I think one of the key challenges which unlocks a lot of coordination capacity is giving people compelling reasons to trust AIs at least in bounded domains, e.g. as genuinely neutral arbitrators.
Yeah, this is one place that things get awfully ‘complex systemsey’, in an exciting but also horrifying way. Some might argue, say, that the Enlightenment ultimately caused a bunch of discoordination and strife (ongoing) in Europe and beyond. Maybe!
But we may need (and be able) to equip people not to coordinate around the most damaging falsehoods (e.g. corrupt populist-demagogic or tech-fantasist) by noticing betrayals and inconsistencies there.
In general the targets of our busting ought to be less identity- and community-foundational and more like political and scientific fads and bubbles. This certainly informs the kind of interventions you should prefer.