Hyperbole aside, it’s worth emphasising that something in the realm of AIFEC is a legitimate full-blown wincon up there with aligned AGI. Indeed, a lot of people’s tacit plan seems to be “align the AI and then let it solve democracy and all the rest”, but one could just as well go in the other direction: “solve politics and ignorance, and then just be reasonable about advanced AI”
To a large extent, a big part of the reason this isn’t tried is that it’s far, far more difficult, fora variety of reasons to make the general public be reasonable and that leading to us take the optimal level of risk, at minimum, compared to aligning AGI, and I’d still say this is correct (politics is more tractable than people on LW thought, but this is mostly downstream of warning shots like Mythos that woke up the executive branch, and to a lesser extent congress without needing to wake up the public.)
On top of that, I do think that before you start swinging the club of truth it’s worth taking a beat to ask how pure your motivations really are. The EA/rationality community has a bit of a history of swinging the club of truth in a more hostile way — “save the drowning child” etc.
I agree that EA/rationality (especially EA) has a bad habit of relying on claiming morally-laden facts as true, and more generally one of the single most important constraints for anyone working in the field of AIFEC is to avoid marking moral/value claims as true/correct.
And I also agree that rationalists tend to be blunter and less persuasive than they should, even when it is true.
The problem is, historically there have been many groups that have weaponised very convincing arguments to get their way. One reasonable response to a seemingly faultless argument for a seemingly crazy conclusion is to throw up your hands and assume you’re being swindled — in other words, epistemic learned helplessness. Even if team AIFEC is actually the good guys, people will be correct to be suspicious. Governments in particular depend a lot on restricting what kind of information is admissible, as does the judicial system.
The other issue is that true claims, especially in AI will tend to be very, very complicated relative to simple but wrong claims, and one of the larger updates I’ve made about the AI revolution is that models have to be very, very complicated, by design, and this means you will always be plagued with bias and very difficult to falsify models.
So yeah, epistemic learned helplessness is common.
What coordination and epistemics do we actually need? One way I like to think about this is in terms of what the channel is that leads from the so-called better angels of our nature to actual global action. Probably most of that comes down to robust democratic oversight and international coordination. These are very difficult! And I do not expect that even the top percentile AIFEC startup moves the needle that much, because these processes are by design pretty resistant to hopping on the latest technological fad.
Maybe worth saying here is I mostly don’t buy this theory of change, precisely because I mostly do think robust democratic oversight and international coordination is far too intractable, and I think it’s more useful to focus on smaller sets of actors like the executive branch or individual members of congress.
I still don’t think the products are good, but the standard you are asking for is both extremely unrealistic and unnecessary.
Some people would rather take AI soon so as to increase the odds of their own immortality or the immortality of their family, even at the risk of destroying all of earth. Some people really do want to tile the universe with hedonium. Some people’s order of preferences is basically “my country wins” > “annihilation” > “enemy country wins”.
Some other useful examples here are people who genuinely believe war is a positive good, because that allows them to exterminate whichever disfavored ethnic group, and you don’t care as much about the resources destroyed as you care about extermination.
More prosaically, people can genuinely disagree with liberalism/democracy being good, or believe that centralization into a dictator, junta or oligarchy is actually just good (populist movements are generally this).
More generally, people faking/genuinely believing that their ethnic group is being exterminated, so we need to go to war against the outgroup, no matter the cost is unfortunately pretty common, and critically the belief does not need to be true in order for it to cause wars.
I reel off these problems not because I don’t believe in AIFEC, but because I do believe in what the better version of AIFEC could be. I’m baffled by how overinvested we are in strategies that, from my perspective, seem specifically geared towards helping frontier labs navigate the acute risk period. I think in an adequate world we’d be pushing simultaneously on every plausible path to victory, at least until we hit the thresholds where they started to trade hard against each other.
Agree that frontier lab epistemics are likely over-represented in AIFEC, but I’d have a different distribution of work that is valuable to work on than you do.
One obvious reason to be long on AIFEC is that it automatically scales with AI capabilities. Some version of it is clearly coming, and will clearly be helpful. And right now in particular there’s some opportunity to steer things, to get certain balls rolling sooner. Even if you’re extremely bitter lesson-pilled, a three month lead could be worth a lot.
It’s worth saying though that this will require people to trust the AIs pretty largely, and frontier labs have obvious and large incentives to do this, but most others don’t and indeed have anti-incentives to trust AIs, and even if the AIs are in fact jusitified in being trustworthy, it’s plausible that people won’t use AI because they incorrectly trust the AIs.
More so than I’d like, AIFEC tends to privilege individual-level agency, whereas I am pretty bought in on group agency being real and very important for risks.
To be fair, that’s because group-level agency without incentives is rare, and most claims of groups acting agentically are describable in terms of their incentives.
Put another way, AIFEC already grants that group-level agency will exist assuming the incentives are right, and it’s correct to limit it to that.
AIFEC doesn’t grapple as much as I’d like with power and leverage, and indeed the frame seems to slightly bounce off them. This is related to problem 4 — AIFEC hasn’t grappled as much as I’d like with conflict theory.
I can plausibly agree with this, though I also think power/leverage is probably overrated as a problem, and this is mostly due to my view that a lot of the more negative results of absolute power/leverage only really hold when you have don’t have a small bit of caring (assuming AIs make the cost of loyalty go to essentially 0 through alignment), and while this does exist (given psychopaths/sociopaths), this is actually pretty rare, and more generally I think the amount of people that have non-zero altruism is large enough such that we can actually have leaders equivalent to the wise/benevolent ones we often see in fiction (and because AI enables very long lifespans, this means that leadership turnover stops, meaning a potential degradation of leadership skill is avoided).
(The main mechanism here is basically the tails coming apart, but in a positive direction).
And finally, my main conclusion is that if you want to work on AIFEC, you should probably focus on smaller groups as your end goal than trying to get large groups of the general public, and I give some reasons for hope on why epistemic efforts focused on small groups could make outcomes much better, even if they don’t involve democracy or liberalism.
far, far more difficult, fora variety of reasons to make the general public be reasonable
Honestly, I’m not sure. Like, yeah, getting everyone to be reasonable seems hard, but aligning AI also seems really hard, and maybe ‘reasonable enough’ is easier than ‘aligned enough’? I mean, I think this hasn’t been tried partly because of severe capacity constraints; I think my optimal portfolio has a bit more in this bucket.
I agree that EA/rationality (especially EA) has a bad habit of relying on claiming morally-laden facts as true,
Right, but part of what I’m trying to get across here is that the framing counts for a lot, see eg yudkowsky on the drowning child. Like, even among the space of true things there’s a lot of degrees of freedom to exploit people’s intuitions.
I mostly do think robust democratic oversight and international coordination is far too intractable, and I think it’s more useful to focus on smaller sets of actors like the executive branch or individual members of congress.
Yeah, my take is: democratic oversight is a larger and more complex challenge, but also I really really don’t want to give up on it. I think people vary in how much they feel like “brief period of pseudo-dictatorship” is an acceptable part of the plan; definitely I’d take that over doom but I would also trade some doom-odds to avoid it.
this will require people to trust the AIs pretty largely
Yeah, I think one of the key challenges which unlocks a lot of coordination capacity is giving people compelling reasons to trust AIs at least in bounded domains, e.g. as genuinely neutral arbitrators.
To a large extent, a big part of the reason this isn’t tried is that it’s far, far more difficult, fora variety of reasons to make the general public be reasonable and that leading to us take the optimal level of risk, at minimum, compared to aligning AGI, and I’d still say this is correct (politics is more tractable than people on LW thought, but this is mostly downstream of warning shots like Mythos that woke up the executive branch, and to a lesser extent congress without needing to wake up the public.)
I agree that EA/rationality (especially EA) has a bad habit of relying on claiming morally-laden facts as true, and more generally one of the single most important constraints for anyone working in the field of AIFEC is to avoid marking moral/value claims as true/correct.
And I also agree that rationalists tend to be blunter and less persuasive than they should, even when it is true.
The other issue is that true claims, especially in AI will tend to be very, very complicated relative to simple but wrong claims, and one of the larger updates I’ve made about the AI revolution is that models have to be very, very complicated, by design, and this means you will always be plagued with bias and very difficult to falsify models.
So yeah, epistemic learned helplessness is common.
Maybe worth saying here is I mostly don’t buy this theory of change, precisely because I mostly do think robust democratic oversight and international coordination is far too intractable, and I think it’s more useful to focus on smaller sets of actors like the executive branch or individual members of congress.
I still don’t think the products are good, but the standard you are asking for is both extremely unrealistic and unnecessary.
Some other useful examples here are people who genuinely believe war is a positive good, because that allows them to exterminate whichever disfavored ethnic group, and you don’t care as much about the resources destroyed as you care about extermination.
More prosaically, people can genuinely disagree with liberalism/democracy being good, or believe that centralization into a dictator, junta or oligarchy is actually just good (populist movements are generally this).
More generally, people faking/genuinely believing that their ethnic group is being exterminated, so we need to go to war against the outgroup, no matter the cost is unfortunately pretty common, and critically the belief does not need to be true in order for it to cause wars.
Agree that frontier lab epistemics are likely over-represented in AIFEC, but I’d have a different distribution of work that is valuable to work on than you do.
It’s worth saying though that this will require people to trust the AIs pretty largely, and frontier labs have obvious and large incentives to do this, but most others don’t and indeed have anti-incentives to trust AIs, and even if the AIs are in fact jusitified in being trustworthy, it’s plausible that people won’t use AI because they incorrectly trust the AIs.
To be fair, that’s because group-level agency without incentives is rare, and most claims of groups acting agentically are describable in terms of their incentives.
Put another way, AIFEC already grants that group-level agency will exist assuming the incentives are right, and it’s correct to limit it to that.
I can plausibly agree with this, though I also think power/leverage is probably overrated as a problem, and this is mostly due to my view that a lot of the more negative results of absolute power/leverage only really hold when you have don’t have a small bit of caring (assuming AIs make the cost of loyalty go to essentially 0 through alignment), and while this does exist (given psychopaths/sociopaths), this is actually pretty rare, and more generally I think the amount of people that have non-zero altruism is large enough such that we can actually have leaders equivalent to the wise/benevolent ones we often see in fiction (and because AI enables very long lifespans, this means that leadership turnover stops, meaning a potential degradation of leadership skill is avoided).
(The main mechanism here is basically the tails coming apart, but in a positive direction).
And finally, my main conclusion is that if you want to work on AIFEC, you should probably focus on smaller groups as your end goal than trying to get large groups of the general public, and I give some reasons for hope on why epistemic efforts focused on small groups could make outcomes much better, even if they don’t involve democracy or liberalism.
Thanks for the detailed comments!
Honestly, I’m not sure. Like, yeah, getting everyone to be reasonable seems hard, but aligning AI also seems really hard, and maybe ‘reasonable enough’ is easier than ‘aligned enough’? I mean, I think this hasn’t been tried partly because of severe capacity constraints; I think my optimal portfolio has a bit more in this bucket.
Right, but part of what I’m trying to get across here is that the framing counts for a lot, see eg yudkowsky on the drowning child. Like, even among the space of true things there’s a lot of degrees of freedom to exploit people’s intuitions.
Yeah, my take is: democratic oversight is a larger and more complex challenge, but also I really really don’t want to give up on it. I think people vary in how much they feel like “brief period of pseudo-dictatorship” is an acceptable part of the plan; definitely I’d take that over doom but I would also trade some doom-odds to avoid it.
Yeah, I think one of the key challenges which unlocks a lot of coordination capacity is giving people compelling reasons to trust AIs at least in bounded domains, e.g. as genuinely neutral arbitrators.