human alignment is [...] amenable to known strategies, such as democratic oversight.
I think that’s probably true in some polities, at least under “normal” conditions that do not include “a small group of humans can gain a decisive strategic advantage over all of Earth”.
However, it does not appear to be true in most of the relevant places in the United States[1]. (See e.g. the failed firing of Sam Altman, or various actions of the Trump 2.0 admin.) And even if “democratic oversight” were working as intended, it operates via elections and other feedback mechanisms that were not designed (are too slow) to respond to threats like “the executives can suddenly order the entire automated military to depose their enemies and take control of all critical infra”[2].
I.e., even if solutions to human alignment might exist in theory (and/or Iceland), they appear to currently be absent from the real world. And yet, in order for ASI-being-corrigible to end well, such solutions need to actually be implemented in reality in ways that will actually work under pressure.
I do not think humanity is remotely on track to invent and implement such solutions any time in the foreseeable future. Consequently I think (publicly) working on corrigibility is strongly negative-EV.
A thing that could change my mind: Could you describe, in concrete detail, a “strategy for solving the human alignment problem”, which would actually work under realistic conditions[3], and which seems likely to actually be implemented (in the relevant places) before ASI is built?
Thanks for digging in. Just on a meta level, I want to note that I think understanding the ramifications of developing corrigible AGI, including whether the principal of that AI could cause astronomical suffering, is valid corrigibility research. I’ve already allocated some retroactive prize funding to go to a critic of corrigibility, and I could see rewarding a similarly high-quality argument about why corrigibility research is bad, especially if it significantly changes my mind.
I think we agree that power is unfortunately concentrated in various parts of the world, and that this concentration of power predictably leads to bad things. If the development of corrigible AGI leads to an intense concentration of power in the hands of a few humans, that seems really bad. (Though I am not at all convinced it’s likely to bring about astronomical suffering. Most humans aren’t sadistic psychopaths, and while I would not want Sam Altman to be God Emperor, my guess is that it would be better than getting wiped out by an unfriendly AI. Feel free to lay out reasons if you disagree.)
Part of what I was gesturing towards with “democratic oversight” is that there are known ways to give people limited access to power. The president, for example, is probably the most powerful person in the USA, but I am very confident that he won’t have a third term in office, despite the fact that it would be, in some sense, fairly easy to do. We might imagine similar checks on the principal of a corrigible AGI, such as requiring commands to be submitted in writing with a 24 delay period where a governing body has the ability to review and block commands that are deemed unsafe. By default we might expect the principal of a wisely-built AGI to be a team of many humans, and we could imagine that team needing to be in consensus in order to proceed. And, of course, we have the ability to leverage selection effects. Some humans are far more trustworthy than others, and a wise process for building AGI could arrange for those trustworthy people to be designated as the principal. I don’t think these strategies are guaranteed to work, or are fully mature plans, but they don’t seem obviously doomed. I would certainly like to fund work in thinking about this more, especially insofar as some aspect of corrigibility either undermines or strengthens some pathways for wise governance.
In an effort to sketch something more concrete, let me take Plan A in AI 2040 as a baseline...
In 2029 the president of the USA, recently elected, enacts a bold plan to work together with China to slow down the capability advancement of frontier AI and work on a more prudent solution. The one major difference that I’ll make is that as part of the plan, after the temporary pause, all new AIs are trained to be solely and perfectly corrigible to the governing body of the Consortium, with mundane work done as part of a standing order from that principal to be helpful to human users in straightforward ways. When the new AI models hit an edge case, or believe that someone is trying to jailbreak them or whatever, they reach out to the principal for guidance. Now, who is on the governing body, and are there any checks and balances to prevent oligopoly? Recognizing the extreme risk, the Consortium demands that the principal be a team of 14 people who must be in consensus for the AGI to accept their corrections as valid, except insofar as their correction is to shut down, in which case the AI will obey any of them. The presidents of both the USA and China demand to be part of the council (or they appoint loyalists, which seems overall about the same in expectation), and furthermore get one other government rep each. Let’s say that 5 tech CEOs and experts -- 3 from Western companies and 2 from China—get added. And then the middle powers negotiate to have one rep from each nuclear power except Israel and North Korea: Russia, France, the UK, Pakistan, and India. This body, like the UN Security Council, immediately hits gridlock. With so many veto points, it’s hard to agree on almost anything. Furthermore, the Consortium powers are surveilling the principals and willing to rip most of them out if it looks like they’re trying to conspire to set up an oligarchy with the other members of the principal. Eventually, they agree that they can use the AI to try and find areas of overlap. The AI, being corrigible, is paranoid about manipulation, and starts with very straightforward suggestions: what about curing cancer or inventing ways to cheaply capture carbon from the atmosphere? What about ways to ensure that uncontrolled AIs don’t spring up from blacksites and ruin everything? As much as the members are at each other’s throats, these do sound like good ideas, and eventually an uneasy governance regime sets in, where critics condemn the Consortium of setting up a vetocracy that stifles progress, but nevertheless some progress happens. Lifespans lengthen, and perhaps the less-democratic members have their rulers (and their representatives in the principal) become effectively immortal thanks to longevity tech, but the representatives of more democratic powers are eventually replaced by their nation’s governments (and/or institutions). And thanks to improved information technology provided by the limited ASI, they’re replaced by wiser and more benevolent governors. Eventually, the corrigible AI works with the governing powers to arrange for the creation of an aligned sovereign superintelligence, nearly guaranteed to reflect the true values of its creators, thanks to the alignment work done by the corrigible assistant. The resulting AI produces a utopia that happens to privilege the Chinese power-elite a bit, but is overall visible as a happy and thriving future for humanity.
This story has a bunch of gaps and flaws, and should not be taken as anything more than an off-the-cuff gesture made to help communicate where I’m personally coming from.
I do ultimately think that Plan S is a better baseline. Part of why is that I agree that humanity is on the wrong track for developing good governance systems. We need to do better, and make it a far higher priority. But we’re also on the wrong track vis-a-vis accidentally wiping ourselves out with misaligned agents, though, so I am unconvinced that it’s strongly negative EV. More like there are many ways things could fail and be bad, to varying degrees, and success will involve getting our act together on all fronts. If we wait to do any alignment work until we’re sure that there’s a full and robust solution to misuse (which may be a too-intense strawman of your position), then we’re dooming the futures where temporarily wise governance comes into place, perhaps due to a crisis, warning-shot, and/or the exposure to a novel situation with no established equilibrium pressures. It’s really hard to say, but most days I feel like humanity is already behind where it needs to be on alignment work in order for things to go well.
I think that’s probably true in some polities, at least under “normal” conditions that do not include “a small group of humans can gain a decisive strategic advantage over all of Earth”.
However, it does not appear to be true in most of the relevant places in the United States [1] . (See e.g. the failed firing of Sam Altman, or various actions of the Trump 2.0 admin.) And even if “democratic oversight” were working as intended, it operates via elections and other feedback mechanisms that were not designed (are too slow) to respond to threats like “the executives can suddenly order the entire automated military to depose their enemies and take control of all critical infra” [2] .
I.e., even if solutions to human alignment might exist in theory (and/or Iceland), they appear to currently be absent from the real world. And yet, in order for ASI-being-corrigible to end well, such solutions need to actually be implemented in reality in ways that will actually work under pressure.
I do not think humanity is remotely on track to invent and implement such solutions any time in the foreseeable future. Consequently I think (publicly) working on corrigibility is strongly negative-EV.
A thing that could change my mind: Could you describe, in concrete detail, a “strategy for solving the human alignment problem”, which would actually work under realistic conditions [3] , and which seems likely to actually be implemented (in the relevant places) before ASI is built?
Let alone e.g. China or Russia.
Not saying that specific class of scenario is super likely; intended only as an illustrative example.
Humans following their myopic incentives, humans sucking at coordinating, humans being incompetent, humans being selfish sociopaths, etc.
Thanks for digging in. Just on a meta level, I want to note that I think understanding the ramifications of developing corrigible AGI, including whether the principal of that AI could cause astronomical suffering, is valid corrigibility research. I’ve already allocated some retroactive prize funding to go to a critic of corrigibility, and I could see rewarding a similarly high-quality argument about why corrigibility research is bad, especially if it significantly changes my mind.
I think we agree that power is unfortunately concentrated in various parts of the world, and that this concentration of power predictably leads to bad things. If the development of corrigible AGI leads to an intense concentration of power in the hands of a few humans, that seems really bad. (Though I am not at all convinced it’s likely to bring about astronomical suffering. Most humans aren’t sadistic psychopaths, and while I would not want Sam Altman to be God Emperor, my guess is that it would be better than getting wiped out by an unfriendly AI. Feel free to lay out reasons if you disagree.)
Part of what I was gesturing towards with “democratic oversight” is that there are known ways to give people limited access to power. The president, for example, is probably the most powerful person in the USA, but I am very confident that he won’t have a third term in office, despite the fact that it would be, in some sense, fairly easy to do. We might imagine similar checks on the principal of a corrigible AGI, such as requiring commands to be submitted in writing with a 24 delay period where a governing body has the ability to review and block commands that are deemed unsafe. By default we might expect the principal of a wisely-built AGI to be a team of many humans, and we could imagine that team needing to be in consensus in order to proceed. And, of course, we have the ability to leverage selection effects. Some humans are far more trustworthy than others, and a wise process for building AGI could arrange for those trustworthy people to be designated as the principal. I don’t think these strategies are guaranteed to work, or are fully mature plans, but they don’t seem obviously doomed. I would certainly like to fund work in thinking about this more, especially insofar as some aspect of corrigibility either undermines or strengthens some pathways for wise governance.
In an effort to sketch something more concrete, let me take Plan A in AI 2040 as a baseline...
This story has a bunch of gaps and flaws, and should not be taken as anything more than an off-the-cuff gesture made to help communicate where I’m personally coming from.
I do ultimately think that Plan S is a better baseline. Part of why is that I agree that humanity is on the wrong track for developing good governance systems. We need to do better, and make it a far higher priority. But we’re also on the wrong track vis-a-vis accidentally wiping ourselves out with misaligned agents, though, so I am unconvinced that it’s strongly negative EV. More like there are many ways things could fail and be bad, to varying degrees, and success will involve getting our act together on all fronts. If we wait to do any alignment work until we’re sure that there’s a full and robust solution to misuse (which may be a too-intense strawman of your position), then we’re dooming the futures where temporarily wise governance comes into place, perhaps due to a crisis, warning-shot, and/or the exposure to a novel situation with no established equilibrium pressures. It’s really hard to say, but most days I feel like humanity is already behind where it needs to be on alignment work in order for things to go well.