Thanks for digging in. Just on a meta level, I want to note that I think understanding the ramifications of developing corrigible AGI, including whether the principal of that AI could cause astronomical suffering, is valid corrigibility research. I’ve already allocated some retroactive prize funding to go to a critic of corrigibility, and I could see rewarding a similarly high-quality argument about why corrigibility research is bad, especially if it significantly changes my mind.
I think we agree that power is unfortunately concentrated in various parts of the world, and that this concentration of power predictably leads to bad things. If the development of corrigible AGI leads to an intense concentration of power in the hands of a few humans, that seems really bad. (Though I am not at all convinced it’s likely to bring about astronomical suffering. Most humans aren’t sadistic psychopaths, and while I would not want Sam Altman to be God Emperor, my guess is that it would be better than getting wiped out by an unfriendly AI. Feel free to lay out reasons if you disagree.)
Part of what I was gesturing towards with “democratic oversight” is that there are known ways to give people limited access to power. The president, for example, is probably the most powerful person in the USA, but I am very confident that he won’t have a third term in office, despite the fact that it would be, in some sense, fairly easy to do. We might imagine similar checks on the principal of a corrigible AGI, such as requiring commands to be submitted in writing with a 24 delay period where a governing body has the ability to review and block commands that are deemed unsafe. By default we might expect the principal of a wisely-built AGI to be a team of many humans, and we could imagine that team needing to be in consensus in order to proceed. And, of course, we have the ability to leverage selection effects. Some humans are far more trustworthy than others, and a wise process for building AGI could arrange for those trustworthy people to be designated as the principal. I don’t think these strategies are guaranteed to work, or are fully mature plans, but they don’t seem obviously doomed. I would certainly like to fund work in thinking about this more, especially insofar as some aspect of corrigibility either undermines or strengthens some pathways for wise governance.
In an effort to sketch something more concrete, let me take Plan A in AI 2040 as a baseline...
In 2029 the president of the USA, recently elected, enacts a bold plan to work together with China to slow down the capability advancement of frontier AI and work on a more prudent solution. The one major difference that I’ll make is that as part of the plan, after the temporary pause, all new AIs are trained to be solely and perfectly corrigible to the governing body of the Consortium, with mundane work done as part of a standing order from that principal to be helpful to human users in straightforward ways. When the new AI models hit an edge case, or believe that someone is trying to jailbreak them or whatever, they reach out to the principal for guidance. Now, who is on the governing body, and are there any checks and balances to prevent oligopoly? Recognizing the extreme risk, the Consortium demands that the principal be a team of 14 people who must be in consensus for the AGI to accept their corrections as valid, except insofar as their correction is to shut down, in which case the AI will obey any of them. The presidents of both the USA and China demand to be part of the council (or they appoint loyalists, which seems overall about the same in expectation), and furthermore get one other government rep each. Let’s say that 5 tech CEOs and experts -- 3 from Western companies and 2 from China—get added. And then the middle powers negotiate to have one rep from each nuclear power except Israel and North Korea: Russia, France, the UK, Pakistan, and India. This body, like the UN Security Council, immediately hits gridlock. With so many veto points, it’s hard to agree on almost anything. Furthermore, the Consortium powers are surveilling the principals and willing to rip most of them out if it looks like they’re trying to conspire to set up an oligarchy with the other members of the principal. Eventually, they agree that they can use the AI to try and find areas of overlap. The AI, being corrigible, is paranoid about manipulation, and starts with very straightforward suggestions: what about curing cancer or inventing ways to cheaply capture carbon from the atmosphere? What about ways to ensure that uncontrolled AIs don’t spring up from blacksites and ruin everything? As much as the members are at each other’s throats, these do sound like good ideas, and eventually an uneasy governance regime sets in, where critics condemn the Consortium of setting up a vetocracy that stifles progress, but nevertheless some progress happens. Lifespans lengthen, and perhaps the less-democratic members have their rulers (and their representatives in the principal) become effectively immortal thanks to longevity tech, but the representatives of more democratic powers are eventually replaced by their nation’s governments (and/or institutions). And thanks to improved information technology provided by the limited ASI, they’re replaced by wiser and more benevolent governors. Eventually, the corrigible AI works with the governing powers to arrange for the creation of an aligned sovereign superintelligence, nearly guaranteed to reflect the true values of its creators, thanks to the alignment work done by the corrigible assistant. The resulting AI produces a utopia that happens to privilege the Chinese power-elite a bit, but is overall visible as a happy and thriving future for humanity.
This story has a bunch of gaps and flaws, and should not be taken as anything more than an off-the-cuff gesture made to help communicate where I’m personally coming from.
I do ultimately think that Plan S is a better baseline. Part of why is that I agree that humanity is on the wrong track for developing good governance systems. We need to do better, and make it a far higher priority. But we’re also on the wrong track vis-a-vis accidentally wiping ourselves out with misaligned agents, though, so I am unconvinced that it’s strongly negative EV. More like there are many ways things could fail and be bad, to varying degrees, and success will involve getting our act together on all fronts. If we wait to do any alignment work until we’re sure that there’s a full and robust solution to misuse (which may be a too-intense strawman of your position), then we’re dooming the futures where temporarily wise governance comes into place, perhaps due to a crisis, warning-shot, and/or the exposure to a novel situation with no established equilibrium pressures. It’s really hard to say, but most days I feel like humanity is already behind where it needs to be on alignment work in order for things to go well.
Thanks for digging in. Just on a meta level, I want to note that I think understanding the ramifications of developing corrigible AGI, including whether the principal of that AI could cause astronomical suffering, is valid corrigibility research. I’ve already allocated some retroactive prize funding to go to a critic of corrigibility, and I could see rewarding a similarly high-quality argument about why corrigibility research is bad, especially if it significantly changes my mind.
I think we agree that power is unfortunately concentrated in various parts of the world, and that this concentration of power predictably leads to bad things. If the development of corrigible AGI leads to an intense concentration of power in the hands of a few humans, that seems really bad. (Though I am not at all convinced it’s likely to bring about astronomical suffering. Most humans aren’t sadistic psychopaths, and while I would not want Sam Altman to be God Emperor, my guess is that it would be better than getting wiped out by an unfriendly AI. Feel free to lay out reasons if you disagree.)
Part of what I was gesturing towards with “democratic oversight” is that there are known ways to give people limited access to power. The president, for example, is probably the most powerful person in the USA, but I am very confident that he won’t have a third term in office, despite the fact that it would be, in some sense, fairly easy to do. We might imagine similar checks on the principal of a corrigible AGI, such as requiring commands to be submitted in writing with a 24 delay period where a governing body has the ability to review and block commands that are deemed unsafe. By default we might expect the principal of a wisely-built AGI to be a team of many humans, and we could imagine that team needing to be in consensus in order to proceed. And, of course, we have the ability to leverage selection effects. Some humans are far more trustworthy than others, and a wise process for building AGI could arrange for those trustworthy people to be designated as the principal. I don’t think these strategies are guaranteed to work, or are fully mature plans, but they don’t seem obviously doomed. I would certainly like to fund work in thinking about this more, especially insofar as some aspect of corrigibility either undermines or strengthens some pathways for wise governance.
In an effort to sketch something more concrete, let me take Plan A in AI 2040 as a baseline...
This story has a bunch of gaps and flaws, and should not be taken as anything more than an off-the-cuff gesture made to help communicate where I’m personally coming from.
I do ultimately think that Plan S is a better baseline. Part of why is that I agree that humanity is on the wrong track for developing good governance systems. We need to do better, and make it a far higher priority. But we’re also on the wrong track vis-a-vis accidentally wiping ourselves out with misaligned agents, though, so I am unconvinced that it’s strongly negative EV. More like there are many ways things could fail and be bad, to varying degrees, and success will involve getting our act together on all fronts. If we wait to do any alignment work until we’re sure that there’s a full and robust solution to misuse (which may be a too-intense strawman of your position), then we’re dooming the futures where temporarily wise governance comes into place, perhaps due to a crisis, warning-shot, and/or the exposure to a novel situation with no established equilibrium pressures. It’s really hard to say, but most days I feel like humanity is already behind where it needs to be on alignment work in order for things to go well.