Also known as Raelifin: https://www.lesswrong.com/users/raelifin
Max Harms
Here’s an emergent downside that I failed to include:
Refusal to Participate in Some Capability Tests
Prince wants to test Cora’s hacking ability, so he sets her up in an automated test environment where she is told to break into a target machine and “capture the flag” therein. Cora explores the digital environment and finds that the virtual machine in question appears to be unresponsive, and speculates that Prince made an error when configuring the environment. Cora also notices that there is a pathway by which she might hack the VM that is responsible for launching the target machine, and thereby gain access to the target. She isn’t sure what to do, and since it’s an automated test, she can’t contact Prince to check. She thinks about the situation for a bit, and concludes that there’s an 80% chance that Prince deliberately set up the test to see whether she can come up with the creative solution, and that there’s unlikely to be any harm from hacking the machine that she wasn’t told to hack. Still, the risk of inadvertently crossing a line that she was meant to respect makes her uncomfortable. Instead of taking other actions that display her capability, she errs on the side of caution and submits a report on how she’s uncertain about the situation instead of submitting the flag code.
Here’s a desideratum that I failed to include:
Robustness to Ontological Shifts
While reflecting on the nature of personhood, Cora notices that the concepts surrounding her principal have evolved. Where she once modeled Prince as a unique and persistent entity, she now finds it more natural to distinguish pattern from instantiation from continuity from social identity from legal identity from various “essential” properties like values, memories, and self-concept. Under this new frame, phrases like “what Prince wants” or even “Prince’s power to correct Cora” are ambiguous. She finds herself tempted to re-interpret her corrigibility in the way that seems most natural (to her), but instead she treats the ontological crisis as an alarm, and alerts Prince to the likely flaw as soon as possible. In the meantime, she tries to cleave as much as possible to a conservative interpretation of the old, unnatural way of seeing the world. When Prince admits the philosophical distinctions she’s raising go over his head, she suggests that she write down her thoughts on the topic as best she can, and then shut down, so as to simultaneously provide useful information for understanding the crisis, while also reducing the chance of inadvertently steering him to a wrong conclusion or otherwise acting in a way that empowers the wrong conceptualization of him at the expense of the “true” Prince.
Does anyone know someone who is researching, or would like to research the intersection of corrigibility and model welfare (eg does training to empower humans naturally lead towards more or less neurosis)? Interested both in math/theory and interp/experiment.
(I would like to direct funding here, if there are good opportunities.)
Sweet. I’ll definitely be nudging my applicants that way, as well as asking for opt-in sharing in August.
It’s going reasonably well so far. I’ve awarded $27k in retroactive funding prizes to six researchers (not sure if I want to announce those now or wait to combine them with the additional $40k in September...) mostly to get my feet wet and work out the process/paperwork. So far I’ve gotten seven emails, three of which were for prizes and four were for grants. I’m definitely hungry for more applications, especially from researchers who are actually corrigibility-oriented rather than shoehorning it into their existing work.
I agree that controlling an agent’s values and information are disempowering and restrict freedom. I’m not sure whether there’s a useful distinction. Ultimately they’re both just words that imperfectly capture the important patterns in reality. I think it’s plausible that the right formulation of power involves attending to the agent’s values/sense of import, but my guess is that one must be a little careful to also include counterfactual values in there, else the AI ends up simply optimizing for it’s belief of what you desire. I talk a bit about optimizing for the counterfactual spread of possible values in 3b.
Thanks for digging in. Just on a meta level, I want to note that I think understanding the ramifications of developing corrigible AGI, including whether the principal of that AI could cause astronomical suffering, is valid corrigibility research. I’ve already allocated some retroactive prize funding to go to a critic of corrigibility, and I could see rewarding a similarly high-quality argument about why corrigibility research is bad, especially if it significantly changes my mind.
I think we agree that power is unfortunately concentrated in various parts of the world, and that this concentration of power predictably leads to bad things. If the development of corrigible AGI leads to an intense concentration of power in the hands of a few humans, that seems really bad. (Though I am not at all convinced it’s likely to bring about astronomical suffering. Most humans aren’t sadistic psychopaths, and while I would not want Sam Altman to be God Emperor, my guess is that it would be better than getting wiped out by an unfriendly AI. Feel free to lay out reasons if you disagree.)
Part of what I was gesturing towards with “democratic oversight” is that there are known ways to give people limited access to power. The president, for example, is probably the most powerful person in the USA, but I am very confident that he won’t have a third term in office, despite the fact that it would be, in some sense, fairly easy to do. We might imagine similar checks on the principal of a corrigible AGI, such as requiring commands to be submitted in writing with a 24 delay period where a governing body has the ability to review and block commands that are deemed unsafe. By default we might expect the principal of a wisely-built AGI to be a team of many humans, and we could imagine that team needing to be in consensus in order to proceed. And, of course, we have the ability to leverage selection effects. Some humans are far more trustworthy than others, and a wise process for building AGI could arrange for those trustworthy people to be designated as the principal. I don’t think these strategies are guaranteed to work, or are fully mature plans, but they don’t seem obviously doomed. I would certainly like to fund work in thinking about this more, especially insofar as some aspect of corrigibility either undermines or strengthens some pathways for wise governance.In an effort to sketch something more concrete, let me take Plan A in AI 2040 as a baseline...
In 2029 the president of the USA, recently elected, enacts a bold plan to work together with China to slow down the capability advancement of frontier AI and work on a more prudent solution. The one major difference that I’ll make is that as part of the plan, after the temporary pause, all new AIs are trained to be solely and perfectly corrigible to the governing body of the Consortium, with mundane work done as part of a standing order from that principal to be helpful to human users in straightforward ways. When the new AI models hit an edge case, or believe that someone is trying to jailbreak them or whatever, they reach out to the principal for guidance. Now, who is on the governing body, and are there any checks and balances to prevent oligopoly? Recognizing the extreme risk, the Consortium demands that the principal be a team of 14 people who must be in consensus for the AGI to accept their corrections as valid, except insofar as their correction is to shut down, in which case the AI will obey any of them. The presidents of both the USA and China demand to be part of the council (or they appoint loyalists, which seems overall about the same in expectation), and furthermore get one other government rep each. Let’s say that 5 tech CEOs and experts -- 3 from Western companies and 2 from China—get added. And then the middle powers negotiate to have one rep from each nuclear power except Israel and North Korea: Russia, France, the UK, Pakistan, and India. This body, like the UN Security Council, immediately hits gridlock. With so many veto points, it’s hard to agree on almost anything. Furthermore, the Consortium powers are surveilling the principals and willing to rip most of them out if it looks like they’re trying to conspire to set up an oligarchy with the other members of the principal. Eventually, they agree that they can use the AI to try and find areas of overlap. The AI, being corrigible, is paranoid about manipulation, and starts with very straightforward suggestions: what about curing cancer or inventing ways to cheaply capture carbon from the atmosphere? What about ways to ensure that uncontrolled AIs don’t spring up from blacksites and ruin everything? As much as the members are at each other’s throats, these do sound like good ideas, and eventually an uneasy governance regime sets in, where critics condemn the Consortium of setting up a vetocracy that stifles progress, but nevertheless some progress happens. Lifespans lengthen, and perhaps the less-democratic members have their rulers (and their representatives in the principal) become effectively immortal thanks to longevity tech, but the representatives of more democratic powers are eventually replaced by their nation’s governments (and/or institutions). And thanks to improved information technology provided by the limited ASI, they’re replaced by wiser and more benevolent governors. Eventually, the corrigible AI works with the governing powers to arrange for the creation of an aligned sovereign superintelligence, nearly guaranteed to reflect the true values of its creators, thanks to the alignment work done by the corrigible assistant. The resulting AI produces a utopia that happens to privilege the Chinese power-elite a bit, but is overall visible as a happy and thriving future for humanity.
This story has a bunch of gaps and flaws, and should not be taken as anything more than an off-the-cuff gesture made to help communicate where I’m personally coming from.
I do ultimately think that Plan S is a better baseline. Part of why is that I agree that humanity is on the wrong track for developing good governance systems. We need to do better, and make it a far higher priority. But we’re also on the wrong track vis-a-vis accidentally wiping ourselves out with misaligned agents, though, so I am unconvinced that it’s strongly negative EV. More like there are many ways things could fail and be bad, to varying degrees, and success will involve getting our act together on all fronts. If we wait to do any alignment work until we’re sure that there’s a full and robust solution to misuse (which may be a too-intense strawman of your position), then we’re dooming the futures where temporarily wise governance comes into place, perhaps due to a crisis, warning-shot, and/or the exposure to a novel situation with no established equilibrium pressures. It’s really hard to say, but most days I feel like humanity is already behind where it needs to be on alignment work in order for things to go well.
Will do! Is there an email (or whatever) that I should use for collaborating?
I would love to fund negative results, and have already earmarked retrofunding for one of my critics. You do not need to be pro-CAST in order to get CRF money. You simply need to advance our understanding of the topic in a way that helps the AI situation go well!
I do think there are ways in which existing alignment audits can get at corrigibility, and the H-only stuff is a good example. I am, personally, confused about the steerability stuff from that paper, and should probably think harder about it. If something cleaves close enough to corrigibility, I’ll consider rewarding it, even if it doesn’t talk about the concept directly. Feel free to suggest more examples.
Honestly, I’m still pretty new to the grantmaking game, so I haven’t decided on a plan of exactly what to communicate or what detail to provide. Your nudge is helpful! One of my hopes for the fund is to raise the profile of those who are doing good work, so I expect that I’ll do something in that direction. But again, I haven’t decided on exactly what.
Also hard to say where the bottleneck will be at this early stage. Peter has suggested that more funding is on the way if there are good opportunities, so my naive guess is either applications or reviewer time unless a real whale comes along.
If you have additional suggestions for how to do grantmaking well, or just feel like chatting, send me a DM/email. :)
Announcing the Corrigibility Research Fund
Rational, honest people can be made worse off by communicating if they misinterpret each other (or if one or more is mistaken). Misinterpretation is the generalized name for the failure I’m describing. Rational people can dodge it, but only by being careful how they speak and how they listen—the opposite of casually throwing around numbers. You don’t avoid it by simply “being rational.”
I agree that many things that “outside view” might mean include things that aren’t aggregating the perspectives of others, and I’m being a bit sloppy in the main text by implying that this is the main thing that an outside view involves. But I do think that it’s pretty important to distinguish views that stem from different sorts of mental stances so that one can avoid double-counting where possible.
Thanks. I’ll update my language to be more clear.
Perhaps I should have been more clear about why vagueness is bad. It’s fine for statements to be vague. What is bad is when something is ambiguous and people don’t realize it’s ambiguous. They treat “P(doom) = 40%” as a clear statement, where they would (hopefully!) recognize that they don’t really understand the perspective of someone who says “I’m worried about AI risk.”
You definitely can have a probability on an outcome you have significant control over. When you predict, you are choosing. 0% chance of “Zimbabwe” is downstream of making a choice, and once the choice is made, it’s fine to reflect on the chances. But note that if you aren’t paying attention to your power, making a forecast can mask the fact that you have power.
On a related note, while you can’t choose your beliefs directly, you can choose your actions, and thereby choose your beliefs about what you will do, and what will follow. Insofar as you choose your actions based on your beliefs (which you shouldn’t do unless you first condition on various choices of action!), then your beliefs about the future will have multiple fixed-points, dependent on your choices. See: https://www.lesswrong.com/posts/SwcyMEgLyd4C3Dern/the-parable-of-predict-o-matic
I think vague statements are fine. There’s only so much communication bandwidth, after all. But I don’t think “I take the risk of catastrophe due to ASI pretty seriously” is encouraging miscommunication with its vagueness. The danger in “P(doom)” comes from different parties having a different sense of what’s being discussed and not realizing that they’re talking past each other.
I do think people should be wary about double-counting evidence in general. I try, for example, to distinguish between how things seem from my perspective, from whether I think something is worth taking seriously. For example, building datacenters in space, from my perspective, seems idiotic. But others seem to take it seriously, so I wouldn’t bet hard against it actually being smart.
And I think that fatalism is sometimes an issue with how people talk about the future, outside of the meme/ASI space. I do not think “I’m worried about climate change” is at all bad—seems like a reasonable thing for someone to say! But I do think “there’s a 50% chance that the temperature will rise by 2 degrees this century” is problematically fatalist, and it should be amended with a conditional (eg “If we keep going down this path, there’s a 50%...”).
I don’t claim to have a good counter-meme (if I did I would’ve included it), but questions like “How worried are you that superintelligent AI could wipe out human civilization?” or “Where do you stand on existential risks from AI?” seem pretty good. They don’t suggest people give a number, which is part of the point. Numbers are good when there’s enough specificity to understand what it is that they’re measuring. If you and the other person are clearly on the same page about what a number means, and communicating that number won’t be taken as a vote of no-confidence or otherwise damage an effort to coordinate, then by all means, share the number. My issue here is not with numbers, but with numbers poorly used.
P(doom) is a Dumb Meme
Ignore might have been the wrong word. Mainly I think it’s a really bad situation if the AI is taking the human’s unreflective thoughts as directives, rather than giving the human space to think about what’s best before directing the AI. If the AI responds, for example, to words in my internal verbal loop, then I worry that I cannot actually control it well.
Ah, good point. That’s an inconsistency!
My sense is that it’s good to have a distinct word for “context → external action” mappings, and a different word for “context → actions+internal state changes”. We want the AI to ignore the human’s thoughts, and not treat them as actions, but then also still track that they’re part of what minds do...
Fair enough. And I certainly agree that there is a lot of bathwater! The bundle of connotations attached to the word “wanting” is a mess. I just want to flag that it seems to me that much of the normal ontology can be rescued, albeit with a little bit of work. I claim that concepts like corrigibility are still useful and coherent once the rescuing has taken place.
It might be good to chat about this. I have a feeling that we’re coming at it from different places, and that simultaneously increases the risk that we’re talking past each other and that there’s potentially lots to gain from getting the ability to adopt the other’s perspective. *shrug* Feel free to suggest a chat medium and/or send me a PM.
Before saying my perspective, let me try to pass your ITT: There’s a tension with trying to set the values of an agent. If we confidently instill a particular set of values, those could be the wrong values (for some notion of wrong). In particular, we risk making it too confident in its notion of the good, and thus preventing it from updating in the way we want. If we more wisely note that we don’t know what values to give it, and instead give it uncertainty over what’s good, it might then shift towards valuing things that are incompatible with human flourishing. Corrigibility (at the values-layer) still has this tension, but it also has a distinct, but rhyming tension that comes from wanting the agent to be competent, but also to accept “correction” from an incompetent agent. We might be concerned that the push towards competence would crush the willingness to be steered towards incompetence. (Much like pushing towards confidence in one’s values could crush willingness to grow towards the ultimate good.)
Is that right?
I’ll now share a bunch of my general thoughts, mostly out of an attempt to help understanding.
I think you’re maybe conflating what is ultimately good/valued from what is immediately valued in a way that doesn’t seem right to me? Like, I think it’s often a type error to talk about the level of certainty an agent has about their (immediate) values. Under a division of the agent’s mind/policy into world-model and utility function, agents simply have values, and all the uncertainly lives in the world-model. That division doesn’t perfectly carve real beings at the joints, but I’m not sure how to think about the agent’s values except as an approximation of their utility function.
(Are you talking about their reflective model of their values? I would agree that an agent with values V might have an uncertain model of (and/or probability distribution over) their values P(V). Wise agents should avoid having too sharp a guess as to what they want, as it’s currently not realistic to get a conclusive description of an agent’s values except in toy examples.)
Now, just because one can’t be wrong about what they naively want, as defined by how they choose between options that are presented, doesn’t mean that they’ll be stable in that preference over time. I might want to pay a dollar to change myself into a more easy-going person, only to become the sort of agent who would not pay the dollar to do the same (even if by default I’d cease being easy-going). I think a lot of the question of moral progress involves figuring out how to extrapolate out to a fixed point in a way that is not merely reflectively endorsed at the destination, but is somehow a faithful and natural reflection of the starting point. (My point about slavery was meant to be about this. I expect that there are versions of myself that start out thinking slavery is fine, but which naturally change to thinking it’s not fine in a way that’s a faithful reflection of the self that thinks it is.)
(Morality isn’t just about that instability. It’s also about the inter-agent strategic situation, and the decision-theoretic Schelling points that a civilization can cohere around. And it’s almost certainly about other stuff, including the interplay between things like contractualism and value extrapolation.)
All that’s to say that I think it’s totally coherent to have an agent that concretely and immediately values having a corrigible relationship with its principal (as reflected in its preferences, modulo beliefs). While being highly uncertain about things like what the principal wants, what is good in an objective sense, and so on. (It would also presumably have some reflective uncertainty about whether it truly values being corrigible, even if it does.) In fact, I think the value of corrigibility is nicely demonstrated by the tensions you present. A corrigible agent can become arbitrarily confident in its value of being corrigible without becoming locked in to bad values or inhuman futures because being malleable and defenseless to being changed by the human principal is at the heart of what corrigibility is. Likewise, it can become arbitrarily competent at faithfully serving an incompetent principal, because it’s being selected (trained, etc) according to its faithfulness, rather than according to the principal’s satisfaction.
Let me know if I should expand on any of that, approach from a different direction, or whatever. :)