Formerly alignment and governance researcher at DeepMind and OpenAI. Now independent.
Richard_Ngo
Mused on it for a bit, and there’s no particularly notable feeling.
As per my reply to Jan above, I’m most interested in groups where having more integrity is a realistic plan—e.g. the early rationalists. I’m not trying to swing the whole alignment community around.
These might be much smaller than the groups you’re thinking of but those groups can grow in influence very fast (again, see the early rationalists). And then they can figure out how to scale in healthy ways as they go along.
Maybe a crux is that I don’t think the current alignment community as an entity is a very relevant actor, because it’s so messed up by lack of internal clarity and integrity that it’s lost the ability to steer. For example, there’s nobody who’s psychologically capable of pivoting even just the Constellation cluster into a weird and risky plan (like going hard on pause advocacy), in part because weird risky plans go too strongly against people’s marginalist intuitions. (Not saying Constellation should do that, just that groups are mainly interesting insofar as they have the capacity to do things like that.)
Is there any writeup you can point me to of the argument you want to defend? I’d like to engage with the best version of it, but I’m worried that it’s amorphous enough that different people will end up defending different versions at different times.
I think something that could’ve been clearer in the last section is that dramatically changing up your life plan is not the only possible immediate next step if someone feels moved by this post—one could also start marginally increasing the robustness of one’s strategy while staying in a roughly similar position in life; e.g. stay at a frontier lab but become more like Leo Gao, increase your everyday integrity and psychological health, etc.
This is what I intended with the section starting “For now, I’ll focus on a few high-level principles for starting to move in that direction.”
However, I did wonder as I was writing it if even that was making too much of a concession—because I do think that impact here is heavy-tailed towards people who are able and willing to carve out unusual paths. (FWIW I get some impression from our online interactions that you could end up as one of them eventually.)
Relatedly, I’d nudge you to think more about what you meant with the phrase “the only possible immediate next step”—seems like there’s a bunch of implicit social stuff underlying it.
For my part, I think there’s some implicit message underlying a lot of my writing that I should try to convey more explicitly, along the lines of “holy shit the world is so malleable and open, let’s gooooo”. (Which obviously is a somewhat risky line of thinking, and I think people have mental safeguards against it for some good reasons which I’m tracking, and some which I’m not. Eppur lo si può muovere!)
There is no getting around the fact that AI will be built by people who build AI. Therefore, the people who build AI should care about X-risk.
As I mention in the post, the object-level benefits/harms of more time are of secondary importance in my mind. My main question is: how do we figure out which people “care about x-risk” in a way that reliably shapes their decisions? Because historically the community has taken people saying the words “I care about x-risk” as far too much evidence that this consideration actually steers their decisions.
Re the rest of your comment: it feels hard for me to pick out the underlying crux underneath all the considerations you raise. I do think there’s something that feels kinda demoralized about it—in particular, the sentence “you wouldn’t be able to dunk on the CEOs” conveys a sense that people who care about this stuff are by default gonna be throwing pebbles from the outside at the real decision-makers.
But part of what this sequence is trying to convey is that thinking clearly is so extraordinarily powerful that the people who are able to do enough of it to get on board with alignment are able to produce orders of magnitude more capabilities progress than equally-smart “mainstream” capabilities researchers even when that wasn’t their goal.
So I am generally extremely optimistic about this community’s ability to exert influence if it stops punching itself in the face.
(P.S. I appreciated your aside in your other comment, and it helped me orient to your main comment in a less defensive way.)
I agree with most of these claims. I think the main target audience of this sequence is people who are close to being able to do the thing John describes, and can be tipped either way by which ideas they’re exposed to—in particular whether they’re given a solid alternative to marginally boosting alignment (or even marginalist altruism more generally) as a model of how to be a rational and good person. I’m also targeting people who can do it, but only at the expense of significant psychological tradeoffs, which might be mitigated by having a better model of what’s going on.
My previous response didn’t really address the fact that I’m advocating for high-integrity strategies from a position where I have a bunch of prestige and money from working at OpenAI.
I do think this should make people more skeptical of both me personally and also the strategies I’m endorsing. In particular, it’s possible that I’m pointing at something which was directionally correct for my past self, but which can be overdone by others. However, on the meta level, stuff like “write very honest retrospectives” seems pretty robustly good.
When I think about what advice I’d give my past self, it does seem difficult to get him past his psychological bottlenecks without very direct exposure to the failures of highly prestigious institutions, plus a bunch of money. But as I alluded to in another comment, this isn’t a reliable way for people to fish themselves out of this mentality.
I think there’s a repertoire of hippie/therapeutic interventions which can manage this fairly reliably, and indeed played a big role for me. So I tend to point people towards them (sleepawake.camp is my strongest recommendation, but there’s all sorts of approaches which can help a bunch—e.g. this one-day workshop in Berkeley next month, circling, body work, psychedelic therapy, and so on). Note that things which are more physically/somatically oriented generally work much better than talk therapy. Doing some of these seems strongly correlated with retaining the ability to reliably update amongst alignment people (I don’t want to defend the epistemics of the wider hippie community, but they do have a bunch of metis about something very important).
Notably, I know a bunch of people who burned out of AI safety and left to pursue a more hippie lifestyle, without needing to first accumulate much money or prestige. I am pretty optimistic about most of those people later on ending up doing much more valuable things than they would have otherwise.
I overall feel better about the world where Daniel Kokotajlo worked at OpenAI for a bit (and probably also Richard although I’m less sure).
My sense is that almost all of the value of us working there came from us (and through us, the alignment community at large) becoming better at handling adversarial dynamics, from being forced to confront them directly.
However, I don’t think this is reliably good—I don’t think either of us planned for that going in, and my sense is that most alignment people at OpenAI became worse at handling such dynamics.
You shouldn’t give me much credit for leaving, btw, the main catalyst was Miles’ team dissolving (I could have stayed, but with much less research freedom). I think I should get more credit for leaving DeepMind in 2020 to do conceptual research at Cambridge, even though I had less money and prestige back then.
The main thing I was thinking is that the original 100x factor seems large enough that even a significantly reduced version of it would still establish the importance of RLHF. But now that I think about it, accounting for overfitting could plausibly bring this number a long way down.
Another thing that was in the back of my mind when I wrote this: my main claim in that paragraph is not that I’m confident about what would have happened in the absence of RLHF, but rather that “these arguments seem very suspect”. One of the main points of this post is that it’s extremely difficult to evaluate counterfactuals in the presence of motivated reasoning. What I can be confident about is that Paul published a paper showing that RLHF had enormous effects soon before ChatGPT came out, and then didn’t mention in his analysis of the effects of RLHF. That seems suspect (in a way that reduces my trust in his reasoning) whether or not the result holds up—if he thought the result didn’t hold up, he should have said so (and ideally explained how and why they released a misleading paper).
Part of why I focused so much on a few key individuals in my post, though, is because a few leaders doing this very well make it much easier for others to improve at this.
So the “most people” thing doesn’t seem relevant; I’m more directly trying to improve the peak than the mean.
On a quick skim, not that dissimilar. I don’t think I’m saying much that Yudkowsky didn’t have at least an intuitive grasp on. A big part of what I meant by “hammering home” was trying to create common knowledge through public statements (like Death with Dignity).
My main response is that I expect that (sub)communities which do assign credit in this way (e.g. assigning credit for steps not taken, or noticing people who turn down job offers) will be far more effective at achieving their goals in the long term. A big part of the point of this sequence is showing how the tradeoffs made to accrue money/power/prestige were often not worth it from the perspective of people actually trying to reduce x-risk. So it’s okay to stay smaller and exclude the people for whom sacrifices of (status/power/fame) are not really sustainable.
Even higher integrity more prescient version of Daniel would not have joined OpenAI in the first place. The problem is such version of Daniel has problem even getting noticed.
Maybe! I do think that there was a gaping hole in the community waiting for someone higher-integrity and more prescient than Daniel (or almost any of the rest of us) to start hammering home the points that I made in this post 5-10 years earlier. Also, LessWrong is still fairly meritocratic—it rewards (many kinds of) good writing no matter who it’s from (as Duncan’s Conor Moreton experiment showed IIRC).
But I think one piece of evidence for your position is that Ben Hoffman was basically this person, and indeed has not been noticed much by the wider community. (Though some of his collaborators—but not him—did receive a bunch of money from Jaan Tallinn in recognition of their contributions. (Note: edited upon confirmation it wasn’t him.))
Now, I could say that Ben has been noticed by a disproportionate number of the people who I respect most. But now we’re starting to talk about worryingly small numbers of people. On the other hand, this whole field exists because of the intellectual foundations laid by a very small number of people, and so if you expect that similar growth is still possible, then credit from those people matters a lot. (The founders of the new paradigm would just need to do a better job of not losing the funding and prestige to newcomers than Yudkowsky did.)
In my own case, I do have a sense that various rationalists were kinda wary of me while I was being more power-seeking, and that’s related to why I didn’t receive the kind of mentorship earlier which I currently have.
Meanwhile I am also trying to give ACS credit as one of the healthiest parts of the alignment ecosystem. However, it’s unclear how much that’s worth.
The reason I’m asking is because if we assume that you’re truthseeking, it’s kind of a weird thing to propose, right?
I like this line of inquiry, thank you. If I were working in a Bayesian epistemology I expect I’d find this argument fairly persuasive.
So I think the core reasons I disagree are related to my post on why I’m not a bayesian. That is: I think of intellectual progress in terms of growing a new ontology/developing a new model of the world. At first, this ontology will be fairly illegible to other people. Gradually, I’ll find ways to get it to generate novel falsifiable predictions, or figure out how it overlaps with other peoples’ ontologies. The more we can find overlap, the more productively we can communicate insights to each other. But the procedure of identifying the correspondence between your ontology and my ontology often requires a bunch of work. Insofar as we’re both doing it, we can “build the bridge from both ends” to find the overlapping sections and then communicate insights about those. This is much harder if one person is doing far more of the interpretative labor.
Treating epistemics as social exchange—trying to see things the other guy’s way on his request, in exchange for him trying to see it your way on your request—doesn’t create true maps; it creates false maps representing a compromise between the parties’ preferred lies.
I think your implicit model of “trying to see things the other guy’s way” is incorrect. In particular, it’s not “adopt one of your beliefs on request”, but more like “spin up a mental sandbox which explores the implications of your belief being true”.
And so the exchange that I’m actually proposing is more like “I will spend compute on trying to figure out ways that your perspective might be more consistent and well-intentioned than I currently expect, if you do the same for me”.
The main way that this might create false maps is if the sandboxes are leaky, or if your perspective has been adversarially optimized to lead me to false conclusions. When I talk about trusting someone, one of the things I mean is “trying to simulate their perspective won’t mess my epistemics up”. E.g. you shouldn’t trust or try to mentally explore the implications of claims that a misaligned superintelligence has given you.
the idea seems to be that if people’s attempts to formulate positive visions get critiqued too vigorously, they’ll get discouraged and give up.
The implied psychological model of Less Wrong authors reminds me of my attempts to teach chess to my five-year-old niece this week
I tried out several possible angles of response, but upon reflection, I’m no longer interested in debate with you on this topic (though I might still engage if you have responses to the rest of this comment).
For what it’s worth, I often support people giving in to their dark sides in low-stakes ways, including in this case.
How did it feel?
This seems to be ignoring my point about why I think this is a conflict theoretic situation.
You mean that “it was initially set on this path when Eliezer demanded author mod powers”? You could interpret this as Eliezer wielding power in ways you don’t like, but you could also interpret it as the mods treating Eliezer’s preferences as evidence about what site norms produce intellectual progress (along with many other people’s preferences) and acting accordingly. I assume you don’t think the latter is a good model; if so, why not?
Because you don’t have nearly as much “sunk cost” invested in the current culture (e.g. was not responsible for cultivating it in the first place), I infer that you can change your mind much more easily about what kind of site culture is more conducive for clear thinking or intellectual progress.
Two responses. Firstly, I acknowledge that you did a lot to cultivate the site culture in the first place, and I’m grateful for that.
Secondly: one of the most difficult parts of rationality seems to be changing one’s mind in the face of sunk costs. Because of that, I agree that it’s easier for me to change my mind on this than it is for you. But do you endorse the extent to which sunk costs are making it harder for you to change your mind?
I take the side of the site mods on this issue, though, for the reasons articulated in my comment above.
I would also bid for you to do the thing that I recommended Zack do above, namely “acknowledging the thing that Habyrka and I are trying to protect, and helping us figure out how to protect it with as few tradeoffs as possible”.
I basically endorse Tsvi’s comment above. It does also seem reasonable that Wei is concerned that this post will continue to be used to criticize him (though this doesn’t seem like a good reason to take down the post). In hindsight, I regret the parenthetical where I added “(a concept which was originally inspired by Wei making a similar comment on a post by Tsvi)”, because I don’t want to associate the concept too strongly with Wei—I think it’s a generically useful concept.
For the record I have not moderated Wei’s comments in any way, and don’t intend to do so.
I would also like Wei to continue writing on LessWrong, though I think my original reply was well within the bounds of reasonable discourse norms, and therefore do not take responsibility for him leaving if he does.
Agree with Tom. To elaborate a bit more: the reason this post focused so much on the idea that the alignment/capabilities distinction has lost its meaning is because that distinction was implicitly or explicitly grounding a lot of other arguments about impact.
I’ll discuss this more in the next post (partly because the comments on this one have been helpful in clarifying some of the concepts I’ve missed so far). But as one brief intuition: the key thing I’m worried about is circular justification loops, where ambiguities in what we mean by “alignment research” allow people to justify all sorts of stuff that never cashes out in the real thing we want. As some toy examples:
Or
I’m not saying that any individual person has or would endorse these circular arguments, but I am saying that the community as a whole has ended up being unable to distinguish them from non-circular arguments.
How might you attempt to prevent this? For example, you could divide work into type 1: directly aimed at solving the hard parts of the alignment problem; type 2: aims to produce more type 1 research; type 3: aims to produce more type 2 work; and so on.
This is obviously a kinda blunt way to classify things, though it’s better than anything we actually had. A very stylized retelling of my post is that:
MIRI and Paul-before-OpenAI (e.g. iterated amplification theory) were doing type 1 research.
Paul at OpenAI claimed to be doing type 2 research, and also debunked a bunch of MIRI’s worldview, and so people started interpreting Paul’s type 2 research as a central example of alignment research.
Paul’s arguments for why doing his type 2 research was a good idea were pretty dodgy but he didn’t write them up at the time so it was hard to publicly critique them.
Dario was claiming to be doing type 3 work (by scaling up the models, which would help type 2 work like Paul’s), and nobody publicly called out how dodgy his arguments were, so people started interpreting Dario as a central example of a person who cares about alignment, which then helped him build Anthropic.
And so now we have this whole edifice of people doing stuff that’s “good for alignment” via these very tenuous chains which they have a lot of reason to do motivated cognition about. I think Paul’s retrospective on why RLHF was a good idea was very far from his usual epistemic standards—for example, the original version literally didn’t even mention ChatGPT! And so, while I know much less than Paul about the details of the situation, I can’t trust him to actually figure out if his reasoning was wrong, and I need to model him as basically just defending his past self.
Also, unfortunately, almost everyone in the field is too conflict-averse to characterize others as doing motivated reasoning, and so we’re stuck in the middle of these insane epistemic distortions without any way to talk about them. So, again, my main contention is not “the current paradigm has led to a discrediting amount of harm” but more like “the current paradigm has led to a discrediting amount of rationalizations for doing things which drove an enormous proportion of total capabilities work, from people who are demonstrating nowhere near the level of integrity necessary for us to trust that they’re reasoning clearly about the counterfactuals, and (relatedly) seem totally incapable of calling out even the people most obviously doing deceptive or motivated reasoning under the banner of alignment”.