Formerly alignment and governance researcher at DeepMind and OpenAI. Now independent.
Richard_Ngo
I think something that could’ve been clearer in the last section is that dramatically changing up your life plan is not the only possible immediate next step if someone feels moved by this post—one could also start marginally increasing the robustness of one’s strategy while staying in a roughly similar position in life; e.g. stay at a frontier lab but become more like Leo Gao, increase your everyday integrity and psychological health, etc.
This is what I intended with the section starting “For now, I’ll focus on a few high-level principles for starting to move in that direction.”
However, I did wonder as I was writing it if even that was making too much of a concession—because I do think that impact here is heavy-tailed towards people who are able and willing to carve out unusual paths. (FWIW I get some impression from our online interactions that you could end up as one of them eventually.)
Relatedly, I’d nudge you to think more about what you meant with the phrase “the only possible immediate next step”—seems like there’s a bunch of implicit social stuff underlying it.
For my part, I think there’s some implicit message underlying a lot of my writing that I should try to convey more explicitly, along the lines of “holy shit the world is so malleable and open, let’s gooooo”. (Which obviously is a somewhat risky line of thinking, and I think people have mental safeguards against it for some good reasons which I’m tracking, and some which I’m not. Eppur si può muovere!)
There is no getting around the fact that AI will be built by people who build AI. Therefore, the people who build AI should care about X-risk.
As I mention in the post, the object-level benefits/harms of more time are of secondary importance in my mind. My main question is: how do we figure out which people “care about x-risk” in a way that reliably shapes their decisions? Because historically the community has taken people saying the words “I care about x-risk” as far too much evidence that this consideration actually steers their decisions.
Re the rest of your comment: it feels hard for me to pick out the underlying crux underneath all the considerations you raise. I do think there’s something that feels kinda demoralized about it—in particular, the sentence “you wouldn’t be able to dunk on the CEOs” conveys a sense that people who care about this stuff are by default gonna be throwing pebbles from the outside at the real decision-makers.
But part of what this sequence is trying to convey is that thinking clearly is so extraordinarily powerful that the people who are able to do enough of it to get on board with alignment are able to produce orders of magnitude more capabilities progress than equally-smart “mainstream” capabilities researchers even when that wasn’t their goal.
So I am generally extremely optimistic about this community’s ability to exert influence if it stops punching itself in the face.
(P.S. I appreciated your aside in your other comment, and it helped me orient to your main comment in a less defensive way.)
I agree with most of these claims. I think the main target audience of this sequence is people who are close to being able to do the thing John describes, and can be tipped either way by which ideas they’re exposed to—in particular whether they’re given a solid alternative to marginally boosting alignment (or even marginalist altruism more generally) as a model of how to be a rational and good person. I’m also targeting people who can do it, but only at the expense of significant psychological tradeoffs, which might be mitigated by having a better model of what’s going on.
My previous response didn’t really address the fact that I’m advocating for high-integrity strategies from a position where I have a bunch of prestige and money from working at OpenAI.
I do think this should make people more skeptical of both me personally and also the strategies I’m endorsing. In particular, it’s possible that I’m pointing at something which was directionally correct for my past self, but which can be overdone by others. However, on the meta level, stuff like “write very honest retrospectives” seems pretty robustly good.
When I think about what advice I’d give my past self, it does seem difficult to get him past his psychological bottlenecks without very direct exposure to the failures of highly prestigious institutions, plus a bunch of money. But as I alluded to in another comment, this isn’t a reliable way for people to fish themselves out of this mentality.
I think there’s a repertoire of hippie/therapeutic interventions which can manage this fairly reliably, and indeed played a big role for me. So I tend to point people towards them (sleepawake.camp is my strongest recommendation, but there’s all sorts of approaches which can help a bunch—e.g. this one-day workshop in Berkeley next month, circling, body work, psychedelic therapy, and so on). Note that things which are more physically/somatically oriented generally work much better than talk therapy. Doing some of these seems strongly correlated with retaining the ability to reliably update amongst alignment people (I don’t want to defend the epistemics of the wider hippie community, but they do have a bunch of metis about something very important).
Notably, I know a bunch of people who burned out of AI safety and left to pursue a more hippie lifestyle, without needing to first accumulate much money or prestige. I am pretty optimistic about most of those people later on ending up doing much more valuable things than they would have otherwise.
I overall feel better about the world where Daniel Kokotajlo worked at OpenAI for a bit (and probably also Richard although I’m less sure).
My sense is that almost all of the value of us working there came from us (and through us, the alignment community at large) becoming better at handling adversarial dynamics, from being forced to confront them directly.
However, I don’t think this is reliably good—I don’t think either of us planned for that going in, and my sense is that most alignment people at OpenAI became worse at handling such dynamics.
You shouldn’t give me much credit for leaving, btw, the main catalyst was Miles’ team dissolving (I could have stayed, but with much less research freedom). I think I should get more credit for leaving DeepMind in 2020 to do conceptual research at Cambridge, even though I had less money and prestige back then.
The main thing I was thinking is that the original 100x factor seems large enough that even a significantly reduced version of it would still establish the importance of RLHF. But now that I think about it, accounting for overfitting could plausibly bring this number a long way down.
Another thing that was in the back of my mind when I wrote this: my main claim in that paragraph is not that I’m confident about what would have happened in the absence of RLHF, but rather that “these arguments seem very suspect”. One of the main points of this post is that it’s extremely difficult to evaluate counterfactuals in the presence of motivated reasoning. What I can be confident about is that Paul published a paper showing that RLHF had enormous effects soon before ChatGPT came out, and then didn’t mention in his analysis of the effects of RLHF. That seems suspect (in a way that reduces my trust in his reasoning) whether or not the result holds up—if he thought the result didn’t hold up, he should have said so (and ideally explained how and why they released a misleading paper).
Part of why I focused so much on a few key individuals in my post, though, is because a few leaders doing this very well make it much easier for others to improve at this.
So the “most people” thing doesn’t seem relevant; I’m more directly trying to improve the peak than the mean.
On a quick skim, not that dissimilar. I don’t think I’m saying much that Yudkowsky didn’t have at least an intuitive grasp on. A big part of what I meant by “hammering home” was trying to create common knowledge through public statements (like Death with Dignity).
My main response is that I expect that (sub)communities which do assign credit in this way (e.g. assigning credit for steps not taken, or noticing people who turn down job offers) will be far more effective at achieving their goals in the long term. A big part of the point of this sequence is showing how the tradeoffs made to accrue money/power/prestige were often not worth it from the perspective of people actually trying to reduce x-risk. So it’s okay to stay smaller and exclude the people for whom sacrifices of (status/power/fame) are not really sustainable.
Even higher integrity more prescient version of Daniel would not have joined OpenAI in the first place. The problem is such version of Daniel has problem even getting noticed.
Maybe! I do think that there was a gaping hole in the community waiting for someone higher-integrity and more prescient than Daniel (or almost any of the rest of us) to start hammering home the points that I made in this post 5-10 years earlier. Also, LessWrong is still fairly meritocratic—it rewards (many kinds of) good writing no matter who it’s from (as Duncan’s Conor Moreton experiment showed IIRC).
But I think one piece of evidence for your position is that Ben Hoffman was basically this person, and indeed has not been noticed much by the wider community. (Though some of his collaborators—but not him—did receive a bunch of money from Jaan Tallinn in recognition of their contributions. (Note: edited upon confirmation it wasn’t him.))
Now, I could say that Ben has been noticed by a disproportionate number of the people who I respect most. But now we’re starting to talk about worryingly small numbers of people. On the other hand, this whole field exists because of the intellectual foundations laid by a very small number of people, and so if you expect that similar growth is still possible, then credit from those people matters a lot. (The founders of the new paradigm would just need to do a better job of not losing the funding and prestige to newcomers than Yudkowsky did.)
In my own case, I do have a sense that various rationalists were kinda wary of me while I was being more power-seeking, and that’s related to why I didn’t receive the kind of mentorship earlier which I currently have.
Meanwhile I am also trying to give ACS credit as one of the healthiest parts of the alignment ecosystem. However, it’s unclear how much that’s worth.
What just happened? Pragmatism and Pessimization
The reason I’m asking is because if we assume that you’re truthseeking, it’s kind of a weird thing to propose, right?
I like this line of inquiry, thank you. If I were working in a Bayesian epistemology I expect I’d find this argument fairly persuasive.
So I think the core reasons I disagree are related to my post on why I’m not a bayesian. That is: I think of intellectual progress in terms of growing a new ontology/developing a new model of the world. At first, this ontology will be fairly illegible to other people. Gradually, I’ll find ways to get it to generate novel falsifiable predictions, or figure out how it overlaps with other peoples’ ontologies. The more we can find overlap, the more productively we can communicate insights to each other. But the procedure of identifying the correspondence between your ontology and my ontology often requires a bunch of work. Insofar as we’re both doing it, we can “build the bridge from both ends” to find the overlapping sections and then communicate insights about those. This is much harder if one person is doing far more of the interpretative labor.
Treating epistemics as social exchange—trying to see things the other guy’s way on his request, in exchange for him trying to see it your way on your request—doesn’t create true maps; it creates false maps representing a compromise between the parties’ preferred lies.
I think your implicit model of “trying to see things the other guy’s way” is incorrect. In particular, it’s not “adopt one of your beliefs on request”, but more like “spin up a mental sandbox which explores the implications of your belief being true”.
And so the exchange that I’m actually proposing is more like “I will spend compute on trying to figure out ways that your perspective might be more consistent and well-intentioned than I currently expect, if you do the same for me”.
The main way that this might create false maps is if the sandboxes are leaky, or if your perspective has been adversarially optimized to lead me to false conclusions. When I talk about trusting someone, one of the things I mean is “trying to simulate their perspective won’t mess my epistemics up”. E.g. you shouldn’t trust or try to mentally explore the implications of claims that a misaligned superintelligence has given you.
the idea seems to be that if people’s attempts to formulate positive visions get critiqued too vigorously, they’ll get discouraged and give up.
The implied psychological model of Less Wrong authors reminds me of my attempts to teach chess to my five-year-old niece this week
I tried out several possible angles of response, but upon reflection, I’m no longer interested in debate with you on this topic (though I might still engage if you have responses to the rest of this comment).
For what it’s worth, I often support people giving in to their dark sides in low-stakes ways, including in this case.
How did it feel?
This seems to be ignoring my point about why I think this is a conflict theoretic situation.
You mean that “it was initially set on this path when Eliezer demanded author mod powers”? You could interpret this as Eliezer wielding power in ways you don’t like, but you could also interpret it as the mods treating Eliezer’s preferences as evidence about what site norms produce intellectual progress (along with many other people’s preferences) and acting accordingly. I assume you don’t think the latter is a good model; if so, why not?
Because you don’t have nearly as much “sunk cost” invested in the current culture (e.g. was not responsible for cultivating it in the first place), I infer that you can change your mind much more easily about what kind of site culture is more conducive for clear thinking or intellectual progress.
Two responses. Firstly, I acknowledge that you did a lot to cultivate the site culture in the first place, and I’m grateful for that.
Secondly: one of the most difficult parts of rationality seems to be changing one’s mind in the face of sunk costs. Because of that, I agree that it’s easier for me to change my mind on this than it is for you. But do you endorse the extent to which sunk costs are making it harder for you to change your mind?
I take the side of the site mods on this issue, though, for the reasons articulated in my comment above.
I would also bid for you to do the thing that I recommended Zack do above, namely “acknowledging the thing that Habyrka and I are trying to protect, and helping us figure out how to protect it with as few tradeoffs as possible”.
I basically endorse Tsvi’s comment above. It does also seem reasonable that Wei is concerned that this post will continue to be used to criticize him (though this doesn’t seem like a good reason to take down the post). In hindsight, I regret the parenthetical where I added “(a concept which was originally inspired by Wei making a similar comment on a post by Tsvi)”, because I don’t want to associate the concept too strongly with Wei—I think it’s a generically useful concept.
For the record I have not moderated Wei’s comments in any way, and don’t intend to do so.
I would also like Wei to continue writing on LessWrong, though I think my original reply was well within the bounds of reasonable discourse norms, and therefore do not take responsibility for him leaving if he does.
I think the “conflict” frame here is useful but can also be over-applied; let me jot down some thoughts in response.
One kind of conflict that seems to arise is between people who are trying to diagnose where and how things went wrong, and the people who don’t want that to happen. This conflict is unfortunately kinda fractal. So for instance, many rationalists are trying to diagnose what went wrong with EAs and “AI safety”, while mainstream AI safety is fairly resistant to this. And then Wei Dai and Vassar and Hoffman and you and me (and Habryka to some extent) are trying to diagnose what went wrong with rationalism, while many rationalists are fairly resistant to this.
Unfortunately, within the group of people trying to diagnose what went wrong with rationalism, we end up often categorizing each other as part of the problem (because people disagree about what the causes are, and because you need to be an extremely disagreeable person to end up in that category, Wei perhaps excepted). This might be reading too much into Wei’s original comment, but I get some implicit message from it like “you’re trying to diagnose what went wrong but you are contributing to the problem by valorizing Yudkowsky’s intellectual clarity”. There are also ways in which Vassar and I each think the other is contributing to the problem (though we have a productive collaboration regardless).
One axis of disagreement about the diagnosis of things going wrong seems to be: how valuable it is to have a positive vision, versus to be able to critique flaws. Part of what frustrated me about Wei’s original comment is that I didn’t claim that rationalists were amazing at philosophy by any given objective standard—indeed, I’ve been one of the main people arguing that the whole foundation of rationalism on bayesianism was a massive mistake. Rather, I intended the phrase “beacon of intellectual clarity” to convey that it was unusually good by the standards of the broader landscape. I also didn’t claim that this intellectual clarity translated well into strategic acumen, and indeed criticized several strategic choices like MIRI’s engagement with prestigious elites.
So I basically want to say: I see that there’s something here you want to protect, and I’ve come to appreciate it much more over time. There’s also something that I (and, I infer, Habryka) want to protect—as implicitly expressed in my post it was something like “intellectual progress is in fact a valuable thing which we can aim for”. Insofar as we’re in a conflict frame with each other, you could view Wei’s original comment as an attack on that thing. But I tried to deescalate in my recent comment, because this doesn’t seem like a useful conflict to be in. I am very on board with identifying Yudkowsky’s strategic mistakes, and indeed my draft sequence on the strategic mistakes made by the major people in the field (including Yudkowsky) is now almost 20,000 words long. However, I would like that to trade off as little as possible against a shared sense that intellectual progress is possible. On my end, that involves reorienting how I interpret Wei’s comments. On your end, it would ideally involve acknowledging the thing that Habyrka and I are trying to protect, and helping us figure out how to protect it with as few tradeoffs as possible.
In case it’s useful:
I have no intention of banning you from commenting on my posts, and indeed hereby commit not to doing so.
My original comment which prompted this came from an emotional stance of frustration, but didn’t actually articulate the underlying generators of that frustration, which then led it to come out in a more accusatory way than it could have. Let me try below to be more direct and sincere.
Where am I at? I guess I have this underlying stance of being really very viscerally excited about my research directions (as we discussed in our call recently). And because of that, I find myself casting around for people to help push them forward, or at least to share that excitement.
You, Wei, are one of the people I most respect intellectually; as I referred to in that one twitter thread, you’re probably in the top dozen people whose opinions I’d be most deferential to if they disagreed with me about some high-level strategic decision. When I find myself unable to bridge the gap to your worldview, this feels on a visceral level like an update towards “maybe there’s almost nobody I can bridge that gap with”. And to be clear, I think that this can be viewed as a skill issue on my end; the feeling I’m tracking is something like “it’s sad that it seems like I need to be so much more skillful in order to get people anywhere near the same page as me”. I intend to express that sentiment more as sadness and less as frustration going forward.
Separately, over the last year or two I’ve been greatly appreciating a bunch of conversations with people who’ve been around this scene from the 2000s. So I do think there’s an alternative stance I could take that’s more like “the old guard has certain extremely rare skills and accumulated wisdom, which are a kind of resource I can draw on”. I did find your comment on my other post helpful; and while I wish it had been more helpful, there’s also a mindset I can adopt where it’s actually much more helpful than I am currently able to recognize. (I’m reminded of an exchange I had with Anna Salamon where I said something like “Anna can you just give me your accumulated wisdom directly” and she responded something like “as you can see from all the fables about people seeking wisdom, that’s not how wisdom works”.)
So maybe another mindset I can have towards your comments is “Wei is one of the only people in the world who was able to both take Eliezer’s out-there ideas seriously while also maintaining an appropriate level of skepticism about MIRI’s strategy (and also not giving up on engaging with the rationalist community); when he talks about the difficulty of philosophy, he is conveying that almost-unique skillset, which I can learn a lot from”. From that perspective, I do feel a sense of gratitude for your engagement (alongside the sadness I previously mentioned). I can’t commit to maintaining a perspective like this in future interactions, but I intend to try and see where it goes.
Is there any writeup you can point me to of the argument you want to defend? I’d like to engage with the best version of it, but I’m worried that it’s amorphous enough that different people will end up defending different versions at different times.