Formerly alignment and governance researcher at DeepMind and OpenAI. Now independent.
Richard_Ngo
I’ve commented on this announcement on twitter and in the Constellation slack (from which I’ve resigned my membership, as per below). To quickly summarize: the AI safety community is structurally blocked from reliably improving the world unless it becomes much better at handling adversarial dynamics. In particular, the most important lesson we can learn from the last decade is how and why “AI safety” was so heavily captured by AGI companies, and what it would have taken to prevent that. The most basic building block is the willingness to call a spade a spade, and honestly discuss what’s going on with AGI companies; I interpret Paul as committing via his association with OpenAI not to do so. Of course Paul is only one part of Constellation, but his worldview has been (and I expect will continue to be) the single biggest influence on the broader strategy that Constellation is pursuing, which suggests that most reasonable thing to do is to disaffiliate.
The text of my tweet:
I am sad and disappointed to hear that Paul is joining OpenAI’s board. Being affiliated with OpenAI has historically led AI safety researchers (including both Paul and myself) to behave in low-integrity ways. I personally was drawn to OpenAI in part by the idea that I could make a difference to the future of AI. However, once there, many of my actions were governed by fear of getting on the wrong side of OpenAI execs. I often found myself making excuses for behavior that clearly contradicted OpenAI’s own stated goal of making AGI go well for humanity. I was far from alone in this—e.g. when the board tried to fire Sam over his deceptive behavior, several senior safety researchers became scared of losing their influence, and so pushed hard to bring him back. Meanwhile, many people kept OpenAI’s misbehavior secret for fear of non-disparagement agreements. (More on all of this in an upcoming retrospective.)
I can’t speak directly for Paul’s motivations. However, his previous work at OpenAI contributed significantly both to their biggest capability advances, and to the capture of AI safety by AGI companies over the last decade, as I recount at length in the blog post linked below. One key factor was the unwillingness of (almost) the entire AI safety community to say things which might offend OpenAI execs. For example, I have not been able to find a single comment critical of OpenAI from Paul during his original tenure there (when he was writing prolifically on AI safety and strategy).
Unfortunately, Paul doesn’t seem to have become significantly more willing to directly and honestly criticize people who he believes are behaving in morally abhorrent ways—see the bland corporate-speak of his statement below. While he speaks directly about the possibility of humanity losing control of the world to AI, he expresses only excitement about OpenAI itself, despite OpenAI being one of the main sources of such risk.
Paul’s announcement comes only weeks after OpenAI models autonomously launched a cyberattack on HuggingFace, and only days after we learned that OpenAI hid details of previous breakouts from the external investigators. It is irresponsible for leaders of the AI safety community—whose judgements many people are relying on—to put themselves in positions which will significantly bias their ability to discuss such incidents. Unless Paul makes strong commitments to openness and honesty (and demonstrates willingness to potentially be fired for that honesty), I expect that the main effect of him joining OpenAI’s board will be to help OpenAI defuse external criticism and further “safety-wash” itself.
I want to note that Paul is a brilliant researcher and a prescient forecaster. Because of that, he’s the closest thing there is to a leader of what I’ll call the “pragmatic AI safety” cluster—which includes the organizations working out of the Constellation offices (like Redwood Research, METR, and Paul’s Alignment Research Center), as well as many people scattered across AGI companies, Coefficient Giving, etc.
External observers are often confused about why so many people are working at AGI companies while professing to believe that those same companies have a double-digit probability of permanently disempowering humanity. In large part, it’s because people in the pragmatic AI safety cluster have failed to follow high-integrity strategies for reducing AI risk, in favor of clever arguments about the benefits of being proximate to power. I am not singling Paul out as less ethical than other prominent figures in this cluster, who are also very conflict-averse in their orientation to AGI companies. However, it is well past time for everyone involved to change course. As one (relatively small) step, I’m therefore resigning my membership of the Constellation offices.
I hope that, going forward, the people who are trying to steer the future of AI prioritize building much more solid foundations of courage and honesty than we currently have.
(Next tweet): Here’s the post I mentioned, which recounts the history of AI safety getting captured by OpenAI and (later) Anthropic: https://www.lesswrong.com/posts/yaz8nx4ogZmiqHzt7/what-just-happened-pragmatism-and-pessimization
Follow-up posts in the series will focus more specifically on the power dynamics at OpenAI when I was there from 2021 to 2024.
An earlier draft of my tweet above started with a condemnation of Paul joining OpenAI. But I took it out because I think it’s important to distinguish gradations of moral responsibility. Paul is still one of the few people best able to orient to the possibility of superintelligence, and is still far more morally serious than any of the capabilities researchers or executives at OpenAI or Anthropic. I think he is badly mistaken about strategic considerations, and that the community should correspondingly discount his judgement. But I continue to hope that he and other thinkers in the pragmatic AI safety cluster will reorient.
The problem here is that creating common knowledge that people are following a given set of norms (and therefore should be treated as coalition members in good standing) becomes far more complicated the more complicated the strategy is.
Leave the company: very simple, easily verifiable, everyone knows (that everyone knows) what happened.
Staying but not working, and doing internal agitation, and donating one’s salary: complicated plan, many parts are hard to externally verify, easy to rationalize.
In particular, I claim that most of the people working on “alignment” at Anthropic and OpenAI are significantly deceiving themselves about the impact of their work. Your plan has no safeguard against people doing that. We could try to build stronger norms of what counts as “real alignment work”, but nobody inside the labs should be trusted to do that, because that’s how we lost the concept of real alignment work in the first place.
Just one quick note (might respond to the rest later): re “striking”, the crucial difference here is that unlike strikers, people who declare that they won’t work will continue (by default) to be paid a massive salary.
That undermines a lot of their credibility.
Conflict-aversion doesn’t work by minimizing total conflict, it instead works by locally moving away from current conflict even when that might increase long-term conflict.
Think about e.g. someone not telling their partner that they forgot to do something they promised to do, because it feels aversive (even though not telling them will definitely lead to a bigger problem later down the line).
This leads me to predict that if someone actually did the “refusing to work” thing in any loud or disruptive way, AI safety people would also complain about that.
Separately, I think refusing to work is more conflict-y than just quitting, but less conflict-y than quitting with a strong public statement. It also provides many more opportunities for the conflict to just fade into the background (e.g. maybe someone decides not to work, the company doesn’t fire them, and then things carry on as usual).
There’s a whole meme about how Googlers for example can just do a few hours of work a week and it’s fine.
The writing was pleasant to read notwithstanding a non-zero LLM score (we’re wrestling with LLM-assisted writing on LW, but felt quite good to read).
I noticed the LLM-ness as well; when I did, I retracted my upvote. I think the norm I want is disclosure of how AI has been used to write the post. One sentence is fine, but the lack of that puts me in an adversarial dynamic with the writer.
I am somewhat dismayed to see how highly-upvoted this post is; I think I have literally never seen a post this poorly-reasoned get 200 karma (before my strong downvote). It reads like a grab-bag of arguments hastily thrown at a strong pre-written bottom line.
To be clear, I’m not saying that Kabir shouldn’t have posted this, I generally am very pro people quickly writing up their takes (though I do think the title is unusually bad). Rather, I’m mainly trying to figure out why the community reaction was so positive.
I might engage more on the object level later but for now my default hypothesis is that people are seeing AI safety people get in public loud conflict with the labs by quitting, are having a knee-jerk conflict-averse reaction, and instinctively supporting anything that might defuse that immediate conflict.
In general we see this pattern a lot with the labs: people do something that is somewhat critical or disagreeable, and then everyone quickly generates a morass of FUD and reasons why, yes, they support people taking action in principle, it’s just that this particular action is a terrible idea that will lead to so much backlash compared with all the other possible actions. In this case, there’s not even the usual “this would cause backlash” step, the post is saying “Dear God, please don’t take [action that I haven’t even argued is bad]”, which at the very least is missing an elephant-sized mood. And in exchange, it’s proposing a strategy which is extremely difficult for outsiders (and maybe even others inside the company) to even notice is happening.
I wouldn’t have had this objection if the post had been written more like “here’s another good alternative to consider (though admittedly I have no examples of it working in the past)”. I am open to the idea that something like Kabir’s proposed strategy is a good idea (though I currently think it’s a pretty bad strategy). I’m also open to the idea that my meta-level analysis is wrong. But there’s definitely something weird going on here that needs meta-level explanation.
reasoning about counterfactuals is hard, and rationalization is probably not rare, but I don’t see a great alternative
My current model of the alignment community is that it has ossified in a similar way as the ML establishment in the 2010s (or the AGI companies in the 2020s).
In general when you talk to someone in those ossified communities about AGI risk, even if you manage to get them to concede on each individual point, the fallback response you’ll get is something like “yeah AGI could be dangerous, and we don’t seem to be on track to solve alignment, but I don’t see what I can do to help much”, and then they go back to working on whatever they were previously doing.
In this case, I am talking to alignment community about how the strategy it has been using has driven a big chunk of capabilities progress, while producing little meaningful alignment progress. And so my response is similar to what I say to capabilities researchers: it is your job to figure out an alternative, or else to go sit on a beach somewhere, rather than continuing to contribute to unhealthy ecosystems that can’t reliably steer towards good things rather than bad things.
Higher education as class commitment
Insofar as they’re calling for a capabilities race (I assume you’re referring to the Situational Awareness memo) my read is that a lot of the reasoning behind it was “this will happen anyway so we may as well make sure it’s a capabilities race with marginally more safety characteristics”. In other words, I think Leopold told himself that calling for a capabilities race was good for safety partly because it was a way of getting safety memes into the ears of more important people. (Also, the consideration “it would be better if the US built AGI than if China built AGI” was considered continuous with “safety motivations” at the time—e.g. I remember people (though I don’t remember who) arguing for this by appealing to the fact that the alignment community had more influence in the US, so US AGI was more likely to be aligned.)
This is based in part on my read of the text, and in part on personally knowing Leopold. I also chatted with him about my critique of the Situational Awareness memo specifically—unfortunately I don’t remember well enough what he said to reliably report back, but in my recollection it was consistent with this read.
Insofar as this happened I think it was bad and self-deceptive reasoning, of course, but the important point is that it’s a similar kind of bad and self-deceptive reasoning that I’m calling out in the rest of this post.
See Wei Dai’s comment here for a rebuttal to that point. In particular, Carl Shulman is one of the people who did the most to lay the ideological foundations for early effective altruism, and is still very respected by leading figures in EA. And when I overlapped with him at OpenAI, Leopold was very active in safety advocacy, and worked on the superalignment team. Some of this was admirable work (and similar to what I would have been doing if I’d been less conflict-averse at the time), so I’m not critiquing that, just noting that he was squarely in the safety camp during his time at OpenAI (and also previously at FTX).
In short, I think that Situational Awareness is not just “safety people” but very central examples of safety people, especially in terms of the social graph around Holden, Dario (whose chief of staff is married to Leopold), Constellation, etc.
Yepp, I like that post a lot. And for the record I’ll also point people to the comment I left linking that post to Knightian uncertainty.
The main thing I’ll add is that people on LW have historically privileged a paranoid viewpoint towards smarter agents. Whereas thinking about the what it would look like for smarter agents to be “on your side” seems necessary for e.g. figuring out good alignment targets, or tracking progress on alignment.
“scale-free agency” and “asymptotic agent foundations”
I like this distinction, thanks.
One important question in my mind is whether you’re taking the asymptote as a single agent gain more and more compute, versus an asymptote as multiple agents simultaneously gain more and more compute.
Most of what you call asymptotic agent foundations does the former. However, since other agents are the most interesting part of most environments, this seems badly-motivated. (One potentially-checkable crux: does your claim, that malevolent superintelligences can’t predictably make logical inductors miscalibrated in the limit, hold up for natural operationalizations of the superintelligences also becoming more intelligent in the limit?)
Whereas latter might preserve many dynamics between agents (since their relative computational power isn’t changing much). It’s these dynamics which I’m focusing on when looking for a theory of scale-free agency.
I do also get some sense that we disagree about how much to update on the existence of optimality results. It would be nice if FixDT had such results, but it doesn’t feel crucial to me. Maybe this is because I’ve entered the field relatively recently, and therefore haven’t had as much experience with nice-sounding theories that can’t be pinned down. On that note, since I’m definitely not one of the traditional agent foundations researchers with a strong track record in the field, I do think that people shouldn’t defer much to my intuitions (though I do think they should take my arguments about scientific precedent quite seriously).
I’ll need to think about most of this for longer (and/or discuss in person). One quick comment in the interim:
I think the converse image of your Sierpinski triangle will better describe the true theory—Knightian all the way down, with various degrees of Bayesian-ness in regions where we are willing to spend lots of compute.
This is an excellent point. Though note that it might just be a matter of flipping your point of view: from inside one of those holes looking out, everything might be Knightian by default with small regions of Bayesian-ness. (More correlated with “simple domains” than “domains where we’re willing to spend a lot of compute”.)
I was just reminded of this meme I made a few months ago, which didn’t merit a place in the post itself, but which I wanted to preserve for posterity:
To what extent are you endorsing the reasons they gave in this post, versus reporting knowledge of other legit unspecified reasons?
From my perspective, the reasons given in this post are bad enough to cast doubt on any other reasoning that PauseAI is doing. In particular:
Attacking people’s character is not aligned with our fundamental commitment to nonviolence, neither in principle, nor in the outcomes it generates.
Characterizing attacks on people’s character as a kind of violence is absurd—the ability to critique and evaluate people’s characters is one of the foundational mechanisms which makes ethical coordination possible.
Not only is this a bad mistake in general, it is specifically one of the key mistakes that has led to the AI safety movement becoming ideologically captured, as I recount in this post.
So it seems important for you (Eliezer) to clarify the extent to which you endorse this blanket prohibition on character-based critiques.
But we think the AI safety community is beginning to change its mind and will put its weight behind a pause if we show them this is the better, viable path. We believe that making it easy for people with different views to come round is of immense strategic importance. PauseAI US’ leadership chooses to alienate people with different views.
None of these sentences is outright wrong, but overall this is a kind of reasoning that’s extremely susceptible to conflict-avoidance. In particular, caring a lot about how much you alienate people with other views makes you very adversarially non-robust, because it’s easy to reason yourself into bending over backwards to be conciliatory. Again, this a mistake that the AI safety community has made constantly with respect to labs; now it seems like PauseAI is making the same mistake with respect to the AI safety community.
It does seem plausible that Holly has over-corrected too far away from these mistakes, but in my mind this is pretty understandable, and far better than just continuing to walk into the same rakes as everyone else.
(Edit: someone on twitter reminded me about the whole “interpreting Seb Krier’s memes as a call for violence” saga, which was fairly unhinged from Holly and allies; I would have written the last sentence less sympathetically if I’d remembered that.)
Thanks for the comment.
There’s something a little strange about how you’re using the phrase “bounded rationality”. One implication is that you’re talking about something which isn’t true of superintelligences. But even a superintelligence will also be a bounded agent, and it’s hard to predict that it won’t get any use out of, say, 12-dimensional planet-sized Tarot cards.
So instead I interpret your use of the term as: there’s stuff which can be described using clean theories (of the kind which scale to superintelligence) which you’re calling agent foundations, and stuff which is more messy and contingent, e.g. because it depends on “specific cognitive limitations”, which you’re calling bounded rationality.
One thing I’m trying to convey in this post is that I think there’s an elegant theory of boundedly rational agents which does apply across many scales, such that many things we currently think of as human-specific phenomena will actually turn out to be facets of this deeper structure. In which case the distinction between “bounded rationality” and “agent foundations” would make less sense than you think.
This seems upstream of other disagreements you mention—e.g. how promising LIEDT is, or how much this work is “traditional agent foundations” (note that it’s being done in dialogue with several central examples of traditional agent foundations researchers, as per various links in the post).
Explaining Knightianism on one foot
Another positive impact of the modeling is it could convince politicians of the need to implement good AI regulations soon
This feels like exactly the same mistake I was critiquing Thomas for making. You can rationalize many things as helpful for convincing politicians, but if you sat down and actually tried to model what things would make politicians update in good directions, this would end up far from the top of your list.
(I do want to flag that it’s worth avoiding “optimizing over politicians” in an overly adversarial way, but I think my claim holds even given the constraint of cooperativeness.)
Even a law like “don’t use AI newer than 9 months old for AI R&D” needs justification to pass in the first place. So I feel like modeling is actually on the critical path here.
This seems like it’s relying on a model of how the legislature works where high-quality modeling is one of the key bottlenecks on laws getting passed. But that is just obviously false, and has been for a long time.
A more defensible version of this claim IMO is “modeling would help the AI safety community feel confident enough to strongly advocate for a ban on RSI”. But again, if you worked backwards from “why is the AI safety community not doing that”, then it’s very unlikely you’d end up at anything like your current strategy. Indeed, one key reason that the AI safety community doesn’t feel empowered to try to ban RSI is that it feels too enmeshed with the AGI companies and is scared of offending them, which becomes worse the more respected community members work there.
(I don’t have strong opinions on whether trying to ban RSI is a good idea, I just want to highlight that your reasoning here doesn’t stand up.)

I’m sharing here a rough draft of a new curriculum I’ve made, called Understanding Intelligent Agency. I’ve written a lot lately about how AI alignment should be trying to develop a new scientific paradigm; this curriculum is my attempt to point directly at what that new paradigm looks like.
I’m sharing this in draft form because I don’t know how long it will take me to get it to a state where I’m happy to promote it widely (at the very least, I’ll be offline for most of the next week). I’m hoping it can be useful to people before that; and please feel free to leave comments or suggestions for additional readings.