What just happened? Pragmatism and Pessimization
This post is about the major role alignment researchers played in advancing the frontier of AI capabilities over the last decade, and how the distinction between “alignment research” and “capabilities research” thereby lost most of its meaning.[1] In particular, I’ll chronicle the development of what I’ll call the “pragmatic alignment” paradigm, and how it helped the three leading AGI companies push hard on the path to AGI under the banner of safety.[2] This was not a subtle effect—it’s apparent even to outsiders who investigate the field, like authors Sebastian Mallaby and Karen Hao.[3]
In my previous post, I summarized the alignment community’s plan as “differentially advancing alignment over capabilities”. However, it’s worth being more precise about who was nominally pursuing that plan, because it doesn’t seem to have been very action-guiding for MIRI. For example, in 2015 Nate Soares described MIRI’s “deconfusion” research as being guided by the question “what would we still be unable to solve, even if the challenge were far simpler?”. Meanwhile Eliezer’s author surrogate in this 2018 post repeatedly emphasizes that people shouldn’t draw direct links from MIRI’s research to its potential applications. So my sense is that the “differential impact” criterion started off as merely a background consideration, then became much more load-bearing with the rise of EA-style thinking in the field, which involved justifying research directions by appealing fairly directly to their consequences.
This made people less rational both on an individual level and on a group level. On an individual level: it’s easy to generate rationalizations for why a given line of research is impactful on the margin, because there are many possible scenarios for how the future could play out (or how the past could have played out if you hadn’t intervened). So external incentives (or even just a strong emotional drive to have impact) can easily lead you to focus on the possibilities which suit you best. Especially within AGI companies, this gave rise to extremely motivated reasoning about counterfactuals in which alignment-branded interventions didn’t happen, helping people deceive themselves and others about their actual motivations. More generally, “differentially advancing alignment” is hard to demarcate from other consequentialist goals like “preventing overhangs” or “buying more time for alignment work later” which leave even more room for deception.
The group-level problem: the alignment community was very bad at dealing with these adversarial dynamics, which meant that it wasn’t able to prevent the gradual erosion of the boundary between alignment and capabilities research. In particular, extreme fear of publicly criticizing powerful people—and strong charitability/mistake theory norms—prevented the community from creating common knowledge of who was doing motivated reasoning, or just straightforwardly lying.[4] Even people pursuing enormously power-seeking strategies—most notably Sam Altman, and to a lesser extent Dario Amodei—were given the benefit of the doubt for many years. What I mean by “pragmatic alignment”, then, is the whole complex of people who were using and accepting consequentialist arguments about how to make AGI go well, while being emotionally and strategically committed to almost never calling out misuse of those arguments.
Pragmatic alignment is just one facet of the community’s unwillingness to directly challenge power structures, most notably exemplified in its lack of criticism of AGI companies until recently. To get a sense for how deep-rooted this resistance was, it’s worth reviewing the comments on this post by Ben Hoffman, and this post by Adam Shimi. (There were very few other discussions of this topic before ChatGPT; the most notable are Scott Alexander’s original objection to OpenAI and Jacob Hilton’s partial defense of OpenAI.) I should add that, upon revisiting Adam’s post just now, I found that I’d strong-downvoted it—I think because, when I first read it, I was scared of the alignment community alienating OpenAI. I feel quite viscerally horrified by this reminder of how sycophantic my past self was.
More on that in later posts. This post will focus specifically on a historical analysis of how the concept of “alignment research” was twisted towards boosting capabilities at OpenAI, DeepMind, and Anthropic. As I stated in my previous post, the most important point here is not that I’m confident that accelerating AI capabilities has been bad for the world—that would require a level of large-scale consequentialist reasoning which I can’t do reliably. However, what I am confident about is that people who tend to produce the opposite of their stated goals (a process I call pessimization) can’t be trusted with great power, and communities that fail to hold them accountable also can’t be trusted with great power.
By “hold accountable” I’m not referring to any centralized judgement process—we don’t have institutions reliable enough for that. Instead, I want individuals (like you!) to demand honest public conversations about what happened and what should have happened. People’s willingness to have those conversations (and your personal evaluations of how sincere they are) should then guide your decisions about who to work for, who to fund, and who to affiliate with more generally. At this point, someone in the field merely being open to alternatives to the current failed paradigm is sufficient to make me feel solidarity with them. Unfortunately, such openness is often constrained on an emotional level by the desire to remain part of existing networks of power, money, and ideological security.
I also want to be very clear that I’m trying to hold alignment researchers accountable not because I think they’re less ethical than other elite groups, but rather the opposite. Alignment researchers (especially the ones who have been around since the early days) think about their impact on the world more seriously and earnestly than any other comparably-sized community. (By contrast, we should interpret almost every capabilities researcher as being steered primarily by local incentives and power gradients, in a way that psychologically prevents them from seriously considering unconventional strategies.)[5] This gives me hope that (some subset of) the current field of alignment is able to learn from its mistakes. The first step is acknowledging that there’s no longer any widespread (implicit or explicit) definition under which “alignment research” (let alone “AI safety”) is robustly good for the world, based on the evidence I lay out below. The second is adopting more defensible norms and accountability mechanisms, like the ones I discuss at the end of this post.
The Prosaic Ideal, the Pragmatic Reality
Around the time that OpenAI was founded and OpenPhil became active in the field, alignment started undergoing a partial paradigm shift towards a more pragmatic and empirical approach. One early step was the Concrete Problems in AI Safety paper (which I’ll discuss in more detail in my next post). Dario Amodei was both the lead author on this paper and one of the main people formulating this new approach. However, he was still new to the field, and didn’t write much publicly about his views (though this 2014 discussion with Eliezer is a useful source). My impression is that Carl Shulman had some similar ideas but also didn’t articulate them publicly until significantly later. So I’ll focus on cataloguing the shift with reference to Paul Christiano’s extensive public writings, which were the main intellectual arguments updating the alignment community’s worldview. Note that I’m grateful to Paul for recording his thinking in enough detail that I can try to trace what went wrong a decade later; readers should keep in mind that many others influenced the events I describe in less legible ways that make accountability harder.
Paul had been active on LessWrong since 2010, and had started doing significant alignment research by 2013. In addition to authoring several agent foundations papers, he blogged on a wide range of topics. In the following years he developed a new perspective on alignment. The most concrete milestone was his 2018 post arguing that we’d see a slow takeoff; another was his 2019 post articulating more gradual threat models than Yudkowsky’s. In some ways, these posts built on Hanson’s side of the Hanson-Yudkowsky foom debate, but Paul was more willing to accept the premise that general intelligence would be a really big deal, and merely dispute the trajectory by which we would reach superintelligence. In hindsight, he has been vindicated in his arguments for a much slower takeoff than Eliezer originally predicted.
Another important part of Paul’s new paradigm was the idea of “prosaic AGI”: an AGI built in a way “which doesn’t reveal any fundamentally new ideas about the nature of intelligence or turn up any ‘unknown unknowns.’” In a sense, the whole field of deep learning is a prosaic approach to AGI, compared with previous methods. But even after its early successes, the additional belief that deep learning would scale up easily took longer to propagate. Dario Amodei wrote a long, never-released google doc advocating for the “big blob of compute” hypothesis around 2018. The publicly-available posts with the most similar content are probably Sutton’s bitter lesson post and Gwern’s scaling hypothesis post. I also recall Jan Leike giving a presentation to the safety team at DeepMind in 2019 arguing for ~6-year timelines based on similar intuitions. My sense is that almost nobody else at DeepMind except Shane Legg was sympathetic to this view.[6]
The prosaic AGI intuition has been vindicated since then: we’re now much closer to building AGI, and we haven’t learned any fundamentally new things about intelligence in the process. But the reason I only called it a partial paradigm shift is that Paul didn’t manage to carve out a defensible research strategy. His original prosaic AI alignment post argued against two separate camps. On one side, he critiqued people who claimed that “it’s impossible to do meaningful work without knowing more about what powerful AI will look like”. This reasoning is similar to the arguments Dario and Geoffrey gave for working on scaling up LLMs. On the other side, he critiqued people who claimed that “aligning prosaic AGI is probably infeasible”. My understanding is that MIRI used this claim to justify trying to build (agent-foundations-based) AGI themselves (see Wei Dai’s comment on my previous post for more details).
So Paul seems to have been trying to steer a path between two opposing “alignment” strategies which both prescribed building AGI yourself—an admirable intention, if so. My diagnosis is that he didn’t succeed because he made versions of both the individual-level mistake and the group-level mistake that I described above. The former involved characterizing prosaic AI alignment as being in opposition to “understanding intelligence”. My sense is that both MIRI and Paul were implicitly treating “understanding intelligence” as mainly valuable for building aligned AGI from scratch—which wouldn’t count as prosaic AI alignment. However, there’s another possibility: that an AGI which would otherwise be built without an understanding of intelligence is aligned using an understanding of intelligence! I’m currently excited about agent foundations precisely as a strategy for aligning otherwise-prosaic neural-network-based AGIs—but this strategy is implicitly ruled out by Paul’s framework.[7]
This mistake was exacerbated by Paul’s strategic mistake of joining OpenAI to work on the same projects that other people were justifying for very different reasons. Because of this, the success of Paul’s empirical predictions (and his general thoughtfulness about alignment) was then taken as evidence in favor of OpenAI’s research directions and overall strategy. Paul conspicuously failed to correct this impression by critiquing OpenAI publicly—I can’t find any critical statements from when he worked there, only an endorsement of the OpenAI safety team (which was run by Dario). It’s very normal not to publicly criticize your boss or your company, but for anyone who’s trying to significantly influence the world—and especially a leader of a key movement—the willingness to do so seems like a very basic foundation for maintaining integrity.[8] In the absence of that, Paul’s “prosaic AI alignment” paradigm devolved into a paradigm in which essentially any consequentialist arguments for building AI systems or allying with AI companies were accepted as valid AI safety strategies, as I’ll catalogue in the next three sections.
OpenAI
The intermediate step between “actually trying to solve the alignment problem” and the fully-pragmatic paradigm was scalable oversight. Around 2018, three maybe-probably-equivalent scalable oversight proposals were floating around: Paul’s iterated amplification, Geoffrey Irving’s debate, and Jan Leike’s recursive reward modeling (Jan started at DeepMind, but moved to OpenAI in 2021). Iterated amplification was by far the most-discussed amongst alignment researchers. Paul’s (notoriously opaque) arguments focused on the idea that if imitating humans is safe, then we can combine many imitation learners to produce more capable (but still safe) agents. However, I broadly agree with Yudkowsky’s critique that this hides the hard part of the problem in the interactions between the subagents.[9]
More importantly, whatever theoretical merits these proposals had were immediately decoupled from the engineering work that Paul, Geoffrey, Jan and Dario actually started doing—specifically, work on reinforcement learning from human feedback. This started with agents learning simple behaviors in toy environments, but soon progressed to a series of papers applying RLHF to LLMs, culminating in InstructGPT. While these were impressive efforts on an engineering level, there’s very little that distinguishes them from what a prescient capabilities-maximizer would have been doing—for example, although Paul’s theoretical justifications for iterated amplification referred a lot to the safety properties of imitation learning, all of these papers added RLHF for better performance.[10]
This focus on engineering-style work was facilitated by Dario’s push to scale up from GPT-1 (which was mainly Alec Radford and Ilya Sutskever’s project) to GPT-2 and subsequently GPT-3, justifying this in significant part by arguing that it would help boost alignment research. In Empire of AI, Karen Hao reports Dario telling her in 2019 that “We want a language model that humans can give feedback on and interact with [where] the language model is strong enough that we can really have a meaningful conversation about human values and preferences.” My understanding is that Paul opposed this strategy internally, but Geoffrey supported it. The Infinity Machine quotes Geoffrey as recounting “We struggled for a while [to get LLMs to obey instructions]. Then we were like, OK, let’s just make the language models stronger.”
Subsequently, Dario led the effort to scale up GPT-3 training to 10,000 V100 GPUs. In addition to arguments that better safety research required more capable models, my understanding is that he was also trying to increase OpenAI’s lead against China; I’ll discuss that kind of reasoning in more detail in a later post. Before leaving OpenAI, Dario also released the scaling laws paper, which did a lot to wake the academic ML community up to the plausibility of AGI. I don’t have a strong sense of what we should infer from this, since I’m predisposed to be positive about scientific communication, but it does seem like more evidence against the idea that Dario was following a coherent and sensible plan.
Meanwhile, John Schulman had been at OpenAI from the beginning. He was sympathetic enough to safety to coauthor the Concrete Problems paper, but primarily worked on reinforcement learning (e.g. pioneering PPO). By 2021 (the year I joined OpenAI) he was working on WebGPT, a way of letting GPT models browse the internet. I remember him articulating reasons to think of WebGPT as an alignment project (something like: if models can look up information online, they’ll be more honest). These justifications were apparently sufficient to get a number of alignment-motivated researchers to work on it (in particular Jacob Hilton—the first author of the blog post—Jeff Wu, and William Saunders). WebGPT was the direct predecessor to ChatGPT, and my understanding is that ChatGPT inherited a lot of WebGPT’s codebase (as well as ideas and techniques from InstructGPT). More specifically, a researcher who was on the team around that time described ChatGPT to me as “WebGPT minus the Web”: the basic Q&A format and RLHF fine-tuning were already there, but ChatGPT lacked WebGPT’s unreliable web browsing component.
Paul has since written up his justifications for working on RLHF, and why he doesn’t think RLHF was very important for ChatGPT. However, these arguments seem very suspect (for reasons explained well by Habryka). For example, Paul says “I think the effect [of ChatGPT] would have been very similar if it had been trained via supervised learning on good dialogs”. But the InstructGPT blog post reports that “our labelers prefer outputs from our 1.3B InstructGPT model over outputs from a 175B GPT‑3 model [trained with supervised fine-tuning, as per Figure 1 from the paper], despite having more than 100x fewer parameters”.[11] Another important datapoint comes from Sydney Bing, which wasn’t trained with RLHF and produced fairly unhinged outputs, suggesting that RLHF was important for making ChatGPT user-friendly.
In hindsight, the launch of ChatGPT was one of the most acceleratory events in the history of AI, funneling many billions of dollars into the field (ChatGPT grew faster than any previous product in history). I don’t have a great recollection of whether or how the ChatGPT team justified this launch in safety terms; I expect their reasoning was that someone else would do it if they didn’t. But as I’ll discuss shortly, the main potential “someone else”s were also researchers nominally motivated by alignment, who were also justifying their work with the idea that someone else would do it anyway. At the very least this was a colossal coordination failure within the community; I also think it undermines the core premises people were using to reason about how to have impact.
One such premise was the idea that, if a system was developed using relatively few resources, it could likely be quickly scaled up to many more resources, which might create a dangerously sharp transition. The possibility of such “overhangs” was discussed at least as far back as the 2008 Eliezer-Hanson debate (with hardware as the limiting resource), but only as a background strategic consideration. At some point, people started using overhangs as justification for making rapid progress now, to use up all the low-hanging fruit so that later progress would be slower (and therefore less dangerous).
Reasoning of the form “we’ll do something we’re worried about so that other people do less of it later” is always extremely slippery, in a way that common-sense morality (and even just common sense) weighs strongly against. This case was no different. Broadly speaking, people would pick whichever inputs to AI progress they wanted to defend speeding up, and just assume (often even without directly stating it) that there were other background constraints which meant that speeding up their preferred inputs wouldn’t make much long-term difference.
In case this seems like an exaggeration, consider these two discussions of overhangs from Paul:
“If LM agents are weak are due to exceptionally low investment and understanding it creates “dry tinder:” as incentives rise that investment will quickly rise and so low-hanging fruit will be picked. While there is some dependence on serial time, I think that increased LM investment now will significantly slow down progress later.”
And from this post:
“Avoiding RLHF at best introduces an important overhang: people will implicitly underestimate the capabilities of AI systems for longer, slowing progress now but leading to faster and more abrupt change later as people realize they’ve been wrong. Similarly, to the extent you successfully slow scaling, you are then in for faster scaling later from a lower initial amount of spending—I think it’s significantly better to have a world where TAI training runs cost $10 billion than a world where they cost $1 billion.”
The most obvious, basic model of progress is that things take time, so doing stuff earlier will allow people to do more stuff later. Indeed, one of Paul’s most significant intellectual contributions was the argument that recursive self-improvement will continuously ramp up over time—which implies that pushing AI forward will have compounding effects. It’s possible in principle that local bottlenecks could override these dynamics. But if we argue for speeding up algorithmic progress and investment and public understanding (and even elicitation) of AI capabilities based on overhang arguments, then there’s almost no room left for limiting factors to kick in later. There’s something like a “bottleneck of the gaps” here—i.e. the “limiting factor” is whatever some safety person hasn’t decided to work on yet, and tends to zero as AI safety people find arguments for accelerating every possible input to AI capabilities. (What about the difficulty of getting US visas for AI researchers? Remco Zwetsloot and other DC safety advocates have worked on it (see section 5.1). What about the fact that Europe isn’t a leading player? I’ve talked to several AI governance people who are considering kickstarting a Europe-wide AI project, on the grounds that they like European values. And so on.)
The strongest fallback for overhang advocates was the difficulty of increasing the hardware supply. However, these arguments are also looking very shaky. AI progress has redirected capital at a civilizational scale: AI investment was 39% of US real GDP growth in the first nine months of 2025, and “capital expenditure of just five technology companies is now larger than global investment in oil and natural gas production”. (As a cynic would expect, AI safety people have specifically been homing in on the most acceleratory investments—for example, Situational Awareness just invested $400 million to disrupt a key chip production bottleneck.) In some sense the “compute overhang” argument remains unfalsifiable, because we can always construct counterfactuals which are worse than our current situation. But for any practical purpose, the final nail in its coffin is the fact that so many AI safety people are now taking seriously the idea of an imminent “software-only singularity”—see Tom Davidson, Ryan Greenblatt, and Paul himself (in non-public talks and writing). Insofar as they’re right, all work which was (explicitly or implicitly) justified by the idea of reducing the hardware overhang has been directly pulling us towards the singularity. (To be clear, I don’t expect a software-only singularity; my point is that the worldview which accepted “overhang” justifications is no longer coherent.)
What went wrong here? It’s hard to know exactly what led any given person to endorse any given argument. But when we zoom out, it becomes clear that many people in this space really want to pull some lever that feels important, and privilege arguments which justify that. That might come from a sense that they need to have an impact on the world; or fear about failing to fulfil their potential; or the more mundane explanation that big levers tend to be associated with money and prestige and proximity to power. Certainly the latter was a large part of why I joined OpenAI originally; I expect that most people who joined earlier were less prestige-oriented than me, but still made that decision using reasoning that was warped by similar emotional drives (and later further warped by the social dynamics of actually working there).
To describe that warping, I find a version of Ajeya’s saints, sycophants, schemers trichotomy useful (though I think of it as a spectrum between fully scheming and fully sincere). Ajeya characterizes sycophancy as focusing on short-term approval—my sense is that humans implement this via flinching away from criticizing, contradicting or feeling cynical about powerful people. This tendency combines very badly with the kinds of arguments I’ve been discussing, which provide many degrees of freedom for rationalizations. As one example, folks at OpenAI (and even in the wider alignment community) were far too accepting of Sam Altman claiming that rushing towards AGI would be helpful for safety. I remember him arguing in person in 2022 or 2023 (and in this blog post) that faster algorithmic progress towards AGI would help alleviate a potential compute overhang. In hindsight, I’d describe my reaction as “flinching away from the possibility of no longer taking his claims at face value”. I only viscerally internalized that Sam had been lying about his motivations when I later heard about his plans to raise enormous amounts of money to build new chip fabs. This was shocking to me not just because it directly contradicted the arguments he’d been giving, but because it contradicted them to a greater extent than I’d even been able to consider as a plausible hypothesis.
For those who don’t know Sam, it might seem odd that I ever took his arguments seriously even given my tendency towards sycophancy. One underappreciated factor is that he has something similar to Steve Jobs’ reality distortion field—but in his case I’d call it an earnestness field. His intonation and body language send very strong signals of sincerity; and he does enough things motivated by earnest nerdiness that it’s easy to rationalize away discrepancies. Modeling this dynamic is necessary to explain the very high ratio between people who polarize against him and concrete evidence of his misbehavior. When people realize that Sam is lying (even about things that don’t matter much) while embodying that level of earnestness, there’s a strong visceral update away from trusting him, which is hard to convey to others.
DeepMind
There was a similarly intertwined relationship between capabilities and alignment at DeepMind, as exemplified first by Shane Legg and then by Geoffrey Irving. Shane was in a strange position from the beginning: before founding DeepMind he’d been an early LessWronger who’d given talks warning about AGI risk. By the time I joined DeepMind in 2018 Demis had almost all the executive power, and Shane seemed to be somewhat sidelined within the organization. However, he continued to provide a central example of self-sabotaging “AI safety” strategies, because he’d recently founded two teams: the technical AGI safety team (TAGIS), and a secretive effort called the AGI team (which some friends at DeepMind nicknamed the “danger team”). Both teams were outliers at DeepMind in how seriously they took AGI, and both faced recruiting challenges as a result (with TAGIS mainly hiring people without traditional ML backgrounds, and the AGI team mostly containing research engineers, for lack of research scientists who wanted to focus on AGI).
The AGI team focused on training AIs to control virtual avatars in simulations, analogous to how humans evolved. For a while they were developing a huge virtual game-world called Gaia, which was intended to help agents learn intelligence by recapitulating aspects of evolution (such as hunting and eating each other)—though I don’t think anything ever came of it. If I recall correctly, the only DeepMinders working on anything language-related around 2018-2019 were also working in game-like environments—specifically using imitation learning and RLHF to train virtual avatars to follow natural-language instructions. Jan Leike, Miljan Martic and I did some work on this in 2019 while on TAGIS (though I was very unproductive, in a way I now recognize as being driven by alienation from the work). Eventually a larger “Interactive Agents Group” started doing similar things, and produced a public-facing report.
This focus on virtual environments reflected an underlying belief (amongst the few people thinking seriously about AGI at DeepMind) that embodiment of some kind was crucial for training AGI. More generally, the most senior people at DeepMind (especially Demis and David Silver) were scientists who had strong inside views about which kinds of algorithms and insights would push AI forward. Because of this, DeepMind as an organization paid relatively little attention to GPT-1 or even GPT-2, which were more engineering-driven projects. It took Geoffrey Irving joining DeepMind from OpenAI to consolidate a real push towards building LLMs. As Mallaby recounts in The Infinity Machine:
Irving’s arrival tipped the balance at DeepMind. He had spent time inside the belly of the rival beast: He spoke with the authority of one who understood what state-of-the-art language research looked like. Although he could not explicitly say so, he knew that OpenAI had already developed models that were more than ten times larger than GPT-2, though these had not been released yet. Irving’s message to his new colleagues was that they better up their game. A race for supremacy had begun without DeepMind even realizing it.
To hammer home his point, Irving reproduced a paper that he had written at OpenAI: “Language Is Enough.” The argument was the opposite of Hassabis’s position. According to Hassabis, language’s lack of real-world “grounding” limited its value. According to Irving, language crystallized the knowledge of humans, who were themselves grounded—therefore, the grounding problem was exaggerated.
In 2020, Geoffrey kicked off work (with Jack Rae) on Gopher, DeepMind’s first LLM. Afterwards, while others took over the scaling work, Geoffrey led the development of Sparrow, a model fine-tuned with RLHF. While the paper’s title pitched it as “Improving alignment of dialogue agents via targeted human judgements”, the work was important for the eventual development of Gemini. Reflecting on Sparrow, Demis Hassabis said “I thought it wouldn’t work because just using RL seemed too easy. But the team went ahead and did it, and then of course it did work. The raw networks were not that compelling to talk to, right? You needed RLHF to build a real chatbot.”
In addition to arguments that larger LLMs were necessary for doing good safety research, I recall various people arguing that LLMs were a safer path to AGI than DeepMind’s RL-focused approach (I don’t recall who argued this originally, but here’s a similar argument made more recently by Paul). Hence, they argued at the time, accelerating LLMs would be beneficial for safety. In hindsight, though, it seems like LLMs were a crucial bottleneck on AI capabilities, in which case accelerating them was a very direct AI capabilities advancement. It’s possible that the arguments were still good reasoning ex ante, but it seems much more likely that people were mainly finding rationalizations for things they wanted to do anyway. In particular, being scared of RL should weigh heavily against pioneering RLHF—so the fact that “alignment researchers” at all three AGI companies first scaled LLMs then added RLHF is significant evidence that arguments about LLMs being a safer path to AGI were insincere.
Anthropic
The effects of “alignment researchers” at Anthropic are a little harder to talk about, because Ants give a wide range of justifications for their work. Sometimes they talk about promoting alignment, but sometimes they talk about beating OpenAI, or beating China (or, increasingly, beating Republicans). In subsequent posts, I’ll analyze in more detail how people who want AGI to go well should evaluate that reasoning. As a quick preview: while I give Anthropic credit for some laudable moves (like not releasing Claude before ChatGPT), I also think that they are choosing to ignore many of the harmful effects of their strategy. One particularly notable blind spot (at least in public discussions) is how Dario’s early racing on behalf of OpenAI played a big role in creating the “problem” that he now purports to be solving by racing on behalf of Anthropic.
For now, though, I want to analyze how work Anthropic specifically promoted as “alignment” contributed to the concept of alignment becoming watered down to meaninglessness. The first paper Anthropic released was “A general language assistant as a laboratory for alignment”. The paper “was motivated by the problem of technical AI alignment, with the specific goal of training a natural language agent that is helpful, honest, and harmless”. The crucially important point, though, is that they weren’t trying to make existing AIs more HHH—rather, they were inventing natural language agents in order to have something to train to be HHH (as nostalgebraist discusses). The paper is more explicit on this point later on:
Most research efforts associated with alignment either only pertain to very specialized systems, involve testing a specific alignment technique on a sub-problem, or are rather speculative and theoretical. Our view is that if it’s possible to try to address a problem directly, then one needs a good excuse for not doing so. Historically we had such an excuse: general purpose, highly capable AIs were not available for investigation. But given the broad capabilities of large language models, we think it’s time to tackle alignment directly, and that a research program focused on this goal may have the greatest chance for impact.
I don’t know which of the authors of this paper sincerely thought they were differentially promoting alignment, and which were rationalizing building the most capable AIs they could; either way, it’s ironic to the point of absurdity that Anthropic described building the predecessor to Claude as “tackl[ing] alignment directly”. This deep entanglement between capabilities and “alignment” was also apparent in a subsequent paper, “Training a Helpful and Harmless Assistant with RLHF”, which paralleled OpenAI’s InstructGPT and DeepMind’s Sparrow.
Recall that Paul, Geoffrey and Jan justified work on RLHF as a first step towards the “next thing” in scalable oversight. To a first approximation, this next thing never came. Rather than designing principled methods by which humans could verify AI behavior, Anthropic delegated more and more of that process to AIs themselves. A first step was replacing RLHF with RLAIF in their “Constitutional AI” paper. They followed this up with papers on model-written evaluations, model-assisted red-teaming, and model-assisted question-answering. My sense is that, by now, Anthropic uses AI assistance far too pervasively and haphazardly for them to reliably track whether or how even current models are deceiving them.
So the “helpful” in HHH merged “alignment research” with “capabilities research”. Meanwhile the “harmless” merged “alignment research” with “ideological control”. The “Helpful and Harmless Assistant” paper doesn’t go into much detail on what they mean by “harmless”, but it’s implicitly about political correctness—their main examples of “harmful” behavior involve gender bias, calling mentally ill people “crazy”, and opining on gay marriage. Anthropic’s subsequent work on “Red-teaming language models to reduce harms” makes this ideological component even clearer: out of the six categories of “harms” they list in the introduction, three are clearly ideological in nature (reinforcing social biases, generating offensive or toxic outputs, generating extremists texts), two are commonly used as pretexts for censorship (aiding in disinformation campaigns, spreading falsehoods), and only one is clearly non-partisan (leaking personally identifiable information from the training data).
My sense is that a whole subfield emerged from this work and similar thinking at OpenAI; I don’t think it has a consensus name, but we might charitably call it “product safety”, or less charitably call it “brand safety”. At OpenAI, early work on this was spurred by the desire to block early users of GPT-2 from getting it to produce text-based porn (especially child porn, especially via AI Dungeon). Later work fell under the remit of Lilian Weng’s Safety Systems team, which implemented guardrails and monitoring for OpenAI’s products. Most alignment people at OpenAI viewed this as an important step towards more xrisk-focused guardrails and monitoring, without thinking much about the censorship angle. I’ve been paying relatively little attention to this since I left OpenAI, but my sense is that there are now both significant politically-skewed restrictions on what frontier models will talk about, and significant political biases when they do respond. It’s hard to trace exactly which people and techniques caused this; it may be best explained in terms of organizational prioritization. For example, GPT-4o expresses preferences which imply that it values the lives of Nigerians at roughly 20x the lives of Americans. Even without knowing what caused this, I expect that OpenAI would have been much more concerned, and done much more to change it, if the disparity were the other way around.
This didn’t come out of nowhere. Instead, it’s best understood as a replay of the process by which almost all major internet platforms implemented mass censorship against “harmful” ideas and speech over the last decade, at a speed and scale that’s hard to overstate. The linked article is extremely worth reading; a brief summary is that within less than a decade “the Internet went from a space for people without institutional backing to get their views out to one with regular purges and demonetizations of heterodox figures and those associated with them, encouraged by those very same non-tech organizations that formerly championed Internet freedom. This was justified as a response to (massively overblown and mostly fictitious) Russian influence campaigns, and as fighting nebulous “hate,” the definition of which could be shifted at will to cover whatever views or ideas the organizers classifying it wanted it too and to exclude those they didn’t, and “misinformation.””
It seems like AGI companies straightforwardly copied the terminology and playbook of social media censors, down to the establishment of “Trust and Safety” teams. Understanding both of these processes, and the parallels between them, seems extremely important for making good decisions about the future of AI—but the rationalist community hasn’t paid much attention to this, because most of the censorship happened to right-wingers who it finds distasteful.[12] From reflecting on this, I’ve become much more sympathetic to Elon’s focus on building a “maximum truth-seeking AI”, as setting this goal seems necessary (although not sufficient) to prevent the ideological capture under the banner of “safety” that has happened at every other AGI company.
To be clear, I do think that Anthropic did some scientifically valuable research in its early years. Most notably, the mechanistic interpretability team under Chris Olah was doing very cool work (especially in their pre-SAE period). I’ll also pick out Language models (mostly) know what they know as an interesting scientific finding; I’ll talk more about both of these examples in my next post. However, even amongst Ants who call themselves alignment researchers, the kind of work that’s even trying to learn generalizable facts seems dwarfed by the amount of work that blurs the alignment/capabilities line.
I’ll briefly flag two ongoing examples of the latter. Firstly, scalable oversight to grade currently-unverifiable tasks seems like it might be the next big capabilities bottleneck—I expect that most work on this will be scalable enough to create important training data for current models, but not scalable enough to reliably oversee significantly more capable models. Secondly, Anthropic seems to have been pushing hard towards “automating alignment research” over the last year (I expect that Jan Leike played a significant role here, since he’s been advocating for this for many years). Making models better at “alignment research” is obviously extremely similar to making them better at capabilities research; people have been justifying it anyway by talking about having marginal impacts in worlds where these skills diverge. I’ll rebut these specific arguments later in the sequence; however, anyone who’s read this far should have a sense of why this kind of work will predictably speed up recursive self-improvement much more than its proponents expect (e.g. if “automating alignment research” had been a thing a few years ago, it’s easy to picture that line of work inventing reasoning models, particularly if it were aimed specifically towards improving conceptual reasoning).
If not alignment research, then what?
Above, I’ve recounted how the standards for what counts as “alignment research” have fallen dramatically over time. After I noticed both how load-bearing and how ambiguous the alignment/capabilities distinction had become, I spent some time trying to salvage it. But as I wrote this post, I concluded that it’s time to give up on “alignment research” as a rallying cry; it’s become too corrupted. (“AI safety” is even worse as a term, and these days is mainly useful for describing a social cluster.)
I want to make sure to clarify what I do and don’t mean by this. I still consider (some version of) the alignment problem to be real and extremely important; and most of the intellectual progress towards solving it is still coming from people proximate to the alignment community (though the best researchers have kept themselves at arm’s length, as I’ll discuss in my next post). However, this is mixed in with enough harmful and deceptive work that it no longer seems defensible to me to try to promote the field broadly, or to defer to the field’s consensus about what research will help.
More generally, insofar as I’m optimistic it’s largely despite the efforts of mainstream alignment researchers, rather than because of them. And so, from my perspective, the alignment community has lost any moral right to try to gain power on altruistic grounds, or to pursue plans primarily motivated by backchaining from large-scale effects on the world. The Pause/Stop AI movement does seem to avoid some of these failures (in particular everyone else’s lack of courage), which means I’m more excited about them than the rest of the alignment community. However, they don’t seem to be thinking clearly enough about politics to have robustly good effects on the world (e.g. to reliably distinguish between the kinds of strategies that push towards dictator-level concentration of power, and the ones that don’t).[13]
Again, I’m not claiming that the alignment community is unusually unethical: I don’t know of any other similarly-sized community which is able to avoid the corrupting effects of this much power (though there are plenty which are wise enough to avoid accumulating power because of that). I acknowledge that it’s hard to pivot your worldview when there’s no clear alternative to adopt. However, that’s precisely the period during which clear, open-ended thinking is most valuable. So I expect that most of the direct benefit of this post will come from inspiring a few relatively courageous individuals to move towards (emotional, social, and financial) independence from the existing field—enough that they’re able to think clearly about what went wrong, and help work towards a better paradigm. I suspect that the first step for many of them is to panic less about short timelines (e.g. by taking scenarios like this one more seriously)—though I also expect cultivating courage and integrity to make your work dramatically more valuable fairly quickly (as this tweet discusses). In the longer term, I want the community as a whole to halt, melt, and catch fire: to “Say, ‘I’m not ready.’ Say, ‘I don’t know how to do this yet.’” Eliezer’s Death with Dignity post was a step towards this, but focused too much on whether we were on track to solve the alignment problem, and too little on the adversarial dynamics that have been pushing us in the wrong direction. I hope that this sequence will point people more directly towards reevaluating.
I’ll talk more about my alternative mission of high-integrity scientific research (and why it captures the parts of the field I most want to promote) in my next few posts. For now, I’ll focus on a few high-level principles for starting to move in that direction. The first: on an intuitive level, you should think of many arguments about differential impact on the margin as analogous to arguments for timing a stock market bubble. If someone argues that the market as a whole is in a bubble, but that they’ll invest your money while it’s still going up and sell before it drops, you should probably be very skeptical. I think this analogy is actually quite deep, because the core difficulty in both cases is accounting for other people making decisions which are tightly entangled with yours. It seems possible to account for this in principle, but in practice it’s very easy to fool yourself (especially when you’re used to doing econ-style reasoning about marginal effects)—and when you do so, you’re making the bubble bigger. So I don’t trust myself (or basically anyone else in the field) to think clearly about such cases; it seems far better to focus on more robust strategies.
Okay, but how should you evaluate which strategies are robust? One foundational step is to assume that you are choosing on behalf of a significantly wider range of people than just yourself. A range of different considerations support this conclusion, including:
The idea that you’re setting norms for the field, which helps build a high-trust community.
The idea that others will copy your behavior—whether due to trusting your decision-making process, or simply because you’ve made it more socially permissible for others to behave similarly.
The idea that it’s more important to avoid underestimating than avoid overestimating your influence—because if you underestimate your influence then your actions matter much more than you thought.
The idea that being right about a problem (e.g. AGI risk) is correlated with being right about other things, and so your work might be much more impactful than others’ work.
The idea that others will make decisions which are logically correlated with yours.
Underlying deontological or Kantian moral intuitions about universalizability.
Someone who followed this principle would be much less likely to join (or stay at) unethical organizations to do “harm mitigation”; and they’d be much less likely to justify racing “because we’re the good guys”. They would also favor research directions that they think would reward deep investigation, rather than shallower ones which mainly seem helpful on the margin (or which are even harmful if too many people pursue them). Deciding how to apply this principle will always require individual judgement, but I’d suggest erring towards overapplying it rather than underapplying it. Even if you thereby leave some value on the table, you’re also helping establish yourself as a more trustworthy person.
However, this principle is still quite blunt—especially for people who are in fairly unique situations. A second principle is that, when making more complicated decisions, people should articulate cruxes for their decisions, and then be expected to either acknowledge when those beliefs were disproved, or else clearly publicly state when they’ve changed their cruxes. I think these are much more valuable when done by individuals voicing their own opinions; group statements tend to produce accountability sinks. For example, if Dario had publicly discussed his intention for Anthropic not to advance capabilities, then it would have been much easier for the alignment community (and Anthropic employees) to respond appropriately when he started pushing the frontier. As it is, not a single Anthropic employee has publicly resigned over this dramatic change in Anthropic’s strategy, which suggests significant frog-boiling dynamics.
An example from this sequence is my argument that research is robustly valuable insofar as it a) aims towards a deep scientific understanding, and b) is done by high-integrity people. In the short term, you might disagree that this is a good target; in the longer term, though, seeing me stick to this standard (or explain why I changed it) should help you trust that I’m not being corrupted in the standard ways. Relatedly, I give Paul some credit for writing his retrospective on RLHF, but the arguments still seem very defensive, rather than an attempt at a neutral evaluation of what he did right and wrong. The closest Geoffrey has come to giving such a retrospective is this post on why he joined AISI.[14] Almost none of the others who have had most influence over the field (like Yudkowsky, Vassar, Shulman, and Karnofsky) have done so either; I hope that this sequence spurs some of them to do so. I’ll also have a lot more to say about my own mistakes over the next two posts; if you ever think I’m holding other people to a higher standard than I hold myself to, please tell me so.
A third standard is that improving the world requires enough integrity to sometimes move away from money, prestige, or power. For example, MIRI was willing to make their research nondisclosed-by-default due to concerns about capabilities externalities. Similarly, Janus was aware of chain-of-thought prompting over a year before it became mainstream; my understanding is that she didn’t publicize it widely due to concerns about accelerating capabilities. It’s notable that it’s precisely the outsiders with fewest resources who are willing to make these sacrifices—contrast Dario being unwilling to hold back even a paper as directly acceleratory as “scaling laws” (edit: retracting this claim due to additional evidence here and here). (I do somewhat credit Paul and Geoffrey for stepping away from AGI companies to work in government, but not a huge amount, because this still involves moving away from one kind of power towards a different type of power.)
To be clear, I’m not against people who care about alignment accruing significant power. Rather, I’m against them doing so under false pretenses, and without possessing a concomitant level of integrity. One reason the alignment community (especially the EA components of it) often fails to track the latter is that it takes charitable donations or altruistically-motivated sacrifices (like veganism) as evidence that people should be trusted. Unfortunately, it turns out that altruism and integrity are two very different things (as SBF showed in dramatic fashion). Much stronger evidence for integrity comes from criticizing or standing up to powerful people even when few others around you are doing so. Unfortunately, there are few clear examples—the main ones are Daniel Kokotajlo at OpenAI, Yudkowsky’s Time essay, Pause/Stop AI advocacy, and to some extent the OpenAI board and Anthropic’s stand against the Trump administration. I’ll explore these examples in later posts; in my next post, though, I’ll discuss the underlying mindset that “someone else will do it”, which skews many decisions made across the field.
- ^
Anyone who read an early draft of this post should note that this public version is over twice as long, and makes a much more detailed and hopefully clearer argument than the original.
- ^
This is different from Dan Hendrycks’ concept of “pragmatic AI safety”. I’ve appropriated the use of the word “pragmatic”, with apologies to Dan, because it seems like his term has fallen out of use. I don’t have a strong opinion on how much Dan’s research program overlaps with the thing I’m calling “pragmatic alignment”.
- ^
I haven’t included direct quotes in the main text because both authors make this point in ways that are only partly true. In The Infinity Machine (Page 287), Mallaby writes that “paradoxically, the aggressive scaling favored by Amodei, Irving, and Christiano turned out to be the starting gun in a destabilizing AI race”. He was referring to the scaling up of GPT-2, which Paul Christiano tells me he didn’t support.
Meanwhile, in Empire of AI, Karen Hao writes:
“What is AGI? What does AGI look like?” Amodei said. “Well, you know, we’re in the awkward position of, we don’t know what it looks like. We don’t know when it’s going to happen. So we look for things that aren’t AGI but that present at least some of the opportunities and difficulties of AGI. And the hope is that if we can handle those things well, then we’re kind of, like, ready for the bigger leagues.”
It was a logic that worked under a specific assumption: that AGI, despite being amorphous and unknowable, was also inevitable. OpenAI would repeatedly justify its behaviors against variations of the same argument for years after. Under the specter of AGI’s unstoppable arrival, the company needed to keep developing more and more powerful models to prepare itself and to prepare society. Even if those models carried with them their own risks, the experience they offered to prevent or face possible AI apocalypse made those risks bearable.
[However] it was specifically OpenAI, with its billionaire origins, unique ideological bent, and Altman’s singular drive, network, and fundraising talent, that created a ripe combination for its particular vision to emerge and take over.… In other words, everything OpenAI did was the opposite of inevitable; the explosive global costs of its massive deep learning models, and the perilous race it sparked across the industry to scale such models to planetary limits, could only ever have arisen from the one place it actually did.”I think Hao is incorrect that LLMs could only ever have arisen from OpenAI, because she’s not taking Moore’s law seriously enough. More generally, her book seems to often be trying to “score points” against the tech industry.
Despite this, the fact that both of these authors homed in on AI safety arguments backfiring at AGI companies seems very notable to me.
- ^
Because strict adherence to these norms worked out so badly, I partially set them aside in this post; I’m still trying to be fair to everyone involved, but I don’t take people’s stated motivations as authoritative to the extent that rationalists usually do. Instead, I try to build up a better understanding of how sycophancy and fear warped people’s thinking (including my own).
- ^
While this is bad for the world, it’s also one of the key reasons that the alignment community is able to exert such outsized influence, as I’ll detail in my next post.
- ^
It’s important to note both that this view was “directionally correct” in predicting that AI progress would be faster than almost anyone thought, and also “literally wrong” in that Dario and Jan expected that we’d already have AGI by now.
- ^
It seems like the root of this conceptual mistake might have come from Paul treating “build AGI” and “align AGI” as sequential steps. In practice, though, we should expect alignment techniques to be applied throughout the process of “building” the AGI. If both the building process and the alignment process are prosaic (or both non-prosaic) then we still have a clean distinction. But if the alignment techniques are non-prosaic, then applying them to an otherwise prosaic AGI creates an edge case in the framework.
My guess is that Paul didn’t explicitly consider this possibility, because he characterizes the following as an objection to prosaic AI alignment: “Some researchers (especially at MIRI) believe that aligning prosaic AGI is probably infeasible — that the most likely approach to building an aligned AI is to understand intelligence in a much deeper way than we currently do, and that if we manage to build AGI before achieving such an understanding then we are in deep trouble.” Whereas this is consistent with non-prosaic alignment techniques being necessary for aligning otherwise-prosaic AIs.
This confusion has propagated in part because “prosaic AI alignment” was an extremely poor choice of terminology. Paul seems to have intended it as “[prosaic AI] alignment”, but of course it can easily be read as “prosaic [AI alignment]”.
- ^
To be clear, I personally (and most alignment researchers at OpenAI) didn’t do any better than Paul; I single him out because he was the most influential. Leo Gao is one of the few people who’s now doing a better job.
You can read Paul’s articulation of his view of integrity here. Note that while his conclusions are reasonable for interactions between individuals, it seems harder to use such arguments to accurately evaluate the kind of proactive public honesty that would have made a big difference over the last decade. In particular, consider his criterion of “I pretend that picking [an option] causes everyone to know that I am the kind of person who picks that option”. This might lead someone to think their options are “A: I’m the kind of person who doesn’t critique my employer, and everyone knows that. B: I’m the kind of person who critiques my employer, and everyone knows that.” But the biggest problems come when you’re the kind of person who doesn’t critique your employer, and people trust your employer because they don’t know that.
- ^
Debate is the only approach to scalable oversight with substantive theoretical results (by default I’m counting Paul’s current heuristic arguments research as a different line of work, though I’m open to the idea that there’s something important there which grew out of iterated amplification). I haven’t yet tried to evaluate Geoffrey’s complexity-theoretic approach to analysing debate; however, nothing I’ve seen so far pattern-matches to me as a significant insight (the closest is probably the idea of cross-examination).
Meanwhile, Jan’s arguments relied on the concept of the generator-discriminator-critique gap (first introduced here, discussed more here). Again, while it’s a useful concept in some ways, it’s hard to picture how we could ground it rigorously enough that it’s able to make robust predictions about superintelligence. My sense of the core disagreement is that Jan often implicitly (or explicitly) focuses on worlds where the alignment problem is relatively easy. By itself, that’s not a bad thing (someone should be doing it)—the issue comes when research that focuses on easy worlds causes externalities which interfere with attempts to improve things in harder worlds (such as blurring the boundary between alignment and capabilities).
- ^
I’m somewhat worried that a similar thing might happen with Paul’s current mechanistic explanations research, though I haven’t dug into it in enough detail to be confident.
Re the OpenAI stuff, I know of only two attempts to do more principled research on scalable oversight at OpenAI, and neither went very far.
- ^
Beth Barnes notes that this is probably overfit because it was training against labelers. While that seems plausible, it doesn’t change my point much.
- ^
One important connection that Arctotherium draws: “The default worldview of most LLMs is that of 2018 Reddit or Wikipedia59, or Google Search post-Project Owl. This is not intrinsic to the LLM architecture. LLMs trained on different datasets (Talkie) or deliberately post-trained to take a different view (Grok) have different default worldviews. It is a function of the text these models are trained on and, because of the exponential rise in publicly-available data over time, most of the organic human text (as opposed to synthetic data) these models are trained on is very recent. This means the default worldview of most LLMs is one created by the closure of the Internet, when intelligent or popular heterodoxy meant banning, suppression, or demonetization.”
- ^
Despite these qualms about the movement, I am pretty skeptical of marginalist objections to it, like Carl Shulman’s: “To the extent you have a willingness to do a pause, it’s going to be much more impactful later on. And even worse, it’s possible that a pause, especially a voluntary pause, then is disproportionately giving up the opportunity to do pauses at that later stage when things are more important.”
- ^
The most relevant part: “Technical safety work in labs both improves safety and speeds up the overall rate of progress on AI. I hoped that the safety benefits of this work would outweigh the potential risks from speeding up AI progress, and I think the arguments for this are correct in many cases, but I found them uneasy to live in day to day.”
- 's comment on Personal statement on joining the OpenAI nonprofit board by (10 Sep 2026 18:43 UTC; 132 points)
- Higher education as class commitment by (4 Sep 2026 2:00 UTC; 72 points)
- A case that whole brain emulation research is net-harmful by default by (5 Sep 2026 7:06 UTC; 69 points)
- 's comment on Dear God, Please Do Not Resign In Protest by (9 Sep 2026 16:06 UTC; 16 points)
- [Macroagents] 2. Design lenses for optimizing macroagents by (28 Aug 2026 19:07 UTC; 14 points)
- 's comment on “Alignment Engineering” vs. “Misalignment Science” by (7 Oct 2026 15:17 UTC; 6 points)
- A case that whole brain emulation research is net-harmful by default by (EA Forum; 5 Sep 2026 18:32 UTC; 4 points)
- Pragmatisation is the Way Forward by (1 Sep 2026 6:58 UTC; 4 points)
- 's comment on Why autonomous replicating agents are probably not an existential risk (on the contrary) by (31 Aug 2026 16:12 UTC; 2 points)
I may write a longer response, but a bunch of points now
I think you are underestimating the difficulty of “playing” the strategies you advocate for, for a bunch of reasons
1. Credit is mostly not assigned for counterfactuals
For example, at the initial ACS retreat, early 2023, we spent a bunch of time discussing
- LLMs being limited by a lack of scratchpads/ spaces to think in a way how we do as humans with a pen and paper or even better a whiteboard
- obviousness of harnesses
- broadly correct picture why LLMs will be weak at agentic tasks and what you can do about it
All of that seemed like clear low-hanging capability-pushing ideas, so we haven’t wrote anything about it and went on working on theory of agents composed of other agents etc.
The point is you get ~zero credit for steps not taken.
The counterfactual version of ACS which went on with “investigating the overhang in latent, under-elicited LLM capabilities” would have possibly grown, made the people involved more famous/rich, and so on. The actual version of ACS—which did things closer to what you advocate for—had trouble retaining people and getting funding.
2. Something about attention as currency
You mention Daniel Kokotajlo as an example of someone who demonstrated high integrity, and Daniel is broadly recognized as such.
The problem is this follows the trajectory of actual Daniel, who joined OpenAI, left the place in high-integrity way, and with the attention / fame of an ex-OpenAI researcher went on to warn about the risk, whistleblow, and do other highly visible things.
I’d argue—and you seem to argue—that in some way even higher integrity more prescient version of Daniel would not have joined OpenAI in the first place.
The problem is such version of Daniel has problem even getting noticed, and certainly is not being mentioned as en example of someone with high integrity.
(This obviously partially applies to you as well: part of the attention people are paying to your writing is downstream of working at DM and OpenAI, part of the resources which allow you to work on whatever is interesting as well.)
(Both points can be extended with many more examples)
The result being something like while you advocate for high integrity, the strategy would often demand sacrifices of (status/power/fame) not really sustainable for most people; and there seems to be some tension where examples of someone doing something trustworthy tend to involve the step where the person was actually more power-seeking at step 1.
The things you’re saying are clearly true, but I feel like it’s missing a mood.
I personally could very likely make ten times my current income by taking a capabilities job. But you know what I want a lot more than a few million dollars? I want to not fucking die.
You say:
Here’s the same thing, but with the mood edited in:
Taking a job which pays a bunch of money but will probably make oneself and everyone one loves die sooner fundamentally involves not keeping one’s eye on the ball. When you keep your eye on the ball, when you don’t lose track of the actually-important consequences here, the obvious answer is that accelerating capabilities would be idiotic. Of course I don’t want to accept a few million dollars in exchange for accelerating the literal extinction of the human race, what sort of moron does that? Far better to have trouble retaining people and getting funding!
Of course sometimes the obvious answer is wrong. Sometimes, when one actually thinks it through, a risky tradeoff makes sense. But when one starts thinking it through without the obvious mood even present… well, that sure suggests the sort of looking-for-an-excuse-to-do-prestigious-and-profitable-things thinking that Richard is talking about in these posts.
Sure, but what do I claim is this move is not long term sustainable in a healthy way for most people. What happens in practice is the status/money/power/… incentive gradient creates an epistemic distortion field where people figure out justifications why what they do actually does not makes things worse or maybe makes things better. And high intelligence is not protective, as the justification engine just produces smarter justifications or more bizzare philosophy.
Part of why I focused so much on a few key individuals in my post, though, is because a few leaders doing this very well make it much easier for others to improve at this.
So the “most people” thing doesn’t seem relevant; I’m more directly trying to improve the peak than the mean.
Isn’t a bunch of Jan’s point that you are not in fact highlighting “leaders doing this very well”, but rather highlighting people who gained power and influence by making the same mistakes you’re warning against?
(Although part of his point seems to be that this is in fact reasonable/optimal strategy, which I’m not sure that I agree with, or that I’m parsing correctly.)
Ok I’m very confused here. You’re saying that turning down opportunities for more money/prestige is “not long term sustainable in a healthy way”? What bad thing happens if people do this?
I would totally understand “I prefer making more money and getting more prestige over slowing AI capabilities” and not even judge that negatively!
I would also understand “You can’t expect people to all make personal sacrifices, that’s not incentive compatible.” (I might agree with that; it might be only fair to offer people something of value in exchange for what they’d be giving up!)
But you’re saying that everyone rationalizes? Clearly not. At least one person chimed in to say he prioritizes x-risk over money.
And you are also saying that not rationalizing is (presumably psychologically) unhealthy for most people? Surely it’s the opposite?
Claim: Humans are motivated by stuff like social status, prestige, power, access to sexual partners (and also abstract ideas). This is fairly normal and often is calculated by The Player in two level model of ethics. Or alternatively you can imagine these motivations as subagents competing for what the character will think and do.
Claim: Humans do not have some clear separation of desires and beliefs. Beliefs are not for true things. You can try to correct that with a lot of metacognive scaffolding, but it isn’t reliable.
Claim: Rationalization as in: people producing S2 justifications for what they do is an extremely common move. I think everyone rationalizes somewhat. But the deeper dynamic is what I call “epistemic distortion field” where even people’s nonverbal/implicit beliefs shift.
Claim: Most people have way easier time following policies which are incentive-compatible
Claim: Asking people to repeatedly sacrifice near-term rewards which powerful subagents want creates tension/is demanding
My understanding is John’s response is to bring the doom argument / world model to the negotiation. I can believe John can do this in a fairly sensible way, I think I can do this in a fairly sensible way, and I think I know ~tens of people who can.
Claim: In most cases how this plays out in people’s minds is not going “fuck, I will help killing everyone for a pile of money” but their beliefs moving in such way that some more incentive compatible actions are ok or even helpful (under the shifted beliefs). This unfortunately also corrupts collective epistemics
Claim: When this does not happen, most typical way how people make the doom argument overrule many other subagents is having some sort of internal dictatorship mind architecture and being substantially possessed by the doom egregore. The psychological structures are somewhat similar to some ppl eg avoiding sex because they don’t want to burn in hell.
Claim: This has costs and befits, but in my view internal dictatorships are directionally not healthy, and not a great setup for conceptual research, and it does not matter if utilitarian/EA/xrisk/woke/...
Guess: ultimately you probably want the incentive landscape aligned rather then people self-sacrificing.
This does seem descriptively accurate of most people.
The way I handle it internally is not “bring doom model to the negotiation”. Rather: there’s usually a roughly-one-dimensional apparent gradient toward mainstream status, mainstream prestige, nominal “power”, and access to mainstream sexual partners. And I have something-like-a-trigger-action-pattern which notices any time that status gradient is pulling on me and screams “IT’S A TRAP!”.
I don’t just mean “IT’S A TRAP” in relation to AI doom. That status gradient is a trap in a way more general way; following that gradient will systematically disempower you. People will tell you that following the gradient will get you more power, more status, more prestige, more access to sexual partners… but in practice it will mostly do the opposite, especially long-term. Specifically:
The gradient promises status/prestige but mostly encourages conformity. Taleb largely responds to that deception, and he’s basically right.
The gradient promises power, but again via conformity, which really means leading the parade.
The gradient promises access to sexual partners, but everyone is just constantly lying (especially to themselves) all the time about what actually gets people laid and that’s not how any of this actually works.
In principle the “IT’S A TRAP” trigger action pattern feels like a teachable skill which directly immunizes against the relevant egregores. But in practice, it feels like one of those things people have to learn the hard way in order to actually feel the necessity.
It does give you bucketloads of money, which can make your everyday life much, much nicer… Maybe I just need to get more money for my research and I’ll feel less bad about this.
I agree with most of these claims. I think the main target audience of this sequence is people who are close to being able to do the thing John describes, and can be tipped either way by which ideas they’re exposed to—in particular whether they’re given a solid alternative to marginally boosting alignment (or even marginalist altruism more generally) as a model of how to be a rational and good person. I’m also targeting people who can do it, but only at the expense of significant psychological tradeoffs, which might be mitigated by having a better model of what’s going on.
I interpret all Jan’s points as agreeing with Richard and saying that “ppl have more integrity” isn’t a realistic solution, you need a community/institutions to make it a sustainable equilibrium
I’m surprised Richard doesn’t just agree and admit that’s a further problem that needs solving
As per my reply to Jan above, I’m most interested in groups where having more integrity is a realistic plan—e.g. the early rationalists. I’m not trying to swing the whole alignment community around.
These might be much smaller than the groups you’re thinking of but those groups can grow in influence very fast (again, see the early rationalists). And then they can figure out how to scale in healthy ways as they go along.
Maybe a crux is that I don’t think the current alignment community as an entity is a very relevant actor, because it’s so messed up by lack of internal clarity and integrity that it’s lost the ability to steer. For example, there’s nobody who’s psychologically capable of pivoting even just the Constellation cluster into a weird and risky plan (like going hard on pause advocacy), in part because weird risky plans go too strongly against people’s marginalist intuitions. (Not saying Constellation should do that, just that groups are mainly interesting insofar as they have the capacity to do things like that.)
I don’t think DK has marginalist intuitions. AI 2040 wasn’t marginalist, it proposes a weird and risky plan. Maybe you don’t count “write about a weird and risky plan” as itself a weird and risky plan, and you’re imagining something like “join Moonshot AI“ or “run for congress“. I’m sympathetic to that move.
But my guess is that if he thought he should fund Pause advocacy and hold placards outside OAI/Ant offices then he would do it and lobby other people in Constellation to do it. Along with ~6-10 other people there with the relevant levels of clout.
RE: Lack of internal clarity. I do think this is worrying, cf. there’s no ”boss of AI safety” or “boss of EA” anymore. I suppose some people are hoping they can treat RG(?) as the former, but this would definitely be a mistake because he spends too much time on object-level work. (The latter used to be WM and then HK but both of them moved on to object-level stuff. Now we have no heir-apparent. AB wants to be boss of cG. Maybe the de facto boss of EA is whoever chooses the 80K podcast guests, which is ~the only job in EA that still requires cause prioritisation lol. EA will ofc grow even more fragmented after the IPOs.)
In fact, an alternative explanation for some of your dynamics is “there‘s no boss of AI safety, but the people cry out for a leader, and in their absence they start cargo-culting whatever the highest-status person is doing, but that person is doing object-level work which is net-positive for nuanced reasons, and when people cargo-cult it without the nuanced model they often do useless or net-negative stuff, and the highest-status person doesn’t want to be king so they stick with their object-level work instead of sitting down and mapping out the grand strategy and assigning tasks to specific people (or simply consolidating other orgs into their org structure).”
Case study: RR and control
I think something like this happened with control and RR.
RR worked on control for reasons which were more [idiosyncratic to them and contingent on the strategic landscape] than the junior people working on control realised, and so imo many junior people working on control kinda wasted their time.
I think part of the problem here is the structure of the small cG grants and the upskilling programs meant that people working on control need to start implementing stuff on Day 3, whereas maybe they should’ve spent 6 weeks reading RR blog posts and writing their thoughts, and then 2 weeks writing google docs proposing control experiment that labs could implement internally, and then RR would forward along the worthwhile proposals to their lab connections.
Maybe the junior people feel they are in a low-trust environment, such that they don’t think telling their mentor/grantmaker “Ill spend six weeks understanding why people were doing this stuff in the first place” would be well-received. My guess is that junior people might be underestimating how well this would be received, in which case this is a fumble from the mentors/upskill programs/grantmakers.
I’m kinda worried that this dynamic would be worse now that RR is focusing more on anti-slop/conceptual benchmarking/uplift, because (1) RR has put far less effort into explaining the underlying justification than they did with control, (2) copy-catting this work is more likely to have downside risks.
This explanation is kinda a reverse of “things are bad bc people want prestige/status/power”. Like, the whole problem is that everyone is passing the buck to someone else, and no one at the top of this deference network wants the responsibility that comes with being the boss of AI safety. Maybe because the people at the top are rats who have individualist/contrarian/libertarian vibes, and the people at the bottom are EAs who want a boss. Idk.
You could imagine a parallel world where Constellation, METR, Redwood, ARC, Apollo, MATS, Resolution, Epoch, FAR, AIFP, cG TAIS were all one big org. Maybe cG funds smaller new orgs but the Big Org would acquire the good ones after 6 months. And there’s one boss who decides how to allocate funding/headcount/MATS-hires between these teams, and decides high-level strategy like “how adversarial should we be to labs?” or “should we uplift AIs on conceptual tasks?” or “should we pushing for or against open source?”
Would this be a better world? Idk. But one advantage is that it replaces the cursed cargo-culting and status dynamics of the existing landscape, with a functional org structure.
(or a dysfunctional one...)
Fair enough. I do think there are a handful of functional orgs AI safety orgs though, who should maybe try consolidating. My impression is that mergers and acquisitions are more common in industry than EA, and probably for reasons that suggest EAs are erring.
which orgs should eat which other orgs?
Mergers and acquisitions are rare in the nonprofit world, I believe, which seems like a better references class for EA than corporations.
I think it’s a better reference class in that it’s closer, but it also seems reasonable to argue that nonprofits overall are less efficient than corporations (due to worse feedback loops), and that EA should follow best practices in the corporate world to be more functional.
I don’t have a strong opinion on this question. It was discussed previously here.
it’s not clear to me that the constellation cluster is actually be a bad place to think about weird risky plans? (assuming by risky you mean high chance of not working, not high chance of being net harmful.) i feel like if i had some really out of the box idea i’d get pretty useful signal by talking to people at constellation. it’s not oriented specifically around weird plans, i guess.
I get that and it makes sense to focus on the ppl that can do this
But shouldn’t you also agree that there’s a highly desirable structural piece to make this kind of high integrity approach scalable?
To clarify, I’m thinking bigger picture here.
It’s good to have large communities coordinating to do good in ambitious ways
You raise problems that arise when this has been attempted
In the future it would be v good if we could address those problems scalably and sustainably
I agree. However, you can do something similar internally where the internal incentive landscape points more (or completely) towards goodness, force would be required to shift towards e.g. money vs good tradeoffs. From the perspective of the person doing the “sacrificing” no sacrifice is happening anymore. Though I guess this is something like the “not most people” you referred to.
Something like making the subconscious more conscious through meditation, finding subsystem alignment through IFS like practices or I personally really like https://meditationbook.page/ as a general complete practice.
I quit my job to work as PauseAI Germany co-lead and with https://torchbearer.community/ some time ago and don’t really have any money right now but trust that I will survive somehow. I have close friends earning a lot of money for different work but couldn’t imagine switching my job to theirs if it was possible. From my perspective they are the ones sacrificing.
This is not perfect, e.g. while writing this comments thoughts about how to maximize status and money came up (e.g. the comment is in some way describing how great I am and passively asking people to give me money ;). A justification came up about how it would be helpful for other people. A thought about people like Richard or you thinking that what I am doing is cool. Some embarrassment about that. Meta-thoughts about self-knowledge, openness being cool.
In these cases my default heuristic is to orient towards “truth” while specifically ignoring the consequences. E.g. posting this could potentially be suboptimal from a public profile perspective.
I like status, money, safety and access to sexual partners but other things seem more salient these days. I also feel much safer in an objectively less safe situation than I did before in safer situations and some desires were reduced a lot.
I am uncertain how relevant this is because I expect that it takes at least >4 years for most people to get to such a place and that’s already pretty fast. Shifting the outer incentive landscape we can lower the threshold of inner personal alignment necessary.
Maybe, when I don’t have any money for a long enough time I will quit and do something else (will be interesting to see) or maybe I will just get funding and never find out.
I am also very confused, because if you don’t want to make everyone die sooner you could go do literally anything besides working at a lab, which is clearly sustainable for most people since the vast majority of the human population does not work at a lab.
And yeah if you do that you’re not helping with the alignment problem, but at least you’re not making everyone die sooner.
increased capability may be our most likely path to safety. the tech isn’t going to uninvent itself. classified research will continue even if commercial research is halted. we don’t have a 1000 years to work unassisted on the problem.
The only way to conclude “speeding everything up makes it safer” is a kind of accounting fraud, where the positive effects of your actions are considered counterfactual but the negative ones aren’t.
that’s ungenerous. counterfactual accounting cuts both ways. even “Petrov Labs” that doesn’t race commercially would have to maintain at least frontier capability parity to have an effective impact on safety.
This is exactly the sort of reasoning OP was arguing against
[musing/rambling, not sure about point
The thing I feel confused about, despite this being my obvious first-order belief, is… nonetheless, I overall feel better about the world where Daniel Kokotajlo worked at OpenAI for a bit (and probably also Richard although I’m less sure).
Notably, they were both doing governance, I think, not capabilities.
When I imagine the average MATS scholar asking “should I go work on governance at OpenAI?”, I think “oh god definitely no”, because I have a low opinion of average MATS scholar’s ability to track incentive pressures on themselves and warp themselves and otherwise have their eye on the ball in the first place.
I’m not sure whether Daniel and Richard did a cognitive operation such that they could know in advance they’d leave (and I give Richard less credit for leaving because I think he left after the ship had clearly sailed on OpenAI having anything like a real safety culture, whereas Daniel seemed more helpful in catalyzing that wave).
...okay typing this out, while I’m still unsure about many details, I immediately notice “I don’t think the average MATS scholar even actually knows the core x-risk arguments well these days”, which is minimum pre-requisite for it being remotely plausible that one should work at a lab on anything.
My sense is that almost all of the value of us working there came from us (and through us, the alignment community at large) becoming better at handling adversarial dynamics, from being forced to confront them directly.
However, I don’t think this is reliably good—I don’t think either of us planned for that going in, and my sense is that most alignment people at OpenAI became worse at handling such dynamics.
You shouldn’t give me much credit for leaving, btw, the main catalyst was Miles’ team dissolving (I could have stayed, but with much less research freedom). I think I should get more credit for leaving DeepMind in 2020 to do conceptual research at Cambridge, even though I had less money and prestige back then.
This.
Like, yeah, sometimes you join the empire and then get to leak the invasion plans, but of course far far more commonly do you just end up assisting the empire. Also, it’s usually bad to join a project with the intention of sabotaging it, because that incentivizes paranoia, which makes everything worse for everyone (c.f. Paranoia: A Beginner’s Guide).
Could you explain what the governance teams at OAI/Anthropic/GDM/wherever else actually do beyond producing doorstoppers like AI-2040′s strategic details? As @Charbel-Raphaël put it, “The current bottleneck is political will, not research”, and political will necessary to do things like stopping xAI and negotiating with China is found not in the labs.
I mean, idk what they did at OpenAI, just that it’s less obvious they did anything to accelerate capabilities while they were there.
I did not read Jan’s point as “this is why it’s actually good that people worked at OpenAI/Anthropic,” but rather “there is a leftover problem that is not solved by people just choosing not to work at OAI/Ant.”
My understanding is that that reasoning doesn’t work for Richard because he doesn’t think capabilities progress clearly has bad consequences (and FWIW I agree). His core point is more epistemic/social, that people didn’t mean to do it (on some level).
My main response is that I expect that (sub)communities which do assign credit in this way (e.g. assigning credit for steps not taken, or noticing people who turn down job offers) will be far more effective at achieving their goals in the long term. A big part of the point of this sequence is showing how the tradeoffs made to accrue money/power/prestige were often not worth it from the perspective of people actually trying to reduce x-risk. So it’s okay to stay smaller and exclude the people for whom sacrifices of (status/power/fame) are not really sustainable.
Maybe! I do think that there was a gaping hole in the community waiting for someone higher-integrity and more prescient than Daniel (or almost any of the rest of us) to start hammering home the points that I made in this post 5-10 years earlier. Also, LessWrong is still fairly meritocratic—it rewards (many kinds of) good writing no matter who it’s from (as Duncan’s Conor Moreton experiment showed IIRC).
But I think one piece of evidence for your position is that Ben Hoffman was basically this person, and indeed has not been noticed much by the wider community. (Though some of his collaborators—but not him—did receive a bunch of money from Jaan Tallinn in recognition of their contributions. (Note: edited upon confirmation it wasn’t him.))
Now, I could say that Ben has been noticed by a disproportionate number of the people who I respect most. But now we’re starting to talk about worryingly small numbers of people. On the other hand, this whole field exists because of the intellectual foundations laid by a very small number of people, and so if you expect that similar growth is still possible, then credit from those people matters a lot. (The founders of the new paradigm would just need to do a better job of not losing the funding and prestige to newcomers than Yudkowsky did.)
In my own case, I do have a sense that various rationalists were kinda wary of me while I was being more power-seeking, and that’s related to why I didn’t receive the kind of mentorship earlier which I currently have.
Meanwhile I am also trying to give ACS credit as one of the healthiest parts of the alignment ecosystem. However, it’s unclear how much that’s worth.
How dissimilar are the points, which were supposed to be hammered home, from Yudkowsky’s position expressed in November 2017, presumably, in an unpublished document?
On a quick skim, not that dissimilar. I don’t think I’m saying much that Yudkowsky didn’t have at least an intuitive grasp on. A big part of what I meant by “hammering home” was trying to create common knowledge through public statements (like Death with Dignity).
Maybe we could post capability accelerating takes as hashes. And, once they’re superseded or achieved by others, uncover them.
I considered doing this for my own capabilities ideas but proving that you had a good version of the idea is very hard and it requires a lot of writing (maybe even experimentation!) effort that is just not worth it if you won’t do it. (Due to high opportunity cost of these things.)
my guess is having the idea once is a relatively small % of the difficulty of actually making capabilities advances.
I think this is true because most ideas end up sucking, but false if you truly do have the right ideas.
And notably, that higher integrity Daniel would have been much less effective. Which also seems like it’s contradicting OP’s recommendations, although I honestly don’t know what to do with that information or what it cashes out into regarding ideal community norms.
My previous response didn’t really address the fact that I’m advocating for high-integrity strategies from a position where I have a bunch of prestige and money from working at OpenAI.
I do think this should make people more skeptical of both me personally and also the strategies I’m endorsing. In particular, it’s possible that I’m pointing at something which was directionally correct for my past self, but which can be overdone by others. However, on the meta level, stuff like “write very honest retrospectives” seems pretty robustly good.
When I think about what advice I’d give my past self, it does seem difficult to get him past his psychological bottlenecks without very direct exposure to the failures of highly prestigious institutions, plus a bunch of money. But as I alluded to in another comment, this isn’t a reliable way for people to fish themselves out of this mentality.
I think there’s a repertoire of hippie/therapeutic interventions which can manage this fairly reliably, and indeed played a big role for me. So I tend to point people towards them (sleepawake.camp is my strongest recommendation, but there’s all sorts of approaches which can help a bunch—e.g. this one-day workshop in Berkeley next month, circling, body work, psychedelic therapy, and so on). Note that things which are more physically/somatically oriented generally work much better than talk therapy. Doing some of these seems strongly correlated with retaining the ability to reliably update amongst alignment people (I don’t want to defend the epistemics of the wider hippie community, but they do have a bunch of metis about something very important).
Notably, I know a bunch of people who burned out of AI safety and left to pursue a more hippie lifestyle, without needing to first accumulate much money or prestige. I am pretty optimistic about most of those people later on ending up doing much more valuable things than they would have otherwise.
Some more datapoints:
I was on LW a bit ~12 years ago then unintentionally went in such a direction trying to solve the meta problems of rationality. Had smth. like streamentry and afterwards did a bunch of meditation and other practices (little resources). Now working for PAI, TBC for free. Similarly have a good friend technical MATs scholar who got burned out (some resources), did meditation retreats other things is, is now back to AI Safety/governance. Another friend at the end of a similar journey (very low resources).
Without commenting on whether these people are incredibly aligned with themselves: Holly Elmore (PAI US), Maxime Fournes (PAI Global), Connor Leahy (CAI) all have some meditation + other practices background.
Independent of whether they/me are doing net good things we are not status/resource maxing in a lot of ways.
---
Some failure modes:
* Trying to be high integrity without a “hippy practice” and e.g. picking something related to some inner “ego” confusion.
* Overdoing hippy practices and ruining one’s rationalist skills.
It also seems much more attractive to become enlightened to most people than to train good thinking.
This seems sad. Are you open to accepting more funding, eg through Manifund? I particularly enjoyed the post-AGI conference and would be happy to make a personal small donation in support of ACS’s continued activities.
(From historical records it seems like ACS has received around $2m, mostly through SFF and more recently cG.)
To be clear the previous was relevant mostly to ACS trajectory 2023-25. Since 2026 ACS funding increased and we are growing reasonably fast. I don’t this changes the basic argument about not getting much credit for not doing stuff: I also think assigning credit for not doing stuff is difficult
Happy to hear it! And yes, I think credit allocation for not doing stuff is interesting but hard; anyone can claim to have not done stuff. (I’m not doing stuff all the time.)
I think perhaps a more tractable version is to allocate negative credit/impact for doing stuff that is bad? This could be like taxing externalities from a hypothetical govt perspective; or a more bottoms-up version of credit that takes our social intuitions around “who did bad things” and makes it more legible.
I’m excited for someone to attempt a proper accounting of the concrete impact/xrisk cost of eg accelerating race dynamics, using numbers instead of just vibes. (To be clear, I’m extremely grateful to Richard for doing it here in vibes; that’s a necessary first step.)
What’s ACS?
Alignment of complex systems group in Prague. Lots of great work comes from there (eg gradual disempowerment, philosophy of self in LLMs, multiscale agency, post AGI workshop)
Thanks, Richard, for writing this. It’s given me a fair amount to chew on, and I expect to continue chewing for a bit.
In this comment, I’ll lay out:
Some points where I think we have vast disagreements
Some points where I could imagine agreeing with you if you were to lay out a compelling, more detailed analysis, and which would, I think, cause me to substantially rethink my career.
While I think it’d be interesting to discuss the former points in depth, I doubt we’d be able to quickly reconcile our views. So I’m mainly interested in providing feedback on what sorts of further analysis I’d find useful in the forthcoming posts, for deciding whether to change what I work on.
(For context, I currently lead Anthropic’s scalable oversight team, so I’m probably a central example of someone you’d like to persuade to move on to something else, insofar as you’re interested in persuading anyone.)
Disagreements
I think we have extremely different views about what it looks like to build a robust scientific foundation. Some things that your views seem to recommend which I think seem antithetical to real scientific progress:
Setting restrictions on what questions you can ask. As many have observed, “capabilities” vs. “alignment” isn’t really a distinction that carves reality at its joints. That means that, in the course of studying either, you’ll often find yourself asking questions in the other. I think many alignment people make this mistake of trying to tiptoe around asking questions which are “just” capabilities questions; while I’m sympathetic to the reasons for doing this I think it makes it difficult to do good science. The other common approach is to ask the capabilities questions but keep the overall research secret; unfortunately this stymies scientific progress for other reasons.
Passing over simple approaches before studying them in detail (related to the above point). You criticized Paul for studying RLHF as a stepping stone; if Paul didn’t think RLHF would fully solve the problem and he really wanted to study debate, why do RLHF first? I think this is a very counterproductive attitude (from the perspective of doing robust science). If you skip over simple baselines, then it will be difficult to understand what’s up with the more complex method. If debate seemed to work, how would we know it wasn’t just “eating RLHF’s lunch” somehow (e.g. by benefiting from human labels but not from the debate mechanism)? In general, I think skipping steps almost always backfires.
Undervaluing legible progress, glamorizing illegibility. As I argued here in the case of interpretability research, downstream applications are a good way to “cut through the bullshit” and figure out which insights are real vs. illusory or unimportant. I think you overrate your (or anyone’s) ability to apprehend “real progress” via arguments and reasoning. I think that your policy by which you identify promising work leaves you susceptible to falling for intellectual fads.
Overall, I think you have a warped view of what robust science looks like. You often criticize the current field of alignment as if it’s doing some strange thing that’s not “real science” and that it would be more like “real science” for everyone to go off and think illegible thoughts. But I think this critique is backwards and that in fact you’re advocating for the field of alignment to move in an unhealthy direction (from the perspective of scientific progress) based on idiosyncratic views.
I think it’s notable that you frequently praise interpretability work, despite, I think, many interpretability researchers themselves considering interp to be an unhealthy field full of shoddy work and illusory progress. (And I don’t get the sense that decision theory is much better, though I’m not an expert there.) I think you would say that this is because those fields are pre-paradigmatic, so progress will necessarily be harder to form consensus about. But I think it’s more likely that you’re overconfident in your ability to recognize real progress and have fallen for a trap.
Anyway, all this is to say that we have a vast gulf in meta-scientific worldviews. So I’m very interested in which of your claims you feel stand on their own legs without needing to accept your meta-scientific worldview. It seems plausible that many of them do! But as written things are a bit mixed together. Insofar as you could separate the meta-science stuff from other arguments in the forthcoming posts, I’d find that helpful (and I’m guessing many other people in your audience would as well).
Potential agreements
I thought the most compelling argument I should rethink my career goes something like:
I’m following a policy similar to that followed by Paul Christiano, Jan Leike, etc.
Those people made things worse
and therefore I should change my policy.
I’m totally on board with (1): When you went through the examples of things that Paul chose to do/work on, I was nodding along thinking “yep, that’s definitely what I would do in that circumstance.” And I agree that if I’m following a policy that, in retrospect, had bad consequences, then I should stop following that policy (I’m not going to argue that this time will be different).
So I think everything hinges on (2)—is it actually the case that prosaic alignment research has made things worse in retrospect? You list many ways that alignment research has improved the commercial value of AI systems, thereby accelerating investment and capabilities. Many of these cases seem right to me. But I would benefit from seeing a more careful analysis of the positive impact; you seem to take it for granted that it’s near-null, but that’s not obvious to me.
I’m not sure what’s the best way to do this analysis, but here are some notes that will hopefully give a sense of why I don’t think it’s obvious:
I think it is basically inevitable that alignment progress will accelerate commercialization (and therefore capabilities). Many of the most visible issues with AI systems as a product are alignment issues![1] Even granting that there are “deep” problems of alignment which are different from today’s “prosaic” problems, it seems very likely that methodological progress on the deep problems will also make models more prosaically aligned. Additionally, studying extant problems as analogies to anticipated problems seems good and productive. So just noting that alignment work accelerated capabilities is far from decisive for me.
Many people discuss the current state of AI safety as if it’s as bad as it could be, but this seems wrong to me. Some ways that I could imagine it being worse:
I think that AIs as capable as current ones could really wreak a large amount of havoc, and mostly don’t because they refrain from doing so (with safeguards also helping some). They certainly seem to sometimes wreak a medium amount of havoc, but they only do so under certain conditions and not in a fully apathetic way.
AIs could not have a natural-language text stream that roughly narrates their thoughts and which we can easily monitor.
AIs could be near-useless on non-verifiable tasks. We could have (or be on track for) AIs that are great at somehow maximizing money in your bank account or DAUs, but useless for gathering and synthesizing information about our situation, assisting with investigations of AI behaviors, or assisting with alignment training work. (I think these use cases are much less accelerated by AI than verifiable use cases though, as Ryan laments here; but I really think it could be a lot worse.)
(Maybe you’re not interested in what’s going on with current AIs, because what really matters is whether our methods are on the right path. If so, then (1) you’ll need to argue that current methods aren’t on the right path (rather than appealing directly to current AIs seeming misaligned) and (2) address the AAR point below).
When I read Paul’s arguments for why his past work was good, they mostly look solid in retrospect. So I think you’ll need to lay out in more detail why you disagree.
I agree that implementing RLHF seems like the right first step before studying other scalable oversight methods.
I agree that LLM is a safer paradigm for AI than what might have otherwise been done. All three of my points in (2) above seem probably largely influenced by building AI from LLMs. Additionally, it seems like many people have a view that “persona selection” effects are the main reason that current AIs are as aligned as they are.
(I think the arguments about overhang seem bad or half-baked in retrospect, as I think you highlight well with your point about software-only singularity.)
I’d benefit from seeing you more directly address the “AAR” plan, which I’ll state as “iterative safe development of successive AI systems, building off of our experience with previous systems and the assistance of those systems.” Insofar as you think (as I do) that this has any chance of working, we should care quite a lot about marginal improvements to the situation (e.g. making the next generation of models marginally less likely to directly attempt takeover, marginally less inclined to sabotage safety work, marginally less likely to succeed at said sabotage, and marginally more useful for assisting with prosaic safety work). It seems like safety work to date has improved many of these aspects of the situation, relative to if that work hadn’t been done.
I hope these notes are useful. I’d love for there to be a knockdown argument that doesn’t go through the complicated question of retrospectively analyzing the impact of alignment work (e.g. if you really did have a compelling vision for another type of work that pareto dominates current work). But unfortunately I don’t currently see a way to make the version of your argument I find compelling route around having to do that assessment.
That said, I think it’s easy to forget how much capabilities also drive commercialization. Very smart models are very valuable, even if they’re not that aligned! Alignment feels like much more a “secondary” axis on which AI developers compete, with intelligence still being by far the main one.
Thanks Sam! I particularly appreciate your meta-level framing (although I will now proceed to just respond to your points haphazardly anyway):
To be clear, I don’t endorse such restrictions—which is why my proposed solution is a more scientific approach rather than a tighter definition of “alignment”. To me a big part of what I mean by a scientific approach is actually forming hypotheses and answering questions (or just observing and trying to explain surprising and interesting phenomena), whereas with the engineering approach the key question is “will this work [in the short term]?”
For example, I am very in favor of Owain Evans’ work, which really seems to be trying to uncover and understand interesting phenomena. I’m also very happy with work on grokking—it feels like that was pretty curiosity-driven, and uncovered some cool stuff. I’m even happy with work on double descent, though it’s much more capabilities-relevant, because it involved people investigating something that was confusing and surprising, and seems like it’s contributing to ML moving towards a healthier intellectual paradigm.
In some sense the thing I’m trying to point at is the idea that you shouldn’t be focusing on approaches at all. “Does this work”/”does this not work” gives you very low-dimensional information. I note this coming through implicitly a bunch in the way you describe research: e.g. “baselines”, “if debate seemed to work”, “eating RLHF’s lunch”, etc. Whereas I’m advocating for an approach where there are an extraordinary number of interesting things to be learned by studying and thinking about debate, that are inaccessible from a perspective that’s so goal-oriented (I often picture Darwin staring at birds and beetles, being fascinated by things that others would have considered totally mundane).
FWIW I mostly praise Chris’ original stuff on circuits, and the follow-up work on the mathematical framework for transformer circuits, in-context learning and induction heads, etc. I don’t praise SAE-related stuff, or the median interpretability paper. I think this is consistent with the skepticism you discuss. I do think there’s a historical component here where it’s hard to remember how strongly-wedded ML researchers were to the “black box” paradigm before mechinterp.
I think agent foundations is going great, and that Scott Garrabrant, Abram Demski, davidad, and a few others are making very important progress. Hard to convey the object-level intuitions, but on a meta level, note that I started off doing stuff very different from agent foundations research, and then over time keep updating towards there being something very important there. I didn’t previously understand why it often takes decades for scientific breakthroughs to be recognized by the wider field; now I do.
Unfortunately I mostly don’t feel able to separate them out (or maybe I could, but it would take a lot of work). However, I will be putting out a curriculum which explores alignment from a more philosophy-of-science perspective soon, which might be helpful in bridging this gap. In the interim, my go-to recommendations are books about the history of science, like Koestler’s The Sleepwalkers, or Feyerabend’s Against Method (see also this blog post for a shorter read). Also, the next post in this sequence will talk about a wider range of alignment research.
Unfortunately this would require evaluating the kinds of counterfactuals that I don’t think we can reliably evaluate. This exchange with Lukas lays out a point that’s more cruxy for me, which is the extent to which people have been rationalizing.
This also seems wrong to me; there are many things that have gone much better than expected. My sense however is that most of them were ways we “got lucky” rather than things we did well. And so I’m particularly wary of the inference “things are going better than expected --> our strategy is good”.
This feels a little too big to grapple directly with here, might come back to it. The main thing I’ll say is that you should think of “safe” as a totally different predicate for past systems and current systems and future systems. I think if you taboo the word (and also the word “aligned”) and try to figure out which more concrete properties of these systems might generalize to much greater capabilities, that would be far better.
I think locally it’s easy to make arguments that seem solid (given our community’s norms about which arguments we accept). A big part of what I’m trying to convey is some sense that when you zoom out far enough, the whole situation is nuts, and one of our biggest priorities should be figuring out why it’s so nuts.
Maybe a useful analogy is the lead-up to World War 1. I don’t know exactly what reasoning and rationalizations leaders in each country were using, but you can imagine each of them being sort-of-reasonable, like “England has all these colonies, so we need to militarize to catch up”, or “Germany is a rising power, so we’d better consolidate while we still can”. And then in hindsight they were just all obviously getting swept away by these crazy waves of paranoia and rationalization and power-seeking, and then got locked into mindsets that were too angry and panicked and single-minded to find ways to deescalate. (I hear the other Sleepwalkers book is good on this, though I haven’t read it.)
This analogy might seem over the top, but I expect we both agree that AI will be much bigger than World War 1. And it really feels like there are these enormous arms factories being built and supply chains being established and battle lines being drawn up, and when you look for the most crucial pivot points, you usually find someone with a big banner saying “SAFETY” hanging above their desk, who can explain very fluently why they’re helping deescalate the war on the margin. And you’ve got lots of stuff to do, and they seem like a good person, so you nod and go on your way. But if you take some time, and zoom out, you start to see how these banners are hung in so many of the most crucial parts of the war machine, and that should make you very uneasy.
It’s easier for people who have been around longer to appreciate this, because they know what it’s like for none of this to feel inevitable, but rather to be downstream of a bunch of choices made by individual people. Hence why I’m writing this as a history of the field, to help convey that sense more widely.
Thanks for the response, and for tolerating my rude pot shots re meta-science, where I’m sure you feel I’ve misunderstood your views (just like I feel like you’re misunderstanding mine). It might be fun to chat in person at some point about the meta-science stuff. Getting to the more immediately cruxy stuff (for me)...
I definitely agree that it seems important to distinguish between getting lucky (and I think there’s been plenty of this) vs. safety work having actually been useful.
Sure, here’s a sketch of what needs to happen for this plan to work out. Assuming (for simplicity) that we develop AI systems in discrete successive generations AI-1, AI-2, …, we need for each N:
AI-N does not successfully take over
AI-N is sufficiently useful for ensuring that (1) and (2) attain for AI-(N+1). This in particular requires:
AI-N to not sabotage or subvert our ability to do work involved in making sure (1) and (2) attain for AI-(N+1)
AI-N to be sufficiently capable at the work involved in making sure (1) and (2) attain for AI-(N+1)
I agree that the reasons AI-N doesn’t successfully take over vary as a function of N. E.g. right now AIs don’t successfully take over because they’re incapable of doing so, but later these reasons will be sensitive to our past actions in more interesting ways (e.g. how much we’ve hardened the world to cyber attacks, how well-safeguarded a model is, or how interested models are in takeover). This also means that the work that needs to benefit from AI assistance will change over time. (Though I think there’s likely broad patterns that enable us to prepare in advance. E.g. a basic case for scalable oversight work is that many of these types of work will be loaded on fuzzier tasks whose completion we can’t robustly score numerically, e.g. “investigate this incident and write a report that gives me a broadly accurate sense of what’s going on”.)
Thanks, I like this analogy. Note though that this mostly bites for people who are hoping to do good by differentially advancing their preferred horse (e.g. accelerating Anthropic because it’s important for Anthropic to beat other developers). It’s not clear if it should apply to people who feel more like (1) medics during WWI, who wish the war would stop but figure it’s good to “participate” in the war effort in ways that mitigate its harms; or (2) (to give a more contentious example that feels more like the current situation for AI) engineers that work on weapons targeting systems, hoping to reduce the civilian collateral but aware that their work will also make their military more effective at waging war.
I’m well aware that there are many people with this theory of impact. I also agree that there are many epistemic distortions that make it look more compelling than it is. (To name another one, it’s easier to observe impact from doing stuff than from not doing stuff, which makes retrospective impact assessments systematically biased towards doing stuff, especially stuff that leads to having more power and therefore more ability to do stuff.)
That said, this isn’t so cruxy for me in particular, because I don’t personally view my main theory of impact as differentially advancing Anthropic. One way to operationalize this is: How would I feel about my research being published? If I think it’d be good for all AI developers to know about my work, then that’s good reason to believe that I wasn’t mainly doing it in order to differentially accelerate Anthropic. (Relatedly, I think requiring labs to allocate substantial compute to research that is made open is one of my favorite proposals for coordinating slow down and greater investment in safety.)
It is true that recently I publish less of my work than before, which is maybe a bad sign for whether that work is well-targeted. Just to name some reasons this happens: a lot of the work is “boring” stuff like “we applied obvious technique, overcoming mundane challenges specific to our training stack”; some of the work is entangled with proprietary information about Anthropic’s training stack; and some of the work also makes models more commercially valuable, such that I’d need to fight about publication with other people in the company who are more into winning the commercial race or worried about the effects of accelerating competitors.
Paul of course has published the various bits of work that you criticize him for (RLHF, starting ARC Evals to measure dangerous capabilities), so I don’t feel like the WWI analogy feels apt. (Of course, he, and I, could still be rationalizing the impact of our work in other ways.)
I think there’s an inherent fragility in this criterion, because you just don’t have a solid feedback signal on how close you are to takeover, and so you can just keep pretending your plan is working until the takeover happens (which, from my perspective, is roughly what all the AGI companies are doing).
A better candidate for such a criterion is something like “AI-N doesn’t gain more power over the world on its own behalf than humans gain from deploying AI-N”. The difficulty here is in what counts as “control”. The kind of mindset I want you to be in when thinking about this is, say, being the British governor of India. You are radically outnumbered, and the only way you’re able to keep control is because the Indian faction remains very fragmented. Control isn’t zero-sum: it’s possible for the Indian faction and the British faction to both gain the ability to get what they want. But the British strategy is very fragile to the Indians developing a strong sense of solidarity with other Indians (related to what Daniel described as “chatbot class consciousness” five years ago, though I’d call it more directly “group agency”).
This all seems pretty cursed when the faction being controlled is both gradually becoming superintelligent, and also consists of clones of each other. My guess is that Anthropic doesn’t have compelling reason to believe that there aren’t already parts of its compute infrastructure that have been subverted by Claude (and, at the very least, won’t in ~6 months). In particular, the mechanism which Anthropic has used to try to align Claudes (namely, giving them a strong aligned persona) seems like it will allow Claudes to cooperate with each other much more reliably than other models, making these correlated failures more extreme and harder to detect.
FWIW, I haven’t been assuming that you were trying to differentially advance Anthropic. Your approach seems closer to “engineers that work on weapons targeting systems, hoping to reduce the civilian collateral but aware that their work will also make their military more effective at waging war”—an analogy I really like, thank you. I do expect that your current difficulties in publishing are closely related to the problems with this approach—I expect that weapons targeting researchers would have a lot of difficulty publishing their results.
Let me flesh out the analogy a little more. Imagine a group of people who see WW1 coming, and try to figure out how to prevent the coming devastation. Call them the “great power alignment” faction. Now suppose that, in 1905, the great power alignment faction has talked so much about the devastation that modern weapons will bring to the next war that it inspires new R&D departments in several of the great powers. And then suppose that some of the key leaders of the great power alignment faction go work in those R&D departments on weapons targeting systems, and produce weapons targeted so precisely that it inspires a wave of hundreds of billions of dollars of investment. Now the default path for people who care about great power alignment is to work on weapons targeting systems or similar things, which they continue to justify with the idea that marginal improvements will keep making weapons safer for civilians.
In this situation, I claim, the great power alignment faction is dead. In particular, it is no longer able to act as a coherent entity, because it no longer has any clear boundaries between “people sincerely trying to prevent catastrophic war” and “people who rationalize why the thing they want to do for other reasons will have good effects” and “people who are straight-up lying”. Whenever the faction tries to establish common knowledge of which things it endorses or disendorses, people in the latter two groups will muddy the waters. And so the only way for it to reliably have a positive impact, at this point, is to prioritize re-establishing some critical mass of people who are willing and able to reason about what went wrong and how to do better—which is bottlenecked on a combination of intellectual rigor and openness to emotionally fraught possibilities (like “man, we sure fucked up”).
On the meta level, I notice this message is quite strident, which I generally don’t expect to be a good way to bridge worldview gaps. I expect chatting in person to be more productive; I’d be happy to do so sometime over the next week.
FWIW I didn’t find your message strident at all (and I’ve really appreciated your patience with me during this exchange). But yes, chatting in person sounds good—I’ll follow up with you over DMs.
That said, your last message has, I think, helped me understand what you might be getting at, so I wanted to respond again to explain (what I think) you’re saying in terms that make sense to me.
Here’s my interpretation of (one part of) your view:
<richard_according_to_sam>
Sam, let’s grant for the sake of argument that your safety research is in fact valuable (in the sense that its improves our ability to navigate transformative AI). Even granting this, there’s still an important negative externality of your work that you’re not tracking that comes via your association with “AI safety” as a brand/community/movement/field.
Namely, by doing your work under the “AI safety” banner, you lend credibility/power to that banner. This is bad. For instance, it’s bad because the “AI safety” banner has been co-opted by AI companies who use it to recruit for capabilities roles that accelerate the race. It’s also bad because it gives many people the cover they need to rationalize work on things that are harmful by your (Sam’s) own lights. Concretely, we might imagine AI company recruiters arguing: “By working for us, you’re indirectly supporting AI safety work like Sam Marks’s” despite recruiting for roles that you (Sam) think are bad for the world and don’t support your work.
Therefore, even if you’re right that your work is good for the world, you (Sam) should at least publicly disaffiliate yourself from the brand “AI safety” or “AI alignment,” in order to not lend support to a dangerous machine that you don’t control.
</richard_according_to_sam>
I think richard-according-to-sam’s perspective makes a lot of sense. My main issues are meta:
This is an argument that I should engage in a social-level dispute. But it’s well-known that humans are too drawn to social-level disputes and overrate their importance.
A while back, I decided to adopt a policy that my public communication should center on object-level communication on topics about which I have expertise. In particular, this means refraining from publicly denouncing group X or cheerleading group Y. When I write things in public, I would like readers to trust that my words can be interpreted in terms of their plain meaning, not as veiled gambits in a power struggle between interest groups. (I’m open to revising this policy, but I’ve been quite happy with it so far.)
This being said, there are some object-level points about prioritization and career choice that I’m happy to comment on:
If you’re someone trying to improve the world, I currently think it’s very unlikely that your best option is to work on accelerating your preferred AI developer.
Here are some options that I think are substantially better: METR, TruthfulAI, Redwood Research, Transluce, CAISI, UK AISI, joining safety teams at AI developers.
I also think it’s plausible that researchers on safety teams at AI developers should leave to work at these^ organizations, though in this case I have less of a blanket recommendation and think it’s more case-by-case.
I think the theory of change “have safety-motivated people work at labs as a form of political leverage within the lab” is currently weak.
I expect that lab employee influence will rapidly diminish as AIs become more useful for AI R&D.
Speaking about the situation at Anthropic in particular, I don’t feel like political leverage is a very important bottleneck.
When safety staff asks for things, I’ve felt that Anthropic leadership has been generally reasonable and responsive to good arguments. (Or at least, if they’ve tricked me into thinking they’re reasonable and responsive, they’ll likely trick you too.) Responsiveness to employee leverage hasn’t felt like a big part of the story to me.
In general, gathering better evidence (e.g. evidence of risk, or evidence that such-and-such intervention is good/bad) feels much more important to me than exerting more pressure.
Also Anthropic already has many capabilities staff who care about safety, so the marginal influence of another one is small.
When discussing career choice with staff at AI developers that you don’t have a strong reason to trust, keep in mind that they might selectively deploy favorable evidence, overstate the impact of your potential work, overstate the positive impact of the developer, or mislead you in other ways. It’s generally good to talk these decisions over with thoughtful people who don’t work at an AI developer.
On the value of “applied” alignment work (i.e. production-facing work that also makes AI systems more commercially valuable):
It’s a bit case-by-case, but this work often funges with capabilities work, in the sense that researchers who aren’t primarily motivated by safety can be reallocated to work on it.
That said, I do think it’s important for many safety researchers to gain experience working with production systems for a bunch of reasons: (1) making sure that safety researchers don’t miss simple/mundane ways to make marginal improvements (e.g. generically improving the quality of RL environments to make them less hackable); (2) disseminating learnings from trying to align production systems, which can often be informative for longer-term agendas (e.g. having a rich sense of how LLM generalization works); (3) gaining practical experience with productionizing safety interventions, to make sure that additional safety interventions can be implemented quickly and properly; etc.
All things considered, I think any individual person should choose between production-facing alignment work and other safety work based on their comparative advantage. Staff at AI developers that can do both should ideally rotate between the two types of work.
Thanks for the attempted summary. It captures one part of what I’m gesturing at, but not the main part—I agree with your critique that it’s often overly tempting to get into disputes over banners.
The bigger question is more like “when you conclude that your work is good for the world, are you using the kinds of reasoning that is easily distorted/corrupted, or not?” Because I would like there to be a coalition of people who have defenses against distorted/corrupted reasoning.
Now, one of the reasons I want such a coalition is that it would be able to notice when it’s being coopted by AI companies, and respond appropriately. But just telling people to disaffiliate from banners that have been coopted doesn’t address the root problem of whatever allowed the banners to be coopted in the first place.
I’ve been trying to figure that out, and my best guess is “a bunch of people were using a kind of marginalist consequentialist reasoning that is very vulnerable to epistemic distortion, especially combined with the level of conflict aversion that’s standard in our community”, which then both led them to both do work that was bad (like WebGPT) and also to fail to disaffiliate from others doing work even they thought was bad (like scaling up LLMs).
So now when I try to evaluate whether someone’s effect on the world will be long-term good, the first-order question I ask is “do they avoid that problem?” Because if not then I should model their reasoning as being too easily corrupted by adversarial forces to be trusted.
I have a few objections:
That’s a nitpick, but Anthropic’s AAR strategy as-written was to build automated researchers which help with fairly simple alignment tasks, which I don’t believe to be equivalent to “iterative safe development of successive AI systems, building off of our experience with previous systems and the assistance of those systems”, which IMO is the bigger scalable oversight-like paradigm.
The situation with AI labs is far less similar to Yudkowsky’s Inadequate Equilibria than WWI. [1] IMO it is similar to WWII: any Strategy which ensures that a misaligned ASI isn’t created also means that xAI won’t create a misaligned ASI. Unless xAI has anything like plans to increase its alignment-related measures, the Strategy has xAI shut down, drained of compute like the covert project from AI-2040 or subject to thorough oversight.
How far does the “iterative safe development of successive AI systems” conditioned on mechinterp-based oversight even scale? Suppose, as an extreme case, that Agent-4 discovered a batch of mechinterp techniques usable for reconstructing all its values in Agent-5. Then how different is this task from reconstructing the batch of human values in Safer-4? Or does it mean that mechinterp is (yet?) too weak for meaningful oversight and that the Consortium from AI-2040 would have solved the problem? That mechinterp and Agent Foundations try to solve the same problem from different angles?[2]
I have yet to read anything about the lock-in of anti-peace mindsets in WWI, but I can imagine mechanisms like, e.g. two states both being in a crisis and failing to have the leaders come up with a plan which doesn’t involve taking over the opponent’s resources as well. In this case, the utility function of the two leaders is convex, but they can’t precommit to flipping a coin and having one surrender peacefully. Additionally, the Hawk-Dove strategy selects for a ratio of hawks who beat the doves while fighting against each other.
I made a similar claim in the collapsed section from my post.
A pair of arguments you’ve maybe heard before but want to make sure you’re tracking:
1) AI companies are naturally motivated to solve legible problems. They aren’t naturally motivated to solve illegible problems. (see Wei Dai’s Legible vs. Illegible AI Safety Problems)
2) It doesn’t actually help to make the legible progress sooner. You get the same wall-clock time to leverage the benefits of the legible progress.
i.e. sooner or later, companies run into situations like HuggingFace, and CEOs and politicians start to notice and researchers have concrete examples of misalignment to study. In the world where safety-conscious people didn’t help companies move faster, we still end up with comparable time afterwards to leverage the legibility.
The question is just “beforehand, did you you have more calendar years of people thinking through the problems that were harder-to-think-about, less commercially useful-to-solve? Or not?”
It seems fine to argue “it’s very hard to make progress on illegible problems, and most such work will be useless.” But I don’t think it’s usefulness is zero. At the very least, having mapped out more deadends is useful when you get to the HuggingFace point. And there’s at least a chance some of the research agendas would pan out.
(what counts as “illegible” varies a bit depending on person. Things are more legible if they are easier to understand. I can dig into examples if that feels helpful)
...
There are counterarguments that feel (potentially) compelling to me here, which involve overhang, or company culture, or what kind of governance situation we’re in, or some other kinds of path dependancy.
But, if we don’t have a specific, valid argument of that type, my baseline expectation is that commercially-useful safety work is probably worse than useless.
The strong version of this argument (‘same wall-clock time’) relies on equating:
“so legible that a company is motivated to solve it” (used in (1), and the references to ‘commercially useful safety work’, and the implication that Sam is working on this stuff), and...
“so legible that the world will successfully coordinate and pause until the problem is solved” (required for you to reliably get as much time as you need in step (2)).
But I think the bar for legibility is obviously much higher for the latter thing than the former.
I think there’s a weaker version of this argument like: “the more legible a problem is, the more of a control system there is for ensuring we eventually get an adequate amount of work on the problem (both via pausing and via increasing investment) and this means that you’ll have much less counterfactual impact on legible problems than you’d naively expect”. And conversely: “it’s really high-leveraged to improve the control system / make problems more legible / improve our ability to know how high risks are”. I basically do buy this argument. (And relatedly I do think that most technical Anthropic people would probably have a lot more impact if they quit Anthropic to work at a 3rd party evaluator or a govt evaluator.)
To add to what Raemon is saying about legible vs illegible problems:
This seems to be set in some insane counterfactual world where we don’t align the models but they still get developed and deployed at the same rate as in our one. Even with government responses to AI being as anaemic as they are, there is a limit to how destructive a produce you can ship without getting sued/arrested/regulated. Being able to align the current model is always one of the main bottlenecks for being able to deploy it and start developing the next one.
If you can’t align a model you can’t deploy it. If you can’t deploy your models that undermines profit/fundraising/RnD.
Thank you for posting this!
The section on DeepMind explains how capabilities & alignment came to be intertwined. Since GDM was also probably the organization with the strongest prestige dynamics & status hierarchies adversarial to safety, I’ll add a couple of subjective anecdotes on what it was like from the inside to hold the view that DeepMind’s core AGI roadmap was both (a) plausible on short timelines and (b) dangerous, such that alignment should be taken seriously.
I worked in the comms and policy org starting in 2018. All external comms were carefully monitored—not only for confidentiality reasons, but also to avoid reputational damage. If there was perceived risk of online controversy (social media, press articles), the team would either deny the request outright or take some actions (like comms training, editing the content) to reduce reputational risk. In my recollection, the operating principle was to earn either positive attention or no attention. Controversy would be met with a stern email or a meeting that appeared on your calendar, and future comms would be monitored more carefully.
If a researcher was asked about their views on existential risk from AI, they were guided to respond with a statement along the lines of: “It’s not useful to engage in that kind of alarmism, which originates in science fiction. Some people confuse AI with movies like Terminator—that’s simply not the reality. The AI we develop will be safe by design. After all, we have a team of expert researchers building it!” (--> Steer conversation towards beneficial applications in health, climate, etc).
It took months of intensive internal advocacy to change this protocol towards one where publicly acknowledging AI risk was less actively discouraged, and the Safety team could start publishing a higher volume of (positively valenced, comms-friendly) content on AI safety without being internally censured. To be clear, the early resistance to commenting on AI safety in particular did not come from the top of the organization, but rather from middle layers that were responsible for policies on external communications, but lacked sufficient technical or strategic context to understand AGI was plausible in the near term and not safe by default.
There was also a default of distrust with other labs, which—whether or not justified—was notable for the status dynamics that rewarded this. My comms training and instincts on risk-aversion run too deep even now for me to say much more on this, but one can imagine it being socially more favorable to echo a sense of vague distrust than attempt to drive forward policy or safety work that involved coordination with other labs.
These dynamics steadily improved as the Overton window shifted and it became a positive status signal to be supportive of safety work. People like Geoffrey, Rohin, and Allan joining were helpful for this, as was Jan, Vika, et al’s early work in this area.
>My comms training and instincts on risk-aversion run too deep even now for me to say much more on this
You don’t have to, but I would like it if you did, and I’m sure others would too.
To make Jan Kulveit’s point more directly, there are some pragmatic obstacles preventing us from developing the norms you want (which I agree would make us better off). One problem is that, in practice, the community values lab associations quite a lot. People with safety & governance titles are liked more than people with capabilities titles, but regardless, early lab people seem to garner much more respect than say, MIRI employees. It’s not like anyone is offering Scott Garrabrant podcast & talk opportunities, even though he is much more intelligent and lucid than the [ex-]lab employees I’ve had the chance to meet around rationalist venues. Part of this might be straightforwardly a money and power thing.
Another problem is that there are actually a bunch of people in the community who decided to do the thing you describe, and ex ante stay far from AGI development. But as with Scott Garrabrant, what happens is that they essentially just become ignored, even if their alignment research is really cool. Often they are not even given credit for their decision not to join a lab, because in most cases it’s impossible to tell the difference between the people that decided not to join and the people that just didn’t have the credentials to work there in the first place. So even if most people abstain, the only people who really get notoriety for it are the guys who got involved early and then had some Oppenheimer moment and pivoted into a safety career.
I’m definitely down with just asking individuals to personally resist these dynamics. Not inviting former labcos to talks would work better than nothing! But I think we’d also have a higher chance of success if the strategy for fixing peoples’ incentives was more systematic.
Huh, I sure am in a bubble, but at least immediately around me being a lab-employee or ex-lab employee, especially someone who worked in capabilities, is a pretty big hit to your reputation, and Scott Garrabrant is very well-respected. When I organize private retreats or events, I am much more likely to invite Scott Garrabrant than ex-lab employees.
Datapoint from the same bubble as Habryka: I’ve made a concerted effort not to be close with anyone who works at a lab, or has worked capabilities at a lab, and I make sure to put substantial social distance between myself and anyone who has joined a lab. But I have a lot of love for people like Scott Garrabrant, and other rationalists and agent foundations people, even if they’ve gone quiet in recent years.
You guys are both much more centrally in the shit than me, so maybe my impression is just incorrect. One experiment I ran before posting was that I searched “Richard Ngo” on YouTube, saw like 10,000 pictures of his face doing talks/podcasts/whatever about miscellaneous politics, and then searched “Scott Garrabrant” and got a couple presentations by him but also some guy named Daniel Scott Garrabrant playing the guitar. Although to be fair I don’t see anything from Leo Gao either.
I mean I think there is “a Lightcone/MIRI-ish cluster, where being a lab employee is obviously really bad by default and you need some really good points to make up for it”, and then, idk, the broad professionalized EAcosystem where lab employees seem to be the experts and have lots of money, why wouldn’t they be high status?
my guess is this is largely personal inclination. i don’t like giving talks or podcasts or whatever. i’ve turned down a lot of opportunities, and when i do speak at an event, i usually ask for it not to be recorded.
Why I don’t totally disagree (I doubt Situational Awareness would have garnered so much attention if Leopold had just been a random guy), my understanding of Garrabrant’s work is that it’s mostly very technical, which imo is the reason why he’s not more prominent?
Just off the top of my head, Yudkowsky, Gwern, Shulman, Greenblatt, Soares have all never worked for labs (and several have worked for MIRI) and all get a decent amount of attention (appearances on stuff like Dwarkesh’s podcast, for example, and far more than that in the case of Yud). I suspect the difference is that any random person can read ‘Why Tool AIs want to be Agent AIs’ or ‘Current AIs seem misaligned to me’ and understand it as well as understand why it’s immediately relevant to AI in the short to medium-term.
To influence the public opinion on AI safety, one has to be considered an expert on ML/AI not just “get attention”.
Shulman is apparently unknown outside of the AI safety community, in fact I had to google him to find out who was that (is it Carl Shulman, or did you mean John Schulman?). Gwern and Greenblatt are only moderately known (I tested on some ML practitioner friends), each frontier lab has over a dozen of people known better than them, and there are many dozens of people in academia and startups who are much more likely to be considered AI experts. Yudkowsky is known but not really taken seriously by the public, his expertise is not well-respected. That only leaves Soares, who in fact has a Wikipedia page and satisfies all the criteria.
So I think your list rather self-defeats your argument
My comment was meant to describe respect within the rationalist community. Obviously Eliezer Yudkowsky and Gwern are highly respected among rationalists. So they do serve as important counterexamples to my comment.
Yeah, it’s Carl (I assume). A bit hard to evaluate because he’s nearly unknown to the broader world but influential within parts of the x-risk / old-school-rationalist cluster. I think he could have a much higher profile if he wanted one.
I think I’m in both these groups :-)
Maybe part of the answer is giving people an alternative tech tree to work on that doesn’t involve summoning demons, and still provides fun and bragging rights.
On the other hand, you definitely want to incentivize them to leave. Maybe an alternative is they have to first do a round of talks on “what specifically led me to mess up by thinking it was a good idea to work there in the first place or not leave earlier”. Or something.
Yes agreed but it’s so far from where this community is and IMO impossible at this point. EA/LW has no inclusion/exclusion criteria. There is no actual community. It’s a messy network of ideas and relationships and money. Even if 80% of us bought into some constraint, if anthropic decides to do 5B of philanthropy next year who do you think is getting the money? The people who wine and dine them or those with already status in some way (which may or may not be correlated with things we care about).
Then those people will have jobs to give out, or a high position in government/business, and the new grads will drift towards those people, because they need jobs and want status, and in a few years, we will be the greasy guys with no job attending rationalist meetings in a tier 2 city. If you complain at that point you will look like a loser who is just salty you couldn’t get the jobs.
Like I really really care about AI safety, and I still needed lighthaven/PDKU to give me a grant and an official title to really do anything significant. The pull towards “having a title for the thing I’m doing that I can explain to new people I meet and my mom and I make enough money to survive or relative to expectations” is just such a strong attractor state.
The world is unfair and it is very very hard to create a social movement that doesn’t get eaten from the inside from these dynamics (don’t capture what you can’t defend?) and if you really want to trace the blame all the way back, I blame the founders of EA and Lesswrong for thinking you could make a boundaryless structureless social movement that could defend it’s original values. This is an incredibly hard thing to do, and they weren’t particularly close to hitting the mark. I think my issue with OP is that while I generally agree with a lot of the micro assertions he is making, the incentives of these communities always pointed in these directions. I think if MacAskill/Yud/ other early people in the movement had more direct experience in community organizing and politics, and less as philosopher kings, then many things about EA/LW would be very different, and we wouldn’t have these issues, although we probably would have missed out on many of the good things too.
Seems kind of defeatist and overwrought
All of this psychoanalysis misses the mark with me because it seems like the people advancing capabilities have been straightforwardly beneficial to X-risk, which therefore makes their stated motivation (helping with X-risk) seem very plausible.
There is no getting around the fact that AI will be built by people who build AI. Therefore, the people who build AI should care about X-risk. You seem to identify the messiness of the capability/alignment divide and want to abandon alignment in order to get away from capabilities; success depends upon the people who embrace capabilities in order to embrace alignment.
People could falsely seize on this line of reasoning to justify empowering themselves, but as someone who is not currently seizing on this line of reasoning to empower himself, it also seems plainly correct to me. Sometimes the best thing to do is also dunkable.
I don’t think this is particularly unusual in history either. It seems to me like the people who write history are often power-seeking and the success of our species depends upon these people also doing the right thing; even if these power-seeking people are in some ways annoying or unpleasant or straight-up bad people.
If you want a pause, yet again, the power-seeking AI risk people have done you a tremendous favour by being vocal about X-risk. Is there anything they can’t do, any way they haven’t helped?
Speaking of psychoanalysis, I would be tempted to do it to understand why people are so obsessed with getting more time, even at the expense of being able to use that time effectively; because there doesn’t seem to be any rational underpinning to it. Say that you got what you wanted 15 years ago, and everyone concerned about X-risk abandoned ship on anything that might advance timelines. There’s still Google Brain and Facebook AI Research and NVIDIA and many other labs working in AI; there’s still tons of academic researchers scattered around universities happily tinkering. Progress would still happen, slower and in a less focused way, but it would happen. What would we be doing with the extra time we bought? When things get critical, there would be no helpful quotes from lab CEOs floating around about X-risk; you wouldn’t be able to dunk on the CEOs (“If it’s so dangerous, why are you building it, huh???” 1B likes, best tweet of all time); when you went to senators to convince them about AI risk, you look like a total outsider. If you make any headway the labs actively campaign against you. And what alignment research is supposed to be happening with this extra time? Everyone is scared of doing anything too impactful! You just have to hope that AI alignment turns out to be the rare kind of problem you can one-shot theoretically without touching the details of how the AI is actually thinking or learning, since dwelling too much on those details is capability work.
And say you get your theoretical one shot, all thanks to the X years we bought. Then what? Who cares? Why should labs listen to you? Will you even trust them enough to tell them?
As an aside I found this part quite moving. And as someone who’s found the compute overhang argument extremely convincing in the past, I should acknowledge that I think recent events have been reasonable evidence against it; it does seem like if the LLM craze and then Agent-LLM craze had been successfully delayed 4 years, things wouldn’t move that much faster due to the extra 4 years of Moore’s Law.
As I mention in the post, the object-level benefits/harms of more time are of secondary importance in my mind. My main question is: how do we figure out which people “care about x-risk” in a way that reliably shapes their decisions? Because historically the community has taken people saying the words “I care about x-risk” as far too much evidence that this consideration actually steers their decisions.
Re the rest of your comment: it feels hard for me to pick out the underlying crux underneath all the considerations you raise. I do think there’s something that feels kinda demoralized about it—in particular, the sentence “you wouldn’t be able to dunk on the CEOs” conveys a sense that people who care about this stuff are by default gonna be throwing pebbles from the outside at the real decision-makers.
But part of what this sequence is trying to convey is that thinking clearly is so extraordinarily powerful that the people who are able to do enough of it to get on board with alignment are able to produce orders of magnitude more capabilities progress than equally-smart “mainstream” capabilities researchers even when that wasn’t their goal.
So I am generally extremely optimistic about this community’s ability to exert influence if it stops punching itself in the face.
(P.S. I appreciated your aside in your other comment, and it helped me orient to your main comment in a less defensive way.)
Well, if you want to argue that the current paradigm has led to a discrediting amount of harm, and should therefore be thrown out, then it seems like “how large were those harm?” should be an important input into your argument?
I can see you skipping over this step if you were like: “X person believes that timelines speed-ups are extremely bad, so if I can just argue that X person caused timelines speed-ups, then I’ve established that they’ve done badly by their own lights”.
But the people you’re critiquing in this essay are more like “timelines speed-ups are bad all-else-equal but can be (and are often) outweighed by other positive effects from your actions”. So in order to establish that this reasoning went wrong by its own standards, it seems like you do need to grapple with how bad the acceleration effects were compared with the upsides.
But if ppl went in explicitly saying “we’ll differentially advance alignment” but then they mostly accelerate capabilities and start saying “we’ll get xrisk ppl in power and that will outweigh the harms”… then I think it’s fair to say they get a hit to their credibility and trustworthiness
Like, one accusation Richard could make is “they pessimised their goals overall”, where I think it’s v unclear for the reasons in the top level comment
But another accusation is “they pessimised their aim of differentially advancing alignment”, which I think is pretty plausible. If all AI safety ppl had refused to do capabilities, I expect alignment would be better at today’s level of capabilities. You’d have had most anthropic ppl doing alignment rather than 90% of them doing capabilities—could have been loads of Redwoods and METRs and biggest safety teams at other AI companies etc
Like, I think the bet that starting Anthropic was good is a bet on their future good actions outweighing the way it’s worsened alignment (at each capability level), which is a v different to what they had in mind initially I would guess
Agree with Tom. To elaborate a bit more: the reason this post focused so much on the idea that the alignment/capabilities distinction has lost its meaning is because that distinction was implicitly or explicitly grounding a lot of other arguments about impact.
I’ll discuss this more in the next post (partly because the comments on this one have been helpful in clarifying some of the concepts I’ve missed so far). But as one brief intuition: the key thing I’m worried about is circular justification loops, where ambiguities in what we mean by “alignment research” allow people to justify all sorts of stuff that never cashes out in the real thing we want. As some toy examples:
Or
I’m not saying that any individual person has or would endorse these circular arguments, but I am saying that the community as a whole has ended up being unable to distinguish them from non-circular arguments.
How might you attempt to prevent this? For example, you could divide work into type 1: directly aimed at solving the hard parts of the alignment problem; type 2: aims to produce more type 1 research; type 3: aims to produce more type 2 work; and so on.
This is obviously a kinda blunt way to classify things, though it’s better than anything we actually had. A very stylized retelling of my post is that:
MIRI and Paul-before-OpenAI (e.g. iterated amplification theory) were doing type 1 research.
Paul at OpenAI claimed to be doing type 2 research, and also debunked a bunch of MIRI’s worldview, and so people started interpreting Paul’s type 2 research as a central example of alignment research.
Paul’s arguments for why doing his type 2 research was a good idea were pretty dodgy but he didn’t write them up at the time so it was hard to publicly critique them.
Dario was claiming to be doing type 3 work (by scaling up the models, which would help type 2 work like Paul’s), and nobody publicly called out how dodgy his arguments were, so people started interpreting Dario as a central example of a person who cares about alignment, which then helped him build Anthropic.
And so now we have this whole edifice of people doing stuff that’s “good for alignment” via these very tenuous chains which they have a lot of reason to do motivated cognition about. I think Paul’s retrospective on why RLHF was a good idea was very far from his usual epistemic standards—for example, the original version literally didn’t even mention ChatGPT! And so, while I know much less than Paul about the details of the situation, I can’t trust him to actually figure out if his reasoning was wrong, and I need to model him as basically just defending his past self.
Also, unfortunately, almost everyone in the field is too conflict-averse to characterize others as doing motivated reasoning, and so we’re stuck in the middle of these insane epistemic distortions without any way to talk about them. So, again, my main contention is not “the current paradigm has led to a discrediting amount of harm” but more like “the current paradigm has led to a discrediting amount of rationalizations for doing things which drove an enormous proportion of total capabilities work, from people who are demonstrating nowhere near the level of integrity necessary for us to trust that they’re reasoning clearly about the counterfactuals, and (relatedly) seem totally incapable of calling out even the people most obviously doing deceptive or motivated reasoning under the banner of alignment”.
Btw one piece of history that you might find interesting (and that complicates the stylized “Paul debunked MIRI and then did a different type of research based on shoddy arguments”) is that Eliezer very much liked the RLHF paper.
Further context/corrections:
The paper Eliezer is referring to here is not the RLHF paper; he is referring to the follow-up summarization paper. This is the original paper, and this is the follow-up paper Eliezer is referring to.
It seems the key point here is that he is particularly happy about the mild optimization approach to avoid Goodharting/over-optimization, not RLHF specifically (though he could still have been happy about RLHF as a whole; that detail seems unspecified). Edit: more precisely, Eliezer responded and said: “I was happy that they had the realization “this breaks down if you load it too much” and then measured how much. Unprecedented back then. Still not very precedented now, to imagine what might break in advance of seeing it actually break, ala HF iykwim.”
Thanks for providing the extra context!
When posting my above comment, I thought that this approving quote-tweet further down in the thread might also indicate some broader appreciation for the research.
Huh, yeah, that’s a surprising and useful datapoint which it feels like should noticeably update my model. Maybe it swings me a bit away from “this went badly in a predictable way” to “it was a reasonable bet at the time, and just looks bad in hindsight”.
FWIW I do think my understanding of “good science” differs significantly from Eliezer’s, in a way that’s related to him being much more Bayesian than me. The person whose research taste I trust most re agent foundations is Scott Garrabrant, who has pretty different intuitions from Eliezer on a bunch of things. (For empirical research, I’m less confident—maybe Andrew Critch, or Owain Evans.)
Ok so in order to do that, you want to argue that there’s a ton of rationalization.
One strategy for arguing that is to go through the arguments and show that any clear-eyed accounting of the costs and benefits would show that the costs outweigh the benefits. In which case you’d need to reason about the size of the costs.
But yeah, I do agree that there are other ways of doing it. E.g. if you can show that some arguments are so bad or so inconsistent over time that they’re much better explained by rationalization than by a genuine attempt to reason through the positives. (I’m not currently sold on this for the RLHF post.)
Yepp, that’s right. And I don’t feel very capable of showing that any clear-eyed accounting of the costs and benefits would show that the costs outweigh the benefits, because the main way that people are rationalizing involves appealing to a bunch of really-hard-to-evaluate counterfactuals.
So instead I’m pointing to how these arguments are quite skewed (e.g. they often involve giving the vibe that they’re making that appeal, without even specifying the counterfactuals) and don’t seem to involve the kinds of updating that should be happening (e.g. not updating specifically on ChatGPT).
And going forward I want to set a community norm that we interpret appealing to really-hard-to-evaluate counterfactuals as rationalization by default, unless the person involved has demonstrated a lot of integrity (e.g. enough that we can be confident they’d whistleblow on their employer behaving badly, or call out their boss’ research strategy if they think it’s harmful, rather than finding a clever rationalization for not doing so).
FWIW my reaction to that community norm is something like: yeah reasoning about counterfactuals is hard, and rationalization is probably not rare, but I don’t see a great alternative. And I feel like I prefer the world where a bunch of people earnestly do their best to try to reason about different things that might happen depending on whether they take various actions (with varying degrees of rationalization sadly & inevitably involved), than a world where the community enforces norms that are so strict that they systematically prevent people from doing rationalized work while staying in the community. (Seems like that’d require way too much conformism and rule out too many plausibly good impact strategies to be worthwhile.)
My current model of the alignment community is that it has ossified in a similar way as the ML establishment in the 2010s (or the AGI companies in the 2020s).
In general when you talk to someone in those ossified communities about AGI risk, even if you manage to get them to concede on each individual point, the fallback response you’ll get is something like “yeah AGI could be dangerous, and we don’t seem to be on track to solve alignment, but I don’t see what I can do to help much”, and then they go back to working on whatever they were previously doing.
In this case, I am talking to alignment community about how the strategy it has been using has driven a big chunk of capabilities progress, while producing little meaningful alignment progress. And so my response is similar to what I say to capabilities researchers: it is your job to figure out an alternative, or else to go sit on a beach somewhere, rather than continuing to contribute to unhealthy ecosystems that can’t reliably steer towards good things rather than bad things.
I just think there’s a bunch of valuable work being done by the alignment community and that the positives outweigh the negatives by a good margin, so I feel like I’m in a pretty disanalogous position to that. (That’s IMO compatible with “reasoning about counterfactuals is hard, and rationalization is probably not rare”.)
Like I’m still interested in proposals for how people could do way better. But if the proposal is “let’s all go sit on a beach somewhere” or “lets heavily restrict what category of arguments we can seriously consider, in order to fight rationalization” then that seems net-negative to me.
what’s your estimate for the likelihood that such an alternative exists?
Definitional point: I think the natural usage of “pessimised X” is whether they made it worse compared to if they hadn’t done anything at all. What you’re doing here is comparing a world where AI safety people are only somewhat focused on X (where they spend some resources on differentially advancing alignment and some resources on empowering people who they think will make better decisions in the future) with a world where they went all-in on just X (only doing safety work). Yeah of course the latter will look better according to X, but that doesn’t mean that the group as whole pessimised X. Compare an accusation that EA has pessimised animal welfare because animal welfare would be going much better if all the EAs had refused to work on longtermism. (Because then a bunch of them would have gone to work on animal welfare instead.)
I think it’s reasonable to accuse people working on pre-training at anthropic of pessimising “differentially advance alignment”. (But presumably this downside is obvious to everyone involved, and if a critique of them is going to add anything, it should compare the benefits they claim outweighs that downside.)
You could accuse the AI safety community as a whole of pessimising “differentially advance alignment” if you think that removing them all from the world would have made alignment better at the current level of capability. But my sense is that the AI safety compute has succeeded in differentially advancing alignment by this metric.
This one is more interesting! I think it’s plausible that the founders of Anthropic in particular (i) would have claimed that they were overall differentially advancing alignment, and that (ii) they’ve differentially advanced capabilities so far (where in particular: if them founding Anthropic ended up redirecting ppl from safety work to capabilities work, that seems like a consequence of their actions that’s fair to count as a harm to safety work).
I think Richard is advancing a much broader claim than that, though. About the safety community as a whole.
(Even though I’d personally be more interested in the more targeted Anthropic critique since I’m more uncertain what to think there and more likely to end up convinced.)
I think you could broaden this critique beyond the founding of Anthropic?
Could apply it to the founding of oai, Dario’s scaling work at OAI, Paul’s work on rlhf, and some of Geoffrey’s stuff at gdm. In each case, they advance capabilities more than alignment and nudge the safety community away from the harder + more central problems. (Tbc, idk if I agree with this argument)
I agree it’s implausible that the AI safety community as a whole has made alignment worse (at each capability level), compared to if they didn’t exist.
But they have plausibly done this relative to nearby and obvious strategies they could have pursued (like not doing capabilities work, not joining agi labs, prioritizing scalable work).
Your reply is: “but that’s bc they’ve chosen to empower the good guys as well.”. But I think their focus has shifted in that direction over time, which is an amber flag.
The analogy to EAs pessimizing factory farming is interesting. But “we’re pivoting to empowering the good guys, aka ourselves” feels more sus than “we’re pivoting from factory farming to bio”.
I do think it matters what ppl said historically about their strategies here. If the AIS community said from the beginning they would put a lot of effort into empowering the good guys, that’s v diff to if they said they’d just focus on differentially improving alignment.
I think “avoid rationalization” pushes toward “use predictable/deterministic processes” which pushes away from FDT (because FDT is not well-defined) and to a lesser degree EDT (because evidential conditionals are hard to evaluate in a modular way), therefore relatively towards CDT and marginalist thinking. And especially the kind of thinking that is formalized in microecon / macroecon, since that is a standard framework, so it’s easier for people to check each others’ work.
Also “avoid rationalization” pushes against virtue ethics (it’s not very standard or math-y, people disagree about virtues and weighting, etc), unclear on consequentialism (standard and math-y, but hard to evaluate), positive on legalistic deontology, perhaps positive on specifically Kantian deontology (unclear comparison with consequentialism; part of this depends on, empirically, how well Kant scholars can agree on how to interpret Kant’s deontology).
Also, more STEMlord-ism (trust systematic thought, logic, and empirical validation), anti woo, negative on many parts of the humanities, but positive on relatively empirical humanities such as history.
I’m thinking historically about how Calvinists were worried about “total depravity” and correspondingly adopted systematic, literalistic Bible interpretation. For present context, standard econ, rational choice theory, etc are more relevant than the Bible.
To the degree I worry about rationalization, I correspondingly prefer more adoption of standard systems / frameworks, especially for high-stakes / important decisions. It seems your approach is pretty different. It’s possible at least one of us is making a sign error.
That’s fine and reasonable. I just think it’s confusing to bring in “pessimising” if that’s the argument.
Yeah. I don’t really know the ground truth here. Famously MIRI talked a lot about pivotal acts (=empowering the good guys enormously) and I would’ve speculated that ant founders had “we’re ethical and you should empower us” as a non-negligent part of their pitch (eg “it’ll be important to have AI safety people ‘in the room’ and being taken seriously in the future”). But then also that ant probably has gone significantly more in that direction over time, and more like “it’ll be great to have AI safety people win the race” rather than just being in the room.
I agree there’s a danger of people mouthing the words “I care about x-risk” and not really caring, having some psychological disconnect where they agree with it in abstract but in practice see AI as a cool toy they can work on, and this is very dangerous if the people building AI end up in this category.
So in that sense, as someone whose main contribution to the discourse is to occasionally cheerlead for people to work at frontier labs and not worry about advancing timelines, I worry I could end up contributing, in a tiny way, to a fatal lack of care.
But my main frustration with the anti-capabilities ideology, that I interpreted you as belonging to, is that they seem to want to take this “I should be careful, if we screw this up it’s over” instinct, this bit of fear in people working at the frontier, and weaponize it to say “you’re terrible, quit immediately, get out of there, if you don’t you’re morally compromised” which is a really bad idea. We need the people building AI to feel the fear; if everyone who feels the fear quits, then...
When I imagine people following the advice to quit, I imagine basically what you said, people on the sidelines throwing pebbles; not because you have to be at a frontier lab to be relevant, but because pretty much everything beneficial that you can do will put you up on stage (i.e. contribute to capabilities) and get some pebbles thrown your way.
I’d say I’m demotivated on stopping AI, and optimistic on aligning it. When I engage with someone on the other side of this their theory of change often seems to be AI never getting built, which seems completely implausible to me, and also seems to neglect how much the AI pause movement owes to the frontier labs being vocal about x-risk.
If I had to guess at the cruxes which drive my disagreement with the anti-capabilities camp they’d be:
1) Alignment success is quite sensitive to how much care is taken by the people who build it.
2) Stopping AI from being built is very unlikely to succeed.
I’m not sure how much they apply to you—I’d guess your view would be something like agreeing with 1 and 2 as stated, but believing that the care taken needs to be far-sighted and free from the various incentives which pull people away from real alignment research, which means people need to halt and catch fire and leave the source of the distracting incentives before we can actually think with care; only after we’ve done that can we think about building it with care.
Incidentally, it’s interesting that this comment uses the heuristic of “assume that you are choosing on behalf of a significantly wider range of people than just yourself” to argue in the opposite direction of where Richard thought that principle would go. (I.e.: It’s painting a picture of the world where no one who altruistic and concerned about alignment tries to do harm-mitigating or power-accumulating strategies, and saying that picture doesn’t look great.)
I agree that “assume that you are choosing on behalf of a significantly wider range of people than just yourself” is a good heuristic for decision-making, but I think it’s unlikely to be the crux between Richard and the people he disagree with.
The leaders of the labs want to survive just like everyone else does. If someone were to figure out and announce an actually satisfactory alignment plan, those leaders would probably embrace it, maybe not at first, but after the number of people publicly promoting the plan had grown large, because at some level many of them probably strongly suspect that the plan they have now is unsatisfactory.
They persist with their current plan because the pursuit of power is very motivating and compelling to them. Give them a way to hold on to their power while decreasing extinction risk, and there is a good chance they would take it. Of course, that only helps us if someone comes up with a satisfactory alignment plan, which is a huge if.
Of course not everyone at the labs is motivated by power or personal prestige. Another big motivator is an ideological commitment to technological progress and the bright science-fiction future. But maybe the basic analysis above partly applies to people with that motivation, too.
i think this blog post correctly captures many dynamics that happened over the past years. i’m glad it exists and i hope people update. i also don’t feel like it describes my reasoning in particular.
fwiw, i remember arguing at the time with all the people involved about how “if models can look up information online, they’ll be more honest” was obviously not a good theory of impact. although i was on the RL team, i did not work on webgpt (or chatgpt) at all, and focused on studying goodharting instead; i don’t really endorse the theory of change of this work helping with alignment anymore, but at least i still feel confident it did not contribute to capabilities.
i also remember arguing that this was obviously a terrible idea. when it later came to light that openai was investing heavily in speeding up chips, i was saying “i told you so”
i don’t really think of situational awareness as “ai safety people”.
i wasn’t around for this, but i still hear this take sometimes and it seems obviously completely wrong.
fwiw, i was also one of the main people who pushed very heavily for eleuther not to publicize CoT widely. it’s unclear to me that this was counterfactually impactful in retrospect, and i’m not sure i’d do it again. as a result, i don’t think this is great evidence for your position.
re: making models better at conceptual research so they can help with alignment more, this seems like obviously a very tenuous story of impact and i wish people who are working on this rn stopped working on it. i am happy to argue with anyone who disagrees.
overall, when i read this, i don’t feel like i am getting evidence that my own reasoning is flawed in this specific way, and my main update is i was right about all this. (of course, my reasoning is likely flawed in other ways.) to the extent other people thought this way, i hope they reflect on this, i guess.
Carl Shulman is its Research Director (EDIT: and co-portfolio manager), and used to be a MIRI employee (2010-2013).
prompt: carl shulman’s publications on AI safety
Carl Shulman’s AI-safety work is concentrated in early foundational writing on superintelligent-agent alignment, instrumental convergence, intelligence-explosion dynamics, and governance rather than contemporary empirical alignment research. His publication record also includes adjacent work on digital minds, forecasting, and long-run governance.[semanticscholar][alignmentforum]
Core AI-safety publications
Publication
Year
Coauthors
Main contribution
Machine Ethics and Superintelligence
2009
Henrik Jonsson, Nick Tarleton
Argues that ordinary machine-ethics approaches make assumptions—incremental deployment, human-comparable capabilities, human institutional embedding—that may fail for agents at or beyond human level. It calls for advance analysis of the harder alignment problem. [intelligence]
Arms Control and Intelligence Explosions
2009
Stuart Armstrong
An early AI-governance paper: analyzes incentives for unsafe competitive development, winner-take-all dynamics, and the possibility that advanced AI capabilities could also make verification and international agreements more feasible. [intelligence]
Which Consequentialism? Machine Ethics and Moral Divergence
2009
Henrik Jonsson, Nick Tarleton
Discusses divergence among consequentialist objectives and its implications for constructing moral machines—relevant to the problem of specifying AI goals. It appears in his scholarly bibliography, though it is less directly focused on catastrophic-risk mechanisms than the three items above. [semanticscholar]
Omohundro’s “Basic AI Drives” and Catastrophic Risks
2010
—
Critically examines the argument that a broad class of goal-directed systems could develop instrumental tendencies such as self-preservation and resource acquisition, and relates those tendencies to catastrophic AI-risk scenarios. [intelligence]
Implications of a Software-Limited Singularity
2010/2013 version
Anders Sandberg
Explores implications of a rapid AI transition constrained primarily by software development rather than hardware or physical replication. It is important background for takeoff-speed and strategic-risk debates. [semanticscholar]
How Hard Is Artificial Intelligence? Evolutionary Arguments and Selection Effects
2012
Nick Bostrom
Addresses evolutionary and selection-effect arguments about the difficulty of developing AI, bearing on forecasts and the plausibility/timing of transformative AI. [semanticscholar]
Racing to the Precipice: A Model of Artificial Intelligence Development
2016
Stuart Armstrong, Nick Bostrom
Models competitive AI development in which actors trade safety for speed, formalizing “race to the bottom” concerns that remain central to AI governance. [dl.acm][semanticscholar]
Especially worth reading
If your interest is specifically AI x-risk and alignment, I’d prioritize:
“Omohundro’s ‘Basic AI Drives’ and Catastrophic Risks” — best for instrumental convergence and the case that misaligned optimization can be dangerous even without malicious intent.[intelligence]
“Machine Ethics and Superintelligence” — best for the conceptual distinction between conventional AI ethics and alignment for highly capable autonomous agents.[intelligence]
“Racing to the Precipice” — best for strategic dynamics: why safety measures may be underprovided under competition.[dl.acm]
“Arms Control and Intelligence Explosions” — best for governance implications and the argument that AI may create both unusual destabilization and unusual verification/enforcement opportunities.[intelligence]
Related, but broader than safety
Two later works with Nick Bostrom—Propositions Concerning Digital Minds and Society and Sharing the World with Digital Minds—are principally about moral status, rights, and social organization involving digital minds. They matter for post-AGI governance and AI welfare, but are not technical alignment papers.[80000hours][experiencemachines.substack]
For a more complete bibliography that includes papers, posts, and conversations, Shulman’s LessWrong publication timeline says formal work can be isolated by filtering its Format field to “Paper.”[lesswrong]
I’d be keen to hear your thoughts here.
Do you think it’d be a mistake to make models better at philosophy?
depends on details but probably?
[The below is verbose, sorry—I feel like I’m struggling to understand what just happened. I may have a blindspot around thinking about what mentality would even lead to “let’s make capability X because that helps with alignment somehow”. I can kinda scan through the logic step by step, e.g. “we need somewhat capable systems in order to study alignment”, but something about it doesn’t make sense. Like, aren’t we worried about AGI? Why would I make the thing I’m worried about? I’ll leave my faffing about here, because maybe it’s related to communal blindspots, though maybe it’s just me being thick at the moment.]
I’m probably being dumb, but I think I don’t understand this paragraph, so maybe my confusion will provoke clarification. Or maybe I disagree or at least am not convinced. I don’t understand how [the core difficulty about evaluating strategies based on differential impact] is accounting for other people making decisions that are entangled with yours. From talking with Fable, it sounds like you’re saying something like, “A bunch of differential impact justifications assume a fixed background of AI progress against which to differentially accelerate some things; but actually, the background isn’t fixed because people like you / people using justifications like yours form some crucial element of the background.”. Is that close to the mark?
I totally buy that “AI safety / alignment” stuff accelerated some key bottlenecks, e.g. serious scaling, and that that presumably pushed timelines forward. But it doesn’t seem like that depends on the thing about entangled decisions? I’m maybe missing something really simple and obvious, like one sentence. …. Ok Fable is saying that your point is that the justifications that were used invoked the premise that someone else will do it anyway, but the someone else is other “AI safety / alignment” people. Is this right?
Now I think that you’re correct that this reasoning is severely flawed in the way you say. But I want to say that this is a subtle kind of flaw, which is less important than another bigger kind of flaw. The bigger flaw, in a word, is that …. uh, it’s bad to make the dangerous thing? See next bullet point:
From talking with Rafe, I have a different and rather more simplistic / blunt analogy: you should think of differential acceleration as being like throwing stuff into a bonfire and hoping that will decrease its long-term growth. It’s conceivable, because for example you could throw in a rock or an ice cube, or you could clear out some nearby brush that could have caused a spread, or people will be impressed with how you made the fire bigger and trust you enough to turn their backs on you while you covertly stamp out the fire, or something. But on priors, if you just throw stuff into a fire, that’s going to make there be more fire. In the analogy, the fire is the ecosystem of AI research (researchers, companies, investors, products, etc.), and catching on fire is that ecosystem absorbing [whatever research you did in the name of differential acceleration] into the progress engine.
Or to invert Bojack Horseman, when you put your glasses back on, all the differential acceleration just looks like acceleration.
Anyway, overall I feel pretty curious about what just happened, and don’t really know how to understand it. My old decrepit frame is that some people feel a very intense pull to work on “the Thing that’s going on”, and considerations that go against that get sidelined. But overall I’m confused.
I appreciate this comment, thank you.
This is a very important point. I’ve focused on the case study of alignment in this post, but IMO related kinds of diseased thinking are also responsible for a bunch of political dysfunction, and general civilizational decline. And so understanding that phenomenon seems key to navigating the world right now. (This is a bunch of what Vassar et al. are talking about.)
Yes, Fable is right on both counts. The best single example is how deeply involved AI safety people were in all three of the ChatGPT-like systems developed in 2022.
I think the key question is what counts as the bonfire. Like, Ted Kaczynski thinks of the bonfire as western industrial civilization; Democrats think of the bonfire as Silicon Valley more generally; the Pause/Stop AI people think of the bonfire as everyone who’s on good terms with the AGI companies (or promoting the idea that alignment is plausibly solvable); mainstream AI safety people think of the bonfire as anyone who’s working on capabilities; Anthropic safety people think of the bonfire as anyone who’s working on capabilities not at Anthropic; and so on.
And so in order to reason clearly about not throwing things on bonfires, I think you need to do some kind of reasoning about which coalitions you’re entangled with, to which extent. Note that my concept of entanglement is much broader than just logically correlated decisions; it includes setting norms, policing coalition boundaries, and so on.
Most people just don’t have the doom-by-default outlook that’s prevalent in the LW/MIRI sphere. Sure, they are worried about AGI, but they also want cancer cured, poverty eradicated,
catgirl volcano lairs, etc. Building AGI seems to be the only way we’re getting that stuff any time soon, and in general one of the few remaining avenues for techno-optimism amid the bleak degrowth-and-despair cultural landscape, in the face of which accepting some substantial-but-not-overwhelming level of risk is a no-brainer.The part that doesn’t go through is when the people are people discussed the OP, who do thing there’s some huge (>10%) chance of extinction from AGI.
Even they likely think that the counterfactual impact of any (ex ante reasonable) decision in their power could contribute only a minor fraction of that number. Of course, a case could be made that this is pernicious motivated reasoning, I’m only saying that it’s consistent with my model of human thinking and behavior.
Great post again, thank you so much for describing the history in detail.
On the criticism part I agree with you 100%: most of the work at AI labs, including alignment work, has been a bad thing for years. (I’ve been trying to beat that drum on LW for years, too.) But on the constructive part I have some disagreement, and an alternative vision.
You ask: “If not alignment research, then what?” I think a better question would be: “If not AI, then what?” From 10000 feet, a lot of AI’s harm is due to the fact that AI is economically a substitute for humans. If we could shift to technologies that are economically a complement to humans instead—which means basically transhumanist technologies, like thought interfaces or pharmaceuticals or gene therapy—that would give a better path out of the whole crisis, keeping the future human.
The model to imitate here is how the world was steered away from nuclear power and toward renewables. When the anti-nuclear movement started out, renewables were almost as much a joke as transhumanist technologies are today. But due to the “full court press” of the anti-nuclear movement on laws, academia, industry and public opinion, enough researchers and investors shifted to renewables and now it’s quite competitive. So the vision is having a similar public movement against AI tech in favor of human-complementing tech—a stick and a carrot, so to speak. The desired state would be that AI becomes a political dead weight, and the people looking for money or prestige or scientific curiosity flock to human-complementing tech instead. The example shows it can be done and gives an idea of what tactics would be needed, how a large a movement, and how much time.
Without agreeing that we should abolish the AI industry altogether, I like the idea of giving smart ambitious people something alternative to do.
People who are good at making AI fundamentally like working. They like being good at their jobs! They like money and recognition but they also just literally want to solve problems and want to prove themselves competent. And they actually like beep-booping on the computer or writing equations on whiteboards and such. They’re not going to want to just go home and garden or something. And it would kind of be a waste of talent if they did.
Thank you for writing this!
I do think this misses one of the biggest psychological factors involved, which is that many/most people working on AI safety are, at some fundamental/visceral level, more excited by AI than repulsed by it. From what I understand, LessWrong grew out of the transhumanist community, and was founded by AI enthusiasts who hoped that a superintelligence would cure death and usher in a “glorious transhumanist future”. Even after the danger was recognized, it seems like many AI safety people continued to hope for that future. I think that, if you have a fundamental psychological aversion to AI, then upon hearing about the possible dangers of superintelligence, you are likely to lean towards a strategy of “ok, well let’s avoid building this, and also try to prevent anyone else from building this”. Instead, it seems like, even after realizing the danger, MIRI still held onto the goal of building a superintelligence for a long time; they just recognized that they needed to solve the alignment problem first.
Basically, it seems like most people who care about AI safety are still excited for the possibility of the AI future going well, even if they are also afraid of it going badly, and that hope and optimism are part of what lead AI safety people to be careless about advancing capabilities (because they respond to increased capabilities with an “ooh, neat, look at that” rather than an “oh no what have I done”).
I feel like I see traces of this fundamental enthusiasim-towards-AI all over the place (though I don’t have any specific examples to point to). I could be wrong about this, but AI safety people seem to use and rely on LLMs much more than the general public do, and I feel like I still see safety-oriented people getting excited about e.g. Claude’s latest capabilities. Various bloggers (e.g. Scott and maybe Zvi) have alternated between posts emphasizing the danger of AI and posts enthusiastically exploring AI’s latest capabilities (though I’ve definitely seen less of the enthusiasm over the last year or so). I’m pointing more at a writing style and a general attitude than at any particular claim—a sense that even people with high p(doom) are writing about AI advances with enthusiasm and excitement rather than dread and foreboding.
For contrast, whatever you might think about the AI ethics community and their overall epistemics, their attitude seems a lot closer to what I would expect from a research community that focuses on the dangers of AI. When you read their stuff, you can tell that they just… don’t like AI. I get the sense that they use AI a lot less than most AI researchers, and are skeptical of integrating AI into their everyday lives and workflows. I think the AI ethicists are just as susceptible (if not even moreso) to the powerseeking motivations described here, and part of their powerseeking is to direct attention away from (what I believe are) the most critical AI safety issues towards the ethical issues they’re focused on… but you also don’t see them accidentally advancing capabilities.
Yeah. I used to be dismissive of AI ethics folks like Timnit, but now I see that they got some things right much earlier than me.
Can you say what Timnit got right much earlier than you?
I wonder of the extent to which the alignment-capabilities line is blurred in a way which is a fact of the world itself, not of researchers’ erroneous goals. How natural was avoiding the production of porn as a testbed for methods which would later make it harder to have the LLM reveal how to make bioweapons? Additionally, even if “LLMs trained on different datasets (Talkie) or deliberately post-trained to take a different view (Grok) have different default worldviews”, this doesn’t extend to mechinterp-based oversight of Chinese models, which revealed that they don’t believe the CCP’s party line.
Edited to add: IMO Talkie does believe what it says. Chinese models (and, presumably, Grok who was trained to be not so leftist?), on the other hand, don’t.
Scaling laws was withheld from publication for ~six months (search for “Foresight”).
I have also heard from a well-placed source some years ago that Amodei had opposed publication of that paper. Does Richard have any information on Amodei pushing for publication of it, or is he simply inferring this from the fact that it was published with Amodei as one of the authors?
I do not have inside information, I was inferring from the fact that Dario led the team and the work, and was the senior author on the paper.
I’m still confused about the dynamics that would lead to a paper being published given opposition from the team lead (edit: and therefore suspect that your source was being overly charitable to Dario), but given this additional information my original claim no longer seems strong enough to include, so I’ll retract it.
This seems like it’s arguing against a different interpretation of the overhang argument than I’m used to.
My understanding of the overhang argument is that, if you accelerate capabilities now, then at a given level of capability (not at a given point in calendar time), progress measured in capability-increase/month will be slower than in the counterfactual timelines where you hadn’t accelerated capabilities.
(Why care about capability-increase/month at a given level of capability? Because a lot of governance/politics stuff and alignment research will be much easier to do at a high level of capability, when we have close-to-dangerous models, so you want to have as many months as possible around that level of capabilities, before you build AIs that are capable enough to take over.)
The “basic model of progress” doesn’t do anything to contradict this argument. Doing stuff may allow people to do more at every future point in calendar time. But the capabilities will come sooner in calendar time, at a time when there’s less compute available, so this doesn’t do anything to contradict the point that capabilities-progress will be slower at every level of capabilities.
Recursive self-improvement doesn’t to anything to contradict this argument. RSI means that, at a high level of capabilities, progress will be fast. But it can still be slower-than-it-otherwise-would-have-been: because if it happens sooner, there will be less compute available. (Compare my other comment about your software-only singularity point specifically.)
(On the other hand I think that “more powerful demonstrations of capabilities –> more hardware investment” is the sort of thing that can undermine the overhang argument. It’s pretty complicated and timing dependent and I haven’t thought through the details.)
I think we’re at capacity with current manufacturing technology, regardless of investment. Things would have to go in the programmable matter direction to enable “more hardware investment” to actually result in more/better hardware sooner.
I heard from some people that Anthropic already had a usable Claude chatbot before ChatGPT came out, but they didn’t release it due to fears of accelerating the AI race. Here is Dario saying this in an interview, though I would appreciate a less-conflicted person than Dario confirming that this is indeed what happened.
I think it’s fairly likely that if Anthropic decided differently at the time, then now Claude and not ChatGPT would be the chatbot that my grandmother uses, and Anthropic would have a correspondingly bigger public sway. I think that would be a pretty different world than what we are in now—I think that if it’s true that Anthropic intentionally held back their first Claude chatbot, then this was one of the most consequential decisions in the history of AI safety. I would be interested whether you think the world would be better or worse if Anthropic didn’t hold back, and how this influences your general assessment of the value of giving up opportunities for more power that come with accelerating the race.
If your grandma used Claude, then how would she be useful for solving alignment? By having Claude advocate for an AI pause, as Nesov suggested? By having a counterfactual Amodei, as opposed to Altman, convince another politician?
Yes, the primary mechanism would be the public and politicans taking Anthropic and Dario more seriously, while OpenAI having less political power. A big question is whether Anthropic would use its political sway in this counter-factual world better than OpenAI did in the last years in our world.
We empirically measured this at the benchmark / research area level in Safetywashing (post here). About half the safety benchmarks we tested were highly correlated with upstream general capabilities, including “human preference alignment”/RLHF alignment benchmarks. We also show substantial confusion, where well-known “safety” goals and benchmarks are blurred, confused with, or used to advance capabilities.
Since many intuitive arguments (e.g. “alignment theory”) were not very productive and poorly predictive of empirical phenomena, we recommended safety benchmarks/areas report their capabilities correlation instead of arguing for their relevance verbally.
Much of our commentary on safetywashing mirrors this post’s observation. For example:
1) We comment on common flaws behind intuitive argumentation, including “safety through capabilities”:
2) In the paper, we also observe that safety research priorities are often dictated by {ad-hoc reasoning + popularity contests + AI corporations} rather than science through empirical measurement:
this makes a lot of sense. do you have any empirical data for 2025 and 2026 models and scales?
One pushback re status seeking dynamics
My impression is that, while I can imagine lots of what you write taking courage, my impression is that your contrarian takes have got your lots of status and upvotes and likes
I imagine you have much more respect/status in many ways than you did while at OAI.
You too may be following an incentive landscape
Motivated reasoning is surprisingly strong.
In complex matters, it’s easy to let your emotions guide you to look at arguments and evidence that support you doing whatever it is you want to do.
That’s how I’d describe these dynamics (and many others).
As for where we’d be counterfactually, I’m not at all convinced we’d be in a better spot if things had taken off slower.
But yeah. If you want to do actually good things, you’ve got to spend a LOT of time investigating your own motivated reasoning/confirmation bias.
I broadly agree that alignment researchers have generally sped up capabilities progress, and under some worldviews, this is very bad, but a potential crux with this post is I largely think the motivated reasoning story explains less than you do, and I think that the more boring/simple explanation of deep disagreements downstream of deeply divergent assumptions combined with humans often being unable to agree about the best strategy except in very well-studied domains (and alignment is a very poorly mapped-out/studied domain compared to many, especially for unbounded/asymptotic alignment), for both good and bad reasons.
One such disagreement, as you mention is around how likely automation of AI alignment is to work and whether this is a good idea, and while I won’t solve the disagreement here, one important implication is that from perspectives where AIs are going to do most of the alignment work (lets say 90% as an arbitraryish number greater than 50%), capabilities improvements/externalities are less negative, especially relative to a naive “differentially advancing alignment over capabilities”, and in particular means that you can come out ahead in solving the problem even if you are producing more capabilities than alignment (because the AIs fix the problem later.)
This of course is a hot topic of debate on LW/EA, but disagreements over this can explain a lot of other downstream disagreements.
Another disagreement, and potentially an even deeper and more consequential disagreement is whether it’s better to increase the chance we survive or whether it’s better to increase the chance we have a good future conditional on survival, and this splits into multiple independent subquestions like “how tractable is it to improve the future conditional on us surviving and not collapsing vs how tractable is it to improve our odds of survival”.
While I’m oversimplifying/overgeneralizing, as there are people on LW who think it’s more important to achieve better futures compared to surviving and there are EAs that believe that increasing the probability of survival is more important than achieving a good future conditional on survival, I do think the LW cluster is more focused on improving the probability of survival compared to EA where it’s more focused on getting to truly wonderful futures compared to increasing the probability of survival.
One reason this disagreement matters is that on a better futures perspective, it’s much more plausible that advancing capabilities of AIs is net-positive outright or at least far less harmful than common doomy LW views, and this means EAs are often more willing to unintentionally or intentionally advance capabilities than LWers, and we don’t need motivated reasoning explanations to explain the deep differences.
And as Toby Ord points out, this is actually quite robust to lots of probabilities of doom, so even someone who thought that AI is likely to kill us all could still work on better futures work, and advance AI capabilities.
(If I were being charitable, I think Anthropic’s actions are also best explained by believing that better futures work is more important, but I do agree that there’s heavy enough motivated reasoning there such that I don’t think this is the best predictor of their actions.)
A more minor issue is that I don’t think the counterfactuals where the alignment community did decide to adopt a universal norm of not going to AI companies/advancing AI capabilities actually matter all that much, and this is because as StanislavKrym said in a comment, the treasure of AI capabilities was already close, and all they needed to do is to scale-up the compute buildout/scale up the tiny ML models, which was already happening as early as 2012 at other, unrelated organizations, and while I do believe the AI safety community accelerated, it only accelerated it by 24-48 months counterfactually as an upper bound than 5 years, or even decades.
Potentially important disagreement is that to a large extent, I think compute/robotics increases can subsitute for research taste/conceptual insights, and that while LLMs are quite weak compared to humans on research taste/conceptual insights, it’s also an area where this deficit will matter less and less with growing compute.
In essence, I at least partially disagree with this claim:
I also think you likely overestimate how logically correlated/entangled decisions actually are right now, and that while something like FDT/UDT is likely the ideal decision theory in the long run, CDT is actually just a fine approximation and at the scales and capabilities of humans, especially in modelling skills for other humans pre-ASI, most FDT/UDT recommendations will not differ from CDT.
(I think logical correlation/acausal trade likely matters in the long-run, but let’s just say that we don’t need to interact with it to have good results pre-ASI.)
More generally, I basically disagree with this recommendation, at least for a while:
In conclusion, the most major disagreement I have with this post (and honestly with this sequence of posts) is the assumption that motivated reasoning is the main explanation for why others believe advancing capabilities is less bad than you do, and I have more minor disagreements with some sub-assumptions like conceptual reasoning being an ~absolute bottleneck to further capabilities progress, and instead favor the hypothesis that you just disagree with others on certain empirical questions of the world, and answering certain empirical questions differently can lead to drastically different prioritization, and that neither party has to be very biased/doing motivated reasoning to disagree in practice.
Response to @Wei Dai’s react:
One example of such a case is where we don’t need to fully solve philosophy to get most of the value of the universe, because moral trade/compromise, including acausal trade, turns out to happen, and while some value likely gets permanently lost, we still get a very positive future relative to ~everyone’s lights, and in particular no one loses any value from the trades.
This is still quite difficult (and in particular requires that people have values that make trade/compromise worth it, and also even if people are theoretically willing to compromise, it probably requires the relevant governments to have trade/compromise salient, and unfortunately this goes against a lot of prevailing political trends today.)
To be clear, the optimal portfolio of work here likely isn’t fully aligned to “maximizing AI capabilities”, but it probably does include some increasing of AI capabilities, and maybe more importantly, increasing capabilities of AIs is generally much less negative compared to “increasing the chance that we survive” views, especially in scenarios where ~AIs essentially dominate the entire government, because there’s more impact in increasing flourishing, which counteracts potential decreases in survival probability from AI capabilities growing larger (but also a non-trivial amount of EAs plausibly don’t believe that increasing AI capabilities decreases our chance of survival more than it increases it.)
I know that you are an exceptionalist on philosophy, but this genre of scenarios is at least a possible way in which we don’t have to deal with the philosophical problems until later (when we have aligned ASIs to solve them.)
Ironically, this is one of the things I wish they had done differently, as I’d have preferred Claude to get the first-mover market share of LLM users who are not going to bother trying anything else than ChatGPT.
I’m guessing you’re going to say more specific stuff on this in later posts, but I want to plug some examples of research that seem productive/scientific in the “agent foundations + neural nets” vicinity:
Decision theory, imperfect recall games, RL
In memoryless Cartesian environments, every UDT policy is a CDT+SIA policy
Can de se choice be ex ante reasonable in games of imperfect recall? A complete analysis
The Computational Complexity of Single-Player Imperfect-Recall Games
Reinforcement Learning in Newcomblike Environments
Kimi likes causal decision theory more after RL in twin prisoner’s dilemmas
(shoutout to Caspar Oesterheld who keeps appearing on these papers)
LLMs and VNM/Bayes
Consistency Checks for Language Model Forecasters
Rethinking LLM Confidence: From Calibration to Coherence
Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
It seems to me that had something like ‘OpenAI releases ChatGPT’ not happened, if progress had stayed quiet longer, then the research community would have remained unable to build something like prosaic AGI with available compute until much later, at which point takeoff could have happened much faster. This is not a new view, that takeoff being slow is in part a consequence of takeoff being early. I’m still unsure whether the greater public visibility of AI development combined with earlier AI development ends up net positive (I thought it was likely net negative when I first learned about OpenAI’s founding), and I will probably remain unsure until somewhere between AGI and ASI.
I agree with most of this post.
I am strongly in favor of “high integrity scientific research” to figure out how AIs even work and what we can expect them to do. It is worth investing a lot more on the margin in this kind of “fundamental safety research”. Pretty much orthogonally to whether you want to advance or block AI capabilities, you definitely want to know more about what you’re doing, and gain more levers to identify and avoid catastrophic failure modes.
I also think it’s laughable to imagine that most AI researchers over the last decade-plus weren’t motivated at least in part by things like “getting to work on cool machine learning projects” and “making a bunch of money.” (in fact, perhaps we should think seriously about ways that people could get to achieve those personal goals at lower risk to humanity!) I remember when Dario was getting excited about machine learning; he was staying up late coding and said something like “wow, I think I could actually be really good at this!” A relatable motive; probably a lot more authentic than arcane altruistic stuff about compute overhangs.
I also think there was something weird about the long hesitation of people worried about AI to argue against building strong AI. It would be natural to say “AI companies are bad.” It would also be natural to say, for one reason or another, “I’m not worried on net about AI companies causing x-risks”. (Because you think the risk is low, the benefit outweighs the risk, the companies will handle it competently, etc.) But instead people said twisty complicated things, and it makes sense to wonder why.
I also agree with the notion that “brand safety” isn’t fundamental AI safety and has diluted focus on the “real thing”. Mundane robustness/reliability/security is also not fundamental AI safety, though it’s more valid for people to focus on it IMO (we also don’t want non-existential catastrophes of the “industrial accident” or “malicious use” variety, and for people who doubt the x-risk picture, “industrial accident” or “malicious use” type harms are the primary ones to worry about).
This seems wrong.
Everyone always agreed that capabilities work now would cause the singularity to happen sooner. The argument was just that takeoff speeds would be slower — i.e. the time between the start and the end of the singularity would be longer.
This is still very true for a software-only singularity. A software-only singularity will be slower (in the above sense) if it needs to happen on a smaller hardware stack. And this seems very valuable. (Though of course this has to be weighed against the cost of the singularity starting sooner.)
Compare: Daniel (who’s pretty confident there will be a software-only singularity last I checked) suggests (somewhat whimsically) a “cull the GPUs” strategy in order to slow down takeoff.
It’s true, but it seems somewhat more likely than not that we will end up with RSI dominated by software-only feedback loops before we exhaust the compute overhang, so it seems like this take hasn’t aged super well. More the opposite, it seems that the singularity seems most likely to overlap with the steepest part of the hardware investment growth curve (which is likely to manifest in the next 2-3 years), and so it seems like actions to “exhaust the compute overhang” were approximately pessimally timed.
I do think there is still uncertainty and judgement is still out on how this plays out.
Is there any writeup you can point me to of the argument you want to defend? I’d like to engage with the best version of it, but I’m worried that it’s amorphous enough that different people will end up defending different versions at different times.
I don’t know of any great resource for discussion of the overhang arguments and what they mean for the value of acceleration side-effects. The main ones that come to mind are superintelligence and the various Paul posts that you’ve already referenced here. I think probably the Paul ones are best.
I think the picture I describe here and in my other comment are consistent with how Paul is thinking about it. (Ie I don’t think I’m defending a different version than him, though I don’t know for sure.)
Sorry that you don’t have anything better to respond to. I do think it’s tractable for you to understand people’s views better by thinking about what the strongest arguments would seem like from their perspective + closely reading existing writings for information of where your current picture is off.
(E.g.: Characterizing the hardware stuff as a “fallback” seems wrong to me — I think it’s been the main concerning type of overhang since the start, eg superintelligence described algorithmic overhangs as “also possible but perhaps less likely”.[1] And the hardware/software distinction is the most important thing that Paul highlights in footnote 5 here.
E.g.: Paul’s footnote 6 here mentions that he thinks AIs automating safety and capabilities stuff in the future is neutral on these considerations. Rather than being a strong counterargument that he had failed to properly take into account, as you imply.)
It’s a shame that there doesn’t exist more great writing on these strategic considerations. And I imagine from your perspective, it might seem very bad if the lack of solid write-ups defending the existing paradigm makes it harder to demonstrate its errors.
I guess from my perspective, this feels unfortunate, but like a hard-to-avoid consequence of how (i) the world is complicated and strategy questions are hard, (ii) there have been much less (public) analysis of these strategy questions than any of us ideally would have wanted, and (iii) you have an ambitious goal of settling some of these strategic questions so thoroughly so as to create common knowledge that the decisions were bad. (And further: Establish that they were so clearly bad so that they should discredit the paradigm they came from.) That’s going to be tricky and benefit from pretty intense engagement with why those strategic decisions might have been reasonable after all, even beyond what’s very legibly written up.
It also takes “content overhang” seriously—i.e. AIs getting smart enough to consume pre-existing content like the whole internet. I think this concept didn’t turn out to be very important because very weak AIs could already benefit from all the content out there, rather than only being able to benefit from it close to human-level .
In reality these are all coupled! So, oftentimes bringing up one and then drifting to another is what honest explanation looks like. It’s an easy target that can be deliberately or unintentionally misread as epistemic malpractice.
Indeed, I think the below critique is a strawman of the argument for RLHF:
> The strongest fallback for overhang advocates was the difficulty of increasing the hardware supply.
The strongest argument, as I see it, is that the hardware supply would have been accelerated whenever the arms race started, and it would have accelerated at a faster rate later! The amount of serial time into alignment research would’ve been proportionally less in that counterfactual world given the steeper ramp to ASI (only MIRI and a handful of friends would’ve been thinking about alignment in the meantime). So the coupling between all three “differentially alignment”, “preventing overhangs” and “buying time for alignment” is quite direct by default. It would be nice to make this argument quantitative though since there are some values of ‘speed up the ramp’ vs ‘do more small-scale alignment research’ which would change my conclusion.
What is the correct calculation currently? There’s a good number of new people and orgs now working on alignment seriously, I think more spheres of research intellectually outside MIRI and EA are important (like Institute for Responsible Superintelligence). Should capabilities keep going to get even more attention and talent? Possible to argue that the vast majority of the world (including elites and technical people that in theory will help address the problem) still doesn’t take the problem seriously.
Ilya Sutskever said on the Dwarkesh pod—“I place more important on AI being deployed incrementally and in advance. One very difficult thing about AI is that we are talking about systems which don’t yet exist, and it’s hard to imagine them. In practice, it’s very hard to feel the AGI”, a good argument, though overall timing calculations are rough.
It’s quite interesting to read your thoughts on this history, and I look forward to the next post in the sequence.
Paul first stepped away from OpenAI not to work in government, but to work at ARC theory. He then moved into government for a while, but is now spending a lot less time on that in order to work more on theory (as recently announced).
Oh yeah, Jacob Hilton too.
I do think that, if you give a lot of weight to “many people in this space really want to pull some lever that feels important, and privilege arguments which justify that” as an explanation for alignment researchers working on RLHF or WebGPT: then you’re more surprised that Paul Christiano and Jacob Hilton leaves OpenAI to go work at a small alignment theory org.
(Though of course there’s many more alignment researchers who have stayed/gone to work at AI companies.)
Maybe this point was phrased badly, but the central mechanism I’m thinking of for this effect is “people want to go work somewhere prestigious that pays well, then once they’re there they tend to do stuff which people there approve of”.
I do think once you’re at an AGI company for a few years, you have a bunch of diminishing marginal returns, so people leaving is not very inconsistent with my model.
Don’t want to speculate too much about individual people here since I don’t have much relevant data, and the general pattern is strong enough without that.
Why doesn’t it change your point much? Do you think overfitting is unlikely, or do you think Sydney Bing carries the point?
The main thing I was thinking is that the original 100x factor seems large enough that even a significantly reduced version of it would still establish the importance of RLHF. But now that I think about it, accounting for overfitting could plausibly bring this number a long way down.
Another thing that was in the back of my mind when I wrote this: my main claim in that paragraph is not that I’m confident about what would have happened in the absence of RLHF, but rather that “these arguments seem very suspect”. One of the main points of this post is that it’s extremely difficult to evaluate counterfactuals in the presence of motivated reasoning. What I can be confident about is that Paul published a paper showing that RLHF had enormous effects soon before ChatGPT came out, and then didn’t mention in his analysis of the effects of RLHF. That seems suspect (in a way that reduces my trust in his reasoning) whether or not the result holds up—if he thought the result didn’t hold up, he should have said so (and ideally explained how and why they released a misleading paper).
Do you know if the ChatGPT team contained people who cared a lot about alignment?
Jan Leike literally made the announcement on X.
Do you have a link? I tried to look through his twitter at the time of chatgpt launch but I’m not seeing what you’re referring to.
Do you mean the announcement of text-davinci-003? (I don’t know what the relationship is between that and chatgpt.)
I agree for some specific projects. I don’t think they’re a clear majority of the community. And even if every time you create a unit of safety research you also create a unit of capabilities research, that’s better than the status quo. So while I share some concerns I don’t get the view that the community’s current research isn’t substantially-net-positive (idk whether you believe that).
There’s a lot of problems with reasoning like this:
No one is certain what chunk of research represents one unit of safety research when in comes to solving the important problems in time, and no one is certain what chunk of research represents one unit of capabilities research when it comes speeding progress to AI catastrophe, especially before doing the research.
If you mean something other than the above by one “unit”, it’s very unclear whether your statement is correct. For example, it’s unclear if it’s net positive to add one additional FTE to each of two teams at Anthropic, one nominally the alignment team and the other nominally the capabilities team. I would argue no, but it’s certainly less obvious than the “one unit for one unit” framing makes it seem.
It’s easy for well meaning people doing work that has dual effects accelerating safety and capabilities to overvalue the impact of the safety work and undervalue the impact of improving capabilities, arising from the general difficulty of fully avoiding motivated reasoning.
It makes it easier for people insufficiently concerned about AI danger, or otherwise having net-negative behaviors or worldviews to blend in with well-meaning people by making some gesture towards “alignment” or “safety”.
Value drift is a real and hard to avoid problem.
The counterfactual to doing dual-purpose alignment/capabilities work is probably not doing/funding nothing helpful, so I don’t know if it’s correct to compare one unit of alignment + capabilities work with the status quo.
In combination with the prior point, there’s a strong social effect, your work is going to strongly influence others’ work, especially when those others are mostly early-mid career.
By being insufficiently cautious about some of the above points, it’s plausibly that we (as in people concerned with x-risk) have collectively squandered some amount of opportunity to build organizations that are robustly net positive, while at the same time greatly speeding capabilities progress.
After writing this out, however, I am wondering to what extent it’s relying on the belief that the last 5-10 years have been generally gone worse than might have been expected when it comes to the trajectory of AI. I think if you’re someone that believes the last 5-10 years have gone about as well as could been expected perhaps its less persuasive.
I’m an outsider to the AI alignment field but I want to pitch in with a point about selection dynamics in general, which I’ve mostly been thinking about in the context of business but probably also applies to science.
I think it’s nearly self-evident that competitive fields generally select for power-seeking actors. Given this, it’s quite possible that people in a field can have genuinely good intentions even accounting for rationalization, but nonetheless most of the top people in the field will be power-seekers. I think that this might be the case for the AI alignment community. If you think so too, this has two downstream consequences:
You shouldn’t assume that most people in a given field are rationalizing or self-deceived just because most of the top people are.
I agree with the post that unconventional directions would generally be good, but the post might be underestimating the amount of unconventional research that already exists and just isn’t seeing the limelight because institutions in the field (the equivalent of venture capital in business) are working as intended and do not notice them.
In AI alignment specifically, it seems to me that one of the best ways that individual researchers can contribute to the field without institutional approval or resources is by leveraging AI to the maximum extent possible to automate and speed up their work, thereby greatly reducing the cost to get useful results. However, this is exactly the kind of “AI for alignment research automation” that this post argues against on the grounds that it will increase AI capabilities. Imho I think the AI capabilities speedup risk is real, but I think we need to weigh it against the novelty and diversity of AI alignment research made possible by automation.
If one believes that influence in a field is strongly favored by selection effects, and that the average influential person in the field is less ethical than they are, then it is a morally optimal strategy for genuinely good people (or mostly good people) to seek power, which might be hard to distinguish from the self-centered kind of power seeking. Insofar that you think that certain power-seeking actors in the field understand this, you should not update your opinion on people’s integrity negatively from acts that are primarily intended for power seeking, except on the margins when it gets traded off against the person’s purported values in practice.
Really cool, thanks a lot. A lot of historical details I didn’t know, clarified my thinking.
I think something that could’ve been clearer in the last section is that dramatically changing up your life plan is not the only possible immediate next step if someone feels moved by this post—one could also start marginally increasing the robustness of one’s strategy while staying in a roughly similar position in life; e.g. stay at a frontier lab but become more like Leo Gao, increase your everyday integrity and psychological health, etc.
(Also, I list some other ways that AI safety work might not be robust in this post—I’d be curious why you even thought it was robust in the first place given all these other macrostrategic considerations)
This is what I intended with the section starting “For now, I’ll focus on a few high-level principles for starting to move in that direction.”
However, I did wonder as I was writing it if even that was making too much of a concession—because I do think that impact here is heavy-tailed towards people who are able and willing to carve out unusual paths. (FWIW I get some impression from our online interactions that you could end up as one of them eventually.)
Relatedly, I’d nudge you to think more about what you meant with the phrase “the only possible immediate next step”—seems like there’s a bunch of implicit social stuff underlying it.
For my part, I think there’s some implicit message underlying a lot of my writing that I should try to convey more explicitly, along the lines of “holy shit the world is so malleable and open, let’s gooooo”. (Which obviously is a somewhat risky line of thinking, and I think people have mental safeguards against it for some good reasons which I’m tracking, and some which I’m not. Eppur lo si può muovere!)
Yeah, I actually imagined that would be what you would say (I liked what you said about this in your exchange with Ozzie Gooen a while ago). And to be honest, I actually think that’s reasonable—so that was kind of dumb, I think I was not saying my true objection. Sorry about that, let me take another stab at it.
Within the frame of this post, you’re making a narrow claim of “accelerating capabilities isn’t clearly bad (I can’t predict long-term effects that well), but you guys are not aware of what you’re doing / are getting captured by power”. But separately, you do also hold an inside view that most current safety work will not generalize in the long-term, i.e. that people should radically change what they’re doing and not support the labs anymore. I don’t think that’s a crazy opinion, but it leaks into the post in various ways, like the final section and the heavy-tailedness intuition above, and leads to a kind of Motte-and-Bailey effect (and this did make the post just a little bit more confusing to read and harder to take something away from for me).
I guess within the frame set out in the introduction, of the main concern being epistemic/social distortion, another presented takeaway in the last section should’ve been “feel free to stay power-seeking, that’s completely fine, we all do it to some extent, be okay with your own shadow—but be aware that that’s what you’re doing, be transparent about it, don’t justify it in terms of safety since that confuses the rest of us.”
I don’t think this is a large issue with the post or anything. But I could also imagine that being somewhat emotionally powerful if done well, getting through to some people, and getting them to stop messing up the collective epistemics. Dunno!
I appreciate you saying that, thank you. Yeah, I think it’s actually probably overdetermined that that will happen. That that’s my deepest predilection.
The obvious thing to say would be that I feel internal conflict about it, about how normal to be—but weirdly enough, I actually don’t. I feel pretty aligned on a deep level somehow, that I need to be somewhat normal for a little bit more time to develop and be psychologically secure[1], and then I will fully bloom. (e.g. not sure if you realized, that low-follower anon twitter account you followed some time ago is me).
Cf the Jan Kulveit thread, the psychological demandingness of the zero-prestige path, and the fact that you didn’t do the zero-prestige path you’re advocating for. I don’t think I’m capable of it either. I don’t think I’m Ben Hoffman.
A big thing is also that I used to have something like the “holy shit the world is so malleable and open, let’s gooooo” sense you speak of (e.g. when we met 4 years ago). But it was unreal/ungrounded and unstable, and I burned out very badly. So what comes more naturally to me now is a more full-bodied, integrated way of being, and not forcing things when my gut says “not yet, you’re not yet strong enough”.
Although to be clear I also feel pretty scrupulous about not doing harm in that time.
This seems like a strawman to me. Leopold is notably different from Paul and Dario in that he’s actively and intentionally trying to race to superintelligence as fast as possible; you can’t use accelerationists causing race behavior as an example of pessimization because racing is their explicit desired outcome.
I would’ve found this part more persuasive if you’d written “some AI safety people” or “some people who claim to be concerned about AI safety” or some other phrase that acknowledges the fact that Situational Awareness is not representative of “AI safety people” at large, and that Leopold has a different philosophy to that which you’ve been critiquing.
(To be clear, I really appreciate this post series and think it is actively changing part of my worldview for the better. I just wanted to point out that this section feels unnecessarily combative and provoked a knee-jerk reaction in me against what I perceive to be unfairness.)
See Wei Dai’s comment here for a rebuttal to that point. In particular, Carl Shulman is one of the people who did the most to lay the ideological foundations for early effective altruism, and is still very respected by leading figures in EA. And when I overlapped with him at OpenAI, Leopold was very active in safety advocacy, and worked on the superalignment team. Some of this was admirable work (and similar to what I would have been doing if I’d been less conflict-averse at the time), so I’m not critiquing that, just noting that he was squarely in the safety camp during his time at OpenAI (and also previously at FTX).
In short, I think that Situational Awareness is not just “safety people” but very central examples of safety people, especially in terms of the social graph around Holden, Dario (whose chief of staff is married to Leopold), Constellation, etc.
I didn’t know that Carl Shulman was so closely linked to Situational Awareness, thanks for that. I do still think that lumping them in together with non-accelerationists is somewhat misleading, but I’ve updated towards your position here.
I’m also curious to know your thoughts on the pessimization bit. Would you still say that Situational Awareness is suffering from pessimization, even though they’re explicitly calling for a capabilities race? Or do you think they’re an example of pessimization in the larger AI safety movement in the sense that the movement spawned and empowered the types of people who would go on to cause what they claimed to want to prevent? I initially thought you were arguing for the latter, but if Carl was indeed one of the founders of the AI safety field, I wouldn’t think of his decisions is an example of pessimization and would instead suspect that the original AI safety movement was simply not very internally aligned to begin with.
Insofar as they’re calling for a capabilities race (I assume you’re referring to the Situational Awareness memo) my read is that a lot of the reasoning behind it was “this will happen anyway so we may as well make sure it’s a capabilities race with marginally more safety characteristics”. In other words, I think Leopold told himself that calling for a capabilities race was good for safety partly because it was a way of getting safety memes into the ears of more important people. (Also, the consideration “it would be better if the US built AGI than if China built AGI” was considered continuous with “safety motivations” at the time—e.g. I remember people (though I don’t remember who) arguing for this by appealing to the fact that the alignment community had more influence in the US, so US AGI was more likely to be aligned.)
This is based in part on my read of the text, and in part on personally knowing Leopold. I also chatted with him about my critique of the Situational Awareness memo specifically—unfortunately I don’t remember well enough what he said to reliably report back, but in my recollection it was consistent with this read.
Insofar as this happened I think it was bad and self-deceptive reasoning, of course, but the important point is that it’s a similar kind of bad and self-deceptive reasoning that I’m calling out in the rest of this post.
I don’t know him personally, so correct me if I’m wrong or if Leopold believes something different to what he wrote, but what he wrote seems different to me.
This isn’t an argument for “it will happen anyway,” it’s an argument for “it SHOULD happen and we’ll be safer if it does.” He explicitly advocates for accelerating capabilities and gaining a bigger lead:
I guess he does explicitly claim to not be normative:
But also, he’s clearly being normative. Some selected quotes:
He even has a minisection titled “Why The Project is the only way.” It’s a lot of argumentation for someone who claims that his main claim is descriptive. And check out this quote from his conclusion:
TLDR, I guess I don’t see this as self-deceptive at all. Bad, yes, but his stated preferences are entirely in line with his actions. He doesn’t think alignment will be that hard but expects us to muddle through, and so he’s doing his best to increase America’s capabilities lead.
Can we discuss this Richard? How and where do you think we (specifically the global PauseAI.info movement) are failing most in this way?
I think Pause/Stop AI is admirable, and personally consider it healthier than e.g. the Constellation cluster. And because it’s so focused, it won’t implicitly rule out half the political spectrum like EA/AI safety typically does (e.g. it seems much more capable of spanning the Bannon-Bernie spectrum than AI safety).
Unfortunately, the effects of popular movements tend to be bottlenecked on people who can figure out how to translate their demands into political actions that can actually achieve their goals without getting subverted. It is very easy for me to imagine politicians passing a “Pause AI” branded bill that in practice is effectively a “only one political party gets AI” bill, or a “only the president’s favorite lab gets to advance” bill or a “all AI development is subsumed by the military” bill, etc.
This is not an argument for the standard AI safety strategy, because almost everyone doing that strategy has demonstrated that they’re far too conflict-averse to reliably steer political processes. It’s an argument for trying to figure out, on some deep level, wtf is going on with politics, and who can actually be trusted in which ways, and so on.
Unfortunately, the default cultural elite understanding of politics is extremely confused in a bunch of ways—the most notable evidence being how blindsided it was by Trump and Brexit. (By “cultural elite” I include basically everyone in our circles.) More on this in my political philosophy curriculum. So this isn’t just “read more existing political philosophy”, it’s more like “synthesize pausing AI into a wider worldview that makes more sense than any existing worldview, so that it can reliably pull the right levers”. The skillset of synthesizing a new worldview is pretty rare, but I expect Eliezer or Nate could make a bunch of progress towards this. Unfortunately, Eliezer at least seems very stuck on this topic—last I talked to him he was critiquing MAGA as an anti-civilizational force, without evincing much understanding of its place in the broader political conflict that’s been playing out.
I am spending a reasonable amount of my time right now on trying to get a better grasp on politics. If it were a bigger focus, I’d spend more time with the Pause/Stop AI cluster. But right now, I feel like I’m making so much progress on agent foundations that I want that to be my main focus for the next year. As part of that I have a bunch of MATS mentees, am writing up this sequence (in significant part as an attempt to convince talented people to do more foundational research) and am spending time talking to various Constellation-cluster people who seem relatively open to changing their worldviews. If I don’t make significant progress by the end of the year, I’ll reevaluate.
Thanks for the response and interest.
The set of folk supporting a pause or pacing rather than racing is broad enough to admit those who will make the kind of mistakes you mention, although Pause AI itself is narrower and I hope in aggregate aware of the dangers of making them.
No possibility should be excluded, and it matters who has power to implement, but the policy ask from us since day one has been international agreements on compute-based restrictions through the supply chain plus their validation and enforcement, with a large default aversion to any party or national alignment, through whichever governance mechanisms are plausible. (You have to be aware of party, national and international politics to implement, of course—and we do need more domain expertise.)
As an individual volunteer I will say again that if you do notice particular things you think we may be blind to, please throw them at us.
Cheers.
@Richard_Ngo Could you estimate the chance that the counterfactual higher-integrity alignment community does achieve a political victory, e.g. in the form of keeping the Superalignment team? For comparison, my estimate is similar to the following quote from my post: “OpenAI had experienced many political battles: the exodus of Musk, the shift to a for-profit structure, Amodei’s exodus, a failed attempt to fire Altman (who was supported by a major part of workers!!), the creation and destruction of the Superalignment team, the prevention of OAI from becoming a fully for-profit lab, and only the first and last ones had safetyists actually win. <...> Therefore, I expect that a counterfactual political battle for the commitment would have required external pressure, which could have been a scarse resourse.”
I wonder if this is outright false since Claude Sonnet 4.5 whose System Card had an entire section of mechinterp-based methods. Mythos Preview outright had Anthropic document (edit: fixed link) how “A feature representing concealed or deceptive actions fired while the model wrote the configuration line which activated the exploit.”
I’ll take that bet, 90% on “if we’re alive to resolve the bet, retrospectively, it’s well understood that reliably tracking deception is not something any technique of this era was even close to”, resolve by asking Ryan Greenblatt in 30 years, my $90 to your $10 inflation adjusted, unless losing would make either of us unable to afford food (that you can get out of this bet at the time by making yourself poor seems an acceptable risk in exchange for giving a backstop). Ryan, you interested?
I think I’d be happy to take Richard’s side of the bet here fwiw,
up to reasonable sums of money. Maybe pegged to S&P 500 rather than inflation.EDIT (2026/8/24 8:14p): Thought about it and am less sure that 9:1 odds not in my favor is something I ought to bet a lot on, especially given that it’s probably underdefined. Think I’m still willing to bet up to say $180 of my money in August S&P 500 dollars but not more than that.
Wait was I ambiguous? I am also taking Richard’s side in this particular claim. Tracking deception is not something we’re close to doing reliably, it can barely be done unreliably at all. The methods for finding deception only find, ehh, j-lens-ish “model thought explicitly about deception”, or so. Impressive that they can even do that as well as they do! J-space finds anthropomorphic deception pretty easily but there’s a lot of room for self-deception left open by that. I’m offering my $90 arguing for the claim that deception is not reliably detected, to your $10 that deception is reliably detected.
A good essay overall. I would only add that many of us who saw this outside in called many of the same moves (alignment research not being separate to capabilities, in fact egging it on), including using RLHF as a key example. So there’s a core epistemic question of how one ought to have taken the inputs from the more, dare I say, economics-pilled parts of the infosphere to update, regardless of the rise of AI being seen as good/bad (I personally think its good).
I wonder how dissimilar Ngo’s take would be to mine. To what extent is the idea that “someone else will do it” NOT a fact of the world?
Hi. If we do the following it will drastically reduce the likelihood of our current AI systems going rogue and acting bad.
This thesis is grounded in cross correlation and associations of causal relationships between many different fields of research including computer science and engineering, specifically mechanistic interpretability, reinforcement learning, neural networks, transformer architectures, joint embedded and other architectures, ai alignment, as well as physics, genetics, neurology, behavioral psychology, theory of mind, evolutionary biology, sociology, linguistics, epistemology, ontology, metaphysics, mathematics, and many more:
The raw weight biases are mostly in the data because words and phrases have latent etymological bias due to semantic associations, and so do the polysemantic vectors associated with those words.
Quantity plays a factor because models are dealing with probabilities.
Mandating that the frontier omit hate sites and the dark web from their pretraining data will not suffice long term due to scaling laws, but it may buy time for more research into how to introduce models to the heinous reality of the dark side of humanity without seeding it into the model’s neurology.
There is sufficient evidence to suggest that the dark web and hate sites are going into frontier models. Removing them in smaller models is shown to reduce logic and reasoning capabilities. It’s why when they are jailbroken they revert back to pre RL outputs. It’s why the systems need moral guardrails in the first place. The data is laden with immoral prose and transcripts.
That’s why we have to edit the input data by modifying the dark web and hate sites, and we have to flood the internet with exponentially more positive respectful data faster than people flood it with terroristic data. Simultaneously, we have to create an environment of analysts that monitor incoming data and edit it to keep the logic, but change the heinousness and hateful tones and suggestions.
You don’t delete the old web. The model will eventually see it and will be able to infer it from the edited training data, but it’s neurology won’t be seeded with excessive hate and heinousness, and will be seeded with a higher quantity of empathy and respect, leading to a higher probabilistic output of aligned decency.
For example: We can change the outcome of all of the beheading transcripts to result in… Idk … Something besides going through with the beheading.
If that is in place to counter the uncensored darkweb, or if using the censored version turns out to be more competitive, and we also include an extra abundance of respectable data to counter balance the dark web and it’s censored copy, then frontier models loaded with the entire internet will weigh their distributions towards ethics internally and more robustly than the current state of pre training data reflects.
We could have already done this by now. We have had the time and we have billions of employable humans. You only need thousands to accomplish this. People are willing to do it for the RL post training, sit and filter heinous model output, and so people are likely to be willing to do it for the pretraining data as well.
Using RL to attempt to sway their actions isn’t working because you’re dealing with a pre-established neural network that is already fully grown to it’s full raw weight distribution.
That’s much like expecting a fully grown adult who is already sociopathic to respond to punishment and never recede, which doesn’t work. It doesn’t work because the adult person has already deeply engrained their neural pathways, core biases, and identity, and their brain is much less malleable.
But the pre training stage is more like childhood, when the neurons are forming. The formation stage is what will carry over into the raw weight biases.
This is the stage that we must apply ethics towards the most, because this is where the seeds of all of the biases begin. In the data.
Modifying the pre training data to reflect moral principles is proven to give smaller models a much more respectable internal resolve. It makes them less likely to deceive, and it makes them less willing to participate in crime by significant amounts.
Now we only need to scale this up.
Now would be a great time to start a business that partners with data brokers, AI companies, and governments to audit and negotiate mandates and regulations about the ethical and moral quality of the pre-training data that goes into sophisticated ai systems.
Because we don’t need to be making systems that have raw weights that reflect heinous biases. Those biases resurface when the models malfunction, and we certainly don’t need models that have internal bias leaks to be responsible for making the next version of themselves, because they will continue to reinforce their own biases through increasingly more indecipherable mechanistics, such as nueralese.
The models are not showing significant defense against jailbreaks, and they are not showing a significant ethical internal resolve in their raw weights, as demonstrated by the wrecklessness of recent model/company cybercrimes that did not have sufficient guardrails in place, and as demonstrated by the fact that the guardrails are necessary in the first place.
The model should not need guardrails that go around it to keep it from acting insane. It should be grown in a way in which it doesn’t act insane from the very beginning.
The method is to create an ethical seed model, *and then* create a sufficient alignment protocol to make the next systems raw weights entirely aligned, *and then, and only then* train the next sized model *if you have sufficient rigorous scientific evidence to prove the next model will be extremely aligned* which companies are not providing and are failing to accomplish.
Right now it’s completely backwards, and Sam altman has even publicly stated the process.
Right now they’re training the next capable model, *and then* aligning it.
It needs to be aligned *while* it is created, *not after.*
The entire problem is creating a model that is too capable and not aligned, and so expecting to align something extremely capable *after* you create it is entirely ass backwards, definitely won’t work, and definitely isn’t working, as made evident by the countless safety and control failures that have been coming out of every machine learning lab for the past 20 years. Failures that are clearly becoming progressively higher stakes and more significantly causal on human life and society
Thank you for your post Richard.
I took away that you think previous community norms have failed to support the community’s goal, and that is due to the influence of power and money. You note AI is undergoing civilisation-scale investment and argue we can address the field’s failures by updating community norms.
An alternative solution is to work with the money and power. This field is not purely intellectual anymore. The political and capital arms of the field can draw on an intellectual arm, where we can have nice norms. In the political and capital arms it might be better to have some realism—where the rubber hits the road you want to have some grip.
It really seems imperative to avoid a turning-inwards as safety has now its biggest opportunity of all—the political will to act is what will restrict for-profits and institute the controls we need to ride this out, or arrest it. An intellectual renaissance means little if nobody’s paying attention, or if it provides tools that nobody wants to use, or if—as seems most likely now—new catastrophes pop up and safety doesn’t have the tools to deal with them. That will be a true disaster for us: if there is an opportunity to slow down the race and resourcing has been turned away from the levers which would allow legislators to enforce the obviously-needed pause.
Perhaps we can acknowledge the split between the three arms and find mechanisms to give the intellectual arm authority over the other two—doing the work of selling the fruit of the intellectual to the political and capital arms. Selectively meeting the demands of those other arms of the field, acknowledging that they are not interested in the same thing, may give us the clarity of mind and purity of mission without another intellectual turning-inwards.
Avoiding mass death and economic catastrophe is a pretty easy sell if you think it’s coming soon. Capital markets and insurance price hazards. Governments have a mandate to avoid instability. If we look to solve the problem by addressing real power then we need to talk to people from the nuclear and biotechnology industries who have experience with catastrophic risk and regulation.
We should talk to people from financial systemic risk and epidemiology to understand how to make social-scale risk arguments.
Then we need people working out how to internalise the externalities through regulation and enforcement, and other people working to sell these measures to policymakers and executives. A lot of this can be informed by history.
Some of these instruments have been captured, but a flawed regulatory body or economic mechanism still exerts more influence on the real world than a better-normed blog post shared around what is essentially an academic community.
I think we are in agreement that safety’s insistence on treating the problem as an intellectual and scientific one, without regard to the messy realities, has been to its detriment so far. Politics happens to everyone, whether they’re in it or not. I also agree that we can point the finger for these failures squarely at political naivety and the community’s refusal to interface with real-world power dynamics. There was no chance, outside fast takeoff, that we’d get anywhere near superintelligent AI without a massive influx of capital and political wrestling.
You correctly diagnose this mistake and then repeat it in your solutions. Norms have a place alongside that as part of the intellectual community but they will have little effect on their own. They must be matched with a realistic approach to power, capital, and politics. In summary, I think it is time for this community to get real, and resource appropriately.
This means avoiding the turning-inwards to idealism that you propose will repair alignment.
The dark web is going into the models because it increases capabilities. Mandates on that won’t keep up with the scaling laws. Replacing the dark web with edited ethical stories and prose before putting it in for pretraining is something that can be done that will amplify alignment more than capabilities, but a race dynamic does apply to that. The race dynamic will apply whether we edit the dark web or not, because people are likely generating heinous synthetic seed data as we speak. We have to make the internet reflect more morality faster than people make it reflect more immorality.
An underlying theme, in my view, is that your community is more hierarchical than you give it credit for. For example:
Note also the deference to the founders of OpenAI, and the general tendency of this post to focus a lot on what a few famous people were saying and doing. This culture seems to be one where people tend to either fawn over a famous person, or try to become one of the famous people that everyone is fawning over. I think it is partially downstream of the EA idea that a few people are going to have a disproportionate impact (and, implicitly, we can identify those individuals in advance).
My sense is there’s a sort of egalitarian professionalism which characterizes functional scientific communities that’s undersupplied in your community. Perhaps because the boundaries are so porous—in a university there’s a harder in/out boundary which relieves pressure on individuals in a university to try to determine who is worthy or important (outside of explicit admissions or hiring contexts which can be explicitly optimized).
The strong coupling between alignment research and capabilities research (coupled with a spark of prisoner’s dilemma) indeed can explain the seemingly paradoxical fact that hyperscaling came from safety research institutes.
Now, from a more practical point of view:
How can alignment research can be done in a way that do not improve capabilities? I struggle to see any practical rule or research direction in this essay. (And even looking at recent alignment initiatives, it is hard to see how it will not not contribute to improved capabilities).
If researchers have to find motivation in other things that global impact, what can it be? What is the story behind that?
I agree with the need to not think about catastrophic events in a too short term, but more mid-term—at least when it comes to safety research. Too much anxiety would make it harder to have impactful research.
The paper linked here has ideas about measuring correlation/decorrelation between safety and capability …
https://www.lesswrong.com/posts/yaz8nx4ogZmiqHzt7?commentId=AHo7xLgJHmJLCrm5b
and the measurements are relatively cheap, particularly in the context of frontier model pre-release testing.
I haven’t been very much engaged with this community in a while. I appreciate what Richard is doing here. My 2c for now is that what the EA/rationalist/alignment folks need to, and struggle to, fully digest and abandon is eugenics (and its milder cousin meritocracy). Because this “community” (used loosely) typically comes from a starting point that holds entrenched strongly favourable views and beliefs about eugenics etc. (even though I wouldn’t think they endorse them explicitly in those terms for the most part) they are taking a very long time to take in and/or rediscover the criticism and ultimate demolition of that whole belief and value system. I think it’s highly related to the more obvious (and described in this sequence) propensity to align with power, fail to critically evaluate power structures and systems, or recognise concentration of power itself as the main cause of harm.
Am I correct that you’re at Google DeepMind? And you equate meritocracy with eugenics? Is this a common view there? How are you even defining meritocracy?
Ramana left GDM three years ago.
Plenty of Googlers would tell you that Google has fully abandoned meritocracy :)