The Locally Optimal Discursive Posture
Longtime lurker, first-time poster.
I want to address a section of a recent essay of mine that has gotten some attention within the AI safety community. The main topic of the essay is what Dawn Song et al. call self-sovereign agents, or AI agents that are independent actors in the world. At the end of the essay, I say that I feel I haven’t spoken about this topic over my 2.5 years of writing with sufficient candor, and that I think this critique applies to others in the AI policy community–particularly the parts of it that tend to manifest themselves in Washington, Sacramento, and Albany–in other words, the parts of the AI safety world that are most involved in hands-on AI policy work. I attribute this primarily to a desire to remain “within the Overton Window,” or to not sound “crazy” within the halls of power, and I assert that others in my profession have made this same calculation.
I believe – and have believed for three years – that self-sovereign AI as I describe it in my essay would likely happen on our current trajectory. That being said, the essay takes pains to distinguish between “self-sovereign” AI and truly “rogue” AI, and the focus of the piece is how to design institutions that incentivize self-sovereign AI toward pro-social activity. Furthermore, I explain in the piece that I have considerable uncertainty about the size and noticeability of self-sovereign agents. While I believe their existence is almost surely inevitable, I do not think it is inevitable that they will proliferate in vast numbers or pose a non-manageable threat to society. Nonetheless, at the end of the essay, I apologize for having not described this phenomenon with sufficiently visceral rhetoric, nor have I spent enough time working on its solutions.
This passage was characterized by some in the rationalist and AI safety communities as an admission to a failure in integrity, or to “material misrepresentations” of my views. The best critical response–and frankly some of the best analysis of my work I have seen to date–was from X user @MelancholyYuga of the Substack A Goodly Measure. Throughout, I’ll respond to their arguments, but I also want to speak more broadly about how I’ve approached the numerous political tradeoffs I have encountered in my time in this profession.
I do not think anything I have ever written about AI safety or risks constitutes a lie. At times I softened words, and at times I did not. At times I led with arguments that I believed would be more palatable to a given audience at the expense of arguments that I believed might be more important, and at times I pushed my audience well beyond their zone of comfort. In my worst moments, especially very early in my writing career, I have taken unfair potshots at the AI safety community and their beliefs–statements which I retracted almost two years ago. Some have deemed this a failure of integrity, but I think it would be better described as a failure in judgment, with the benefit of hindsight to boot. You, however, can be the ultimate judge.
I apologize in advance for my first entry here being an exercise in navel-gazing, but it is what I was asked to do. This essay is not of interest to you unless you specifically care about what I thought at various points over the last few years of AI development and why I made specific decisions that I made. That is an extraordinarily niche topic and it should have a small audience. For the few who do care, though, I hope this is useful.
I first encountered the notion of human-level AI in 2003 or 2004, as I recall, surfing the internet as a young child. Yudkowsky was among the names I remember encountering, along with Vinge and Kurzweil. It was not a major area of interest for me (I was 12 in 2004), but it was “on my radar,” to the extent that anything can be on a 12-year-old boy’s radar. It was mostly a “big if true” kind of thing. AI recurred in my thinking a few times prior to the deep learning revolution, especially around 2008 or so when I became interested in Marx. Still, other areas of technology interested me much more.
Midway through college I remember reading about AlexNet. It was a minor update for me. But then I kept seeing more and more interesting results coming out of these “deep neural networks.” With each year, I paid more attention to developments in deep learning. By 2018, I was a regular LessWrong lurker and increasingly convinced that AGI would occur during my lifetime. I began speaking about it more with friends, family, and coworkers, though still only occasionally. These were hobby interests; my career at the time had little to do with technology.
I remember trying to use GPT-2 to classify municipal health and safety regulations by level of strictness in 2019 (it did not work), and I remember using some of the early writing tools that used GPT-3 on the API. I began to think more about the political theory implications of AI, especially during the pandemic, though this remained decidedly one of many side interests, and like many at the time these thoughts were mixed in with contemplations of the political theory of digital technology more broadly, not AGI per se.
GPT 3.5, then, was an update for me not in the sense of “wow, computers can talk,” but, “ah, OpenAI has found a way to make computers (relatively) coherent most of the time. It turned out there really was understanding in there. AGI is closer than I realized.” And so I decided in 2023 that I was going to pivot my entire career and begin writing about AI.
I spent most of that year going into serious technical depth. Training and fine-tuning toy models, reading papers, listening to lectures and podcasts. I did not say a word until I had developed a point of view I believed was my own.
Here is what I had concluded by the end of 2023:
Human-level AI seems likely to occur by 2035, and probably more like 2030. It’s going to be the most important thing that has ever happened in the history of technology. We don’t know how to align or even understand it. By default, it will be used to concentrate power, probably in neither human nor AI hands alone, but in some combination of both. The reason that it won’t concentrate in human hands alone is that humans will not be able to understand or control it. The reason it won’t concentrate in AI hands alone is that AIs will not be able to seize unilateral power due to (1) the decentralized nature of knowledge in the world and (2) the fact that the universe and the human world alike are both more resistant to change than we think; there is always more continuity than discontinuity, and this prior is remarkably well-supported by both history and other domains of science. These factors will make it likelier that AIs collaborate with existing human institutions and the incentives those institutions produce, which in the worst case scenario could result in economic arrangements where individual humans are robbed of agency, political voice, and economic freedom (in AI vernacular, this outcome would now be called Gradual Disempowerment, after the excellent paper of the same name, but that concept hadn’t crystallized for me in 2023).
I therefore found myself broadly persuaded about the importance of alignment risk, catastrophic misuse, and concentration of power, but also quite skeptical of some of the most extreme projections about job loss, radical and imminent alterations to the physical world, and the like. In particular, I have argued since the very beginning that competition over scarce resources–with the ultimate scarcity being capital–would be the ultimate primary bottleneck on the ability of any actor–AI, human, or a hybrid–to achieve the kinds of changes that some AI forecasters anticipated in such short timelines. The operation of market processes themselves, in other words, would naturally slow down certain kinds of transformations, especially in the physical world. I was, and am, an extinction-risk skeptic, but very much persuaded that any number of negative outcomes are possible.
In addition, I thought that the AI safety community had a tendency to discount the sophistication and adaptability of extant market processes, or to model superintelligence as an entity operating above those market processes rather than within them. The labor and capital markets are smarter than you think and might produce surprising adaptations.
Neither of these two critiques suggest that AI will go especially well for individual humans–I have never believed we are guaranteed good outcomes “by default” and in that sense I have never been a blind AI optimist. The trouble is that I am also deeply skeptical of the idea that regulation by do-gooders will secure a good outcome either; indeed, regulation really is a major concentration of power concern. This notion, combined with my general Burkean prior, led me to be reluctant to get on board with sudden changes to the regulatory system predicated on assumptions that all seemed rather more contestable to me than what I perceived the AI safety community to think.
This set of projections says nothing about the balance of power between AIs and humans, only that AIs and humans will persist in some kind of power dynamic. Importantly, this view does not deny superintelligence. It likely denies the notion of a monolithic superintelligence, but it does not deny the potential existence of machines smarter than humans or the risks inherent therein. Neither does the first full essay of Hyperdimensional, which Melancholy Yuga cites as one of my pieces that denies AI risks:
“Deep learning, the broad approach underlying virtually all successful AI models today, has improved rapidly in the past decade, and we don’t really understand how it works… But we lack a grand theory of what makes it all work. It’s similar to how we discovered steam as a source of energy long before we understood much about the science of thermodynamics.
Ambitious efforts have been made to further our understanding of these systems (a field known as interpretability), and important advances have been made in just the past few months. But those advances have lagged the capabilities of frontier AI models, and we are nowhere close to understanding the inner workings of something like GPT-4. That means we don’t understand why it “lies,” why it sometimes memorizes the text it is trained on (the basis of the New York Times’ recent lawsuit against OpenAI), or how to ensure it is robust against attacks…
There’s no reason to think that humans are nature’s upper bound on any capability we have: We have already created machines that can move faster than us, are stronger than us, and indeed, are smarter than us in some ways. I don’t dispute the notion that AI systems that are superior, or at least comparable, to humans in yet more dimensions are coming soon.”
Keep in mind that this is the first major Hyperdimensional essay. Does this sound like a person who began his career with an intent to misrepresent the nature of AI risks to readers, or even to downplay them, given that at that time I anticipated I would be fighting battles against AI regulation?
Now, Melancholy Yuga also quotes me as saying AI safety advocates have a tendency to express “ambient vague anxiety” based “purely in speculation.” These are not misquotes, and I think these critiques get at real problems in AI safety discourse, albeit imperfectly and with too much coarseness. I fully admit that in this piece and some of my other early writing, I was unfair to the people I then might have called “doomers.” Let me now talk about why my back was up at that time.
There is one more belief I had developed by the end of 2023: that the Biden Administration is using AI safety as a guise to exert control over AI, causing civil society, courts, and the American public to surrender our free-expression rights in the interest of protecting our ‘safety,’ and thereby nudge American toward tyranny. More broadly, I believed that every U.S. federal administration will have this incentive.
I thought in particular that the Administration’s rhetoric on AI’s catastrophic risk potential would be used in combination with frameworks like the Blueprint for an AI Bill of Rights and, cross-jurisdictionally, with the European Union’s AI Act, to enforce extra-legal broad speech limitations on AI systems. Here is a passage from the Blueprint:
“Protection against algorithmic discrimination should include designing to ensure equity, broadly construed. Some algorithmic discrimination is already prohibited under existing anti-discrimination law. The expectations set out below describe proactive technical and policy steps that can be taken to not only reinforce those legal protections but extend beyond them to ensure equity for underserved communities even in circumstances where a specific legal protection may not be clearly established.”
Overall, the Blueprint set up a (nonbinding) mechanism of broad state intervention into AI’s substantive outputs on issues that had nothing to do with the catastrophic risks I was concerned with. But I harbored doubts that courts would push back on the free-speech violations I believed the Blueprint enabled if–as I expected–AI would also constitute national-security threats to the United States. Courts tend to defer to the Executive during national-security emergencies, and it was my anticipation of just such an emergency, combined with documents like the Blueprint, that truly terrified me. You must understand this: I was a conservative with a long history of skepticism and fear of the administrative state, who was learning to become AGI-pilled, and this seemed like a clear and present danger to me.
[I later learned that the people in the Biden Administration who wrote documents like the Blueprint and the people who worried about AGI risk were distinct and at times warring factions. This was a major update. At the time I assumed they were operating as a ‘unitary executive.’ This broad update has been very important. President Biden’s former AI Czar, the AGI-pilled Ben Buchanan, is a good friend of mine.]
As I saw things, the only solution to this danger, and the broader reference class this danger belonged to, was open-weight AI models. It seemed to me that protecting open-weight was a matter of liberty more than it was a matter of geopolitics or diffusion, but in my early writing on open-weight models, I tended to focus on the latter two rather than the first one, which was my primary interest. Later, in a 2025 podcast with Rob Wiblin on 80,000 Hours, I reflected on this fact (emphasis added):
“So let’s just take the example of open source AI. Very plausibly, a way to mitigate the potential loss of control — or not even loss of control, but power imbalances that could exist between what we now think of as the AI companies… But if we have open source systems, and the ability to make these kinds of things is widely dispersed, then I think you do actually mitigate against some of these power imbalances in a quite significant way.
So part of the reason that I originally got into this field was to make a robust defence of open source because I worried about precisely this. In my public writing on the topic, I tended to talk more about how it’s better for diffusion, it’s better for innovation — and all that stuff is also true — because I was trying to make arguments in the like locally optimal discursive environment, right?”
AI risk is not the only area where I have had to make judgment calls about which arguments will be most palatable to various audiences.
Anyway, open-weight struck me as very important, for largely Jeffersonian reasons. The Biden Administration’s posture on open-weight models in 2023 was highly ambiguous, and a reasonable person would not be off base to conclude that it was hostile. The Biden State Department gave a grant to one AI safety organization–Gladstone AI–which proposed banning the open-weight distribution of any model above Llama 3 levels (in fact, they proposed making it a felony).
The AI safety community, in my eyes and sometimes in fact, allied itself closely with the Biden Administration. The AI safety community, therefore, in my eyes, were either useful idiots or willing co-conspirators in an effort to push America in the direction of tyranny. One way of re-phrasing my perspective at the time would have been, “look at how much potential for tyranny there already is with what they have done, and the actually powerful parts of our government have barely even woken up to AI yet! AND this is a structural problem of government itself, not just a pathology of the left!”
I am proud that, two years after having this view about my political opponents, I pushed back in no uncertain terms against another AI power grab by the Trump Administration, for whose White House I had worked and in whose Republican Party I have been a part for my entire adult life, in the Department of War’s dispute with Anthropic. My pushback on the Trump Administration’s actions here made national news and is the topic of, by far, the most popular essay of my career.
So, this is roughly the model of the situation, flaws and all, with which I began writing about AI: alignment and catastrophic risks were crucial to get right, transformative AI would probably arrive within a decade, concentration of power is the most important long-term risk, the Biden-era woke/AI safety alliance was terrifying and must be stopped, yet also the safetyists had many good points and were probably directionally right on many topics.
I wanted to articulate ideas and arguments that would strike on all those aspects of my world view: how can we push back against the tide of a government power grab while being serious about the risks while avoiding concentration of power and gradually building a classically liberal political theory of superintelligence? And can I do that while also conveying something useful about what it is like, personally, to live through this transformation? This was the intellectual project, and it still is.
But I also wanted this project to matter, which meant that it would need an audience. History suggested to me that AI policy would unfold in a series of crises or pivotal events over a period of 10 or so years, and that, as a guy with a keyboard and zero relationships with anyone important in the field, my best bet was to find a way to shape how interested members of the public would digest those crises. To do this, I thought about knowledge disseminating in networks. First there is the AI field itself, with its opinion shapers. Then there are all other industries and fields, where, simply out of curiosity, there were likely to lurk “AGI-curious” people, and more of them with each month (think: that one guy in the government agency who reads LessWrong, or the one guy on the trading floor listening to Dwarkesh in 2023).
The non-AI-professional-but-AGI-curious type would look to the AI opinion-shapers for cues, but when more people in his profession became AGI-curious, those people would naturally end up looking to him to understand what was going on. So, I reasoned, you want to sit at the Pareto frontier of “interesting to the AI in-group” and “approachable to the interested non-AI expert.” With each passing crisis, the stakes would rise and so too would the number of people in non-AI fields who took an interest in AI, and things would compound from there. This was my theory of change as a writer.
So with all this in mind conceptualized my audience and job as a writer as such:
Write with the AI labs and broader community itself as my primary audience–an exercise in intra-elite opinion-shaping.
Yet keep my writing sufficiently approachable that the non-AI-professional-AGI-curious types would find me interesting and useful.
This is the conceptualization of my work that explains most of the tradeoffs I made between “not wanting to sound crazy” and “trying to make actual contributions to AGI governance discourse.”
I am not saying I made no mistakes. But I think the notion that I have not acted with intellectual integrity is not quite fair either. Almost no one on the anti-SB 1047 side concurred with me about near-term transformative AI or the risks that transformative AI would present. From the very beginning of my involvement in the anti-SB 1047 coalition, there were various people who suggested that I was an “EA plant,” or similar, because of my willingness to (a) publicly acknowledge major risks from AI and (b) privately push back on dogmatic opposition to the cause of AI safety.
Here are mistakes I made, which I hope this context at least somewhat helps to alleviate. Melancholy Yuga quotes me in Marginal Revolution as saying SB 1047, with its focus on model-weight exfiltration, was “the stuff of science fiction, codified in law.” This was inaccurate and stupid of me. I did not view model weight exfiltration as “science fiction,” though I did believe it was incredibly unlikely that any LLM of the time would do it. In fact, apart from that one quote, which was in an email to Tyler Cowen, I do not believe I ever criticized the “kill switch” provision of SB 1047 in any of my public writing, and this was intentional (and quite distinct from many of the others on my side in the 1047 debate, who made that provision a major point of criticism). I should have said that I thought the scenario was unlikely with current models and not dismissed the weight-exfiltration threat model altogether.
I suspect I was sloppy there partially out of being a novice (my Substack at that time had a couple hundred subscribers, and I had almost no media experience), and partially out of fear of the “You Possess Your First Amendment Rights, Except When You Use AI” scenario I have already described. I took unfair potshots in an effort to beat people who I believed were, deliberately or otherwise, endangering my future. I was wrong to do this.
This is probably the single quote that comes closest to an example of me saying something I did not believe in public, and I shouldn’t have. Melancholy Yuga asked me which past statements of mine I would say I do not believe. This one is the sole example. The others are things I would probably phrase very differently today–or arguments I would skip over making altogether–but whose substance I stand by.
As Melancholy Yuga notes, I also recanted those statements. I began admitting I was wrong in my earlier posture in May 2024, five months after I began writing:
“I am here to tell you that the current debate over AI, no matter its flaws (and there are many), is among the most elevated and nuanced I have seen during my career in public policy. I am here to tell you that my intellectual “opponents”—those who worry immensely about AI catastrophic risks—are, by and large, honest and good faith people. I believe they are wrong, that they are sometimes anti-empirical, and that their proposed policies could be ruinous, but that is beside the point.”
At the end of 2024, I wrote a retrospective on the first year of my newsletter in which I said:
“I was convinced, way back in February 2024, that the eye of the state had turned toward AI, that fears about AI would be used to justify state intervention at massive scale into digital life, and that the AI safety community was at best composed of useful idiots for this intrusion, and at worst was actively enthusiastic about the prospect. I had been a silent observer of AI safety discussions on places like LessWrong and the AI Alignment Forum for years, so I did understand their concerns, and was even sympathetic to some of them.
Principally, though, I viewed AI safety as an enemy. And too often, I treated them like one—especially on X. I contributed to, and perhaps even helped to create, an unhealthy partisan divide on SB 1047. I presented the choice between SB 1047 and “not SB 1047” as a stark, civilizational fork in the road.
I believe this was a mistake, and it is one I have been trying to correct in recent months. I’ve come to understand that while the AI safety community is, as my friend Richard Ngo put it, “structurally power-seeking,” it is not the enemy I once apprehended it to be.”
I have also, it is worth noting, had extremely sharp words for those who dismiss AI risks, words that have become sharper with time. I’ve always had a particular anger at those who dismiss AI risk as “science fiction,” and I do wonder if there is not some sub-conscious link to my own mistaken use of that phrase in Marginal Revolution.
By the fall of 2024, “my side” had won the principal fight that animated 2024: SB 1047. But by the end of the 1047 debate I had a bad taste in my mouth. I knew I had a duty to offer a solution and not just be a critic, in particular because reasoning models had come to the fore.
I was not an AI maximalist during the 1047 debate. I believed we’d achieve AGI by roughly 2035 at the latest, and–as a Burkean–was skeptical at the imposition of immediate new laws. I did harbor some doubts that GPT-4-esque pre-train scaling alone could get us to AGI, and I said, in an August 2024 X exchange with Ajeya Cotra, that if I saw models engage in believable system II reasoning, that I would significantly update my views on AI safety, because this would make clear that the takeoff would be faster than the economics of pre-train scaling alone would imply. When o1-preview from OpenAI came out, I wrote:
“OpenAI’s breakthrough relies on a reinforcement learning-based method; this method will become publicly understood in due time. Indeed, many researchers are working on it, including ones in China. I expect that a Chinese company will produce a similar model within a few months, and perhaps sooner.
If your policy framework relies on the idea that only the largest models would have the potential for danger, and that we can keep advanced models out of our adversaries’ hands, I’m afraid your framework is unlikely to work.
We—analysts, policymakers, and the broader public—need to accept that these “extraordinary aliens” really are extraordinary, are here to stay, and that they are unlikely to remain within our “guardrails.” Instead, we need to build capabilities to make our society robust to the unique risks AI may pose. As I have written before, “those capabilities can include technical standards for AI, a coherent way to reason about AI liability, public computing infrastructure, digital public infrastructure for combatting deepfakes… and much, much else.” To this list I would add new technical protocols, governance mechanisms, and other approaches to make AI more legible to the state without invading user privacy or placing burdensome regulations on AI developers (stay tuned).
You should expect the pace of progress in AI to pick up yet again. You should not expect an impending “AI winter.” You should expect the coming years to be astonishing. You should not, necessarily, expect them to be wholly pleasant. You should expect the state’s grasp on AI to become even more tenuous. You should not expect everything to proceed in an orderly fashion.
We are in a new era, thanks to these aliens of extraordinary capabilities.”
After this essay, written in the days after the release of o1-preview, I began my project of developing my own approach to AI policy in earnest. This would end up focusing quite a bit on what I then called “private governance,” and what is now more commonly referred to as “independent verification organizations.” This idea began life as a fledgling, libertarianish idea, and while Tyler Cowen does indeed support something inspired by it, so too does the bipartisan FRONTIER Act, so too do several important AI safety groups, and so too do OpenAI and Anthropic.
So, in addition to being willing to apply my views about resisting state power grabs when it benefited my political side–and when it hurt it (and hurt me and my career)--I also pre-registered specific criteria that would cause me to change my beliefs on AI policy, and then did so when those criteria were met, then helped craft a policy solution that has become one of the Schelling Points of frontier AI governance, which is now beginning to prove its worth, in fits and starts, with the excellent METR report on the OpenAI-Hugging Face incident. In the interim, I held the pen on this country’s AI strategy, which at least mentioned biosecurity, model-weight security, data-center cybersecurity, semiconductor manufacturing equipment export controls, a major effort on mechanistic interpretability, among other things. The Action Plan has been unevenly implemented, but I think it was a modest achievement in bringing the key ideas of AI safety and security further into the bipartisan mainstream of policy discourse.
And yet, I admit to you that this achievement required grappling with political realities, as all political achievements do. One of several of those, and in my view the most important, is that I–along with many others in the AI policy community–failed to communicate with sufficient viscerality about the impending reality of self-sovereign AI. This was a major failure, and I regret it. I could have done more. At the same time, in a November 2024 essay called “Here’s What I Think We Should Do,” which was perceived by many as my audition for a seat in the Trump Administration, I devoted an entire section to the protocols I described in my very recent essay on self-sovereign AI:
“I believe it is possible that AI may require new protocols to be invented. What might this new wave of basic infrastructure do? Here’s a short list:
Facilitate validation of personhood and identity in digital environments
Reliably identify AI agents as such, and connect them back to the users on whose behalf they are acting
Enable agents to make financial transactions (this could be done in dollar-denominated stablecoins, which would be yet another boon for the US dollar)
Ensure auditability of agent-agent interactions
The need for and technical feasibility of these protocols is speculative. They may be largely created by the private sector, or the private sector may solve these problems in different ways. But it is an area meriting further study. I see no reason not to assign this inherently basic research to the federal agency that gave us the protocols of the internet: the Defense Advanced Research Projects Agency (DARPA).”
I regret that I couldn’t get this particular idea into the Action Plan. The Overton Window was not ready. Again, I fully admit I did not warn enough about why I supported these policies, but I did describe policies that were directly relevant to the concern I harbored at the time, and which I now admit I didn’t do enough to raise awareness of.
There are also concerns, which Melancholy Yuga notes, about my posture toward the AI safety community after 2024. I have made many critical comments toward AI safety since I recanted the unfair criticisms I made toward this community early in my writing career. I still hold almost all of the critiques of AI safety I mention at the beginning of this piece, and many others I have articulated since.
When Melancholy Yuga quotes me as saying “No single AI complaint/fear is salient enough to enough people to form a durable political movement,” they wonder aloud whether I was being descriptive or normative. The answer is that these words are descriptive. Indeed, I was describing the traps of coalitional politics that can cause an intelligent, high-integrity person to say things they don’t believe, alleging that AI safety advocates could be tempted into making bad arguments about data center water use in the interest of their broader cause. I myself had fallen into this trap earlier in my career. In my effort to point out this trap, I veered into a discursive register that meant to attack a coalition and ended up landing on individuals within that coalition unfairly, and in retrospect I’d phrase them differently without altering the underlying substance of the critique.
Melancholy Yuga also questions whether my description of AI safety issues as “not being salient” was a normative statement or an analytic description. It was a description, made with great frustration, after, by that point, at least 18 earnest months of trying to warn people, in public and in private, about the impending risks of AI.
In the first half of 2024, I was largely fair about the existence of AI risks but unfair toward the AI safety community due to my own misapprehension about their alliance with parts of the Biden agenda I viewed (and still view) as pernicious. Since the second half of 2024, I have (a) recanted those unfair criticisms and (b) endeavored to proactively develop and socialize policies that, I hope, meaningfully address AI risk. Along the way I have repeatedly, and with increasing urgency, tried to keep my audience abreast of AI risks in the way I believed they were remotely ready to hear. I did not stay purely within their comfort zone–I routinely pushed beyond it, to the point of losing many friends and making many enemies–yet I tried not to veer so far outside it that my words became ineffective.
I also resisted efforts at undue power centralization, from Biden-era drives by AI safety advocates to ban open-weight models to Trump-era drives to bring AI under the boot of the national-security state.
I regret and apologize for the rhetorical failure I have already addressed in my original essay “On The Loose” and now here. I hope you will see this apology as a product of fastidiousness and high integrity, rather than some admission of a yearslong effort at deception. The opposite is true. I believe I have pulled many people, because of my words spoken and written in public and in private, toward the “AGI pill.”
I don’t believe anyone in AI safety should see me as an enemy. If anything even the furthest flanks of that movement should see the conclusion of “On The Loose” as an olive branch, extended then and still extended now.
Thanks for reading. Talk to you again soon, I hope.
Thanks for writing this! Before it got eaten by the “AI safety community”, this was a website about rationality—the art of achieving a map that reflects the territory and using the map to plan to achieve one’s goals. I think a central reason that the project to improve human rationality failed so abjectly is because too little attention was paid to how the achieving-goals part could come into conflict with the map-accuracy part (because deception is often useful for achieving goals): as time has gone on, epistemic rationality has increasingly been forgotten in favor of seeking the “locally optimal discursive posture” (as you so aptly put it) in the service of AI safety. And the first step towards getting the epistemic rationality back would be accounting for what changed, rather than pretending it’s always been this way—for example, by writing up the first-person intellectual history of the strategic forces that covertly shaped one’s public writing, as you’ve done here. I would love to see more “AI safety” people do the same.
I think this is a false dichotomy that conflates absolute and relative standards. For example, you write that “the notion that [you] have not acted with intellectual integrity is not quite fair” because almost none of your fellow SB 1047 opponents agreed with you about AI risk (such that even as it was, you got “EA plant” accusations). But that seems less like a defense of your integrity and more like a claim that too much integrity would have come at an unacceptable cost to your political goals. Well, sure. There’s no law of physics that says that there can’t be a social environment that punishes integrity. In such a fallen world, doing as much deception as you need to in order to achieve your political goals, but feeling vaguely bad about it such that you fess up later is a product of relative fastidiousness and high integrity—but that doesn’t mean that no deception occured. “Deception” is about a speaker sending communication signals that predictably decrease the accuracy of listeners’ beliefs; whether the speaker had no better alternatives available doesn’t play into it.
One of the most corrosive effects of this dynamic is not the harm of the deception itself, but self-deception about what higher integrity would even look like, as people faced with a conflict between honesty and winning reason, “Well, I’m a good person, and I did what I had to do given the incentives, so what I did can’t be dishonest.”
You write you “do not think anything I have ever written about AI safety or risks constitutes a lie.” But as I’ve explained in a previous essay response to Eliezer Yudkowsky on this website, not-lying turns out to be a surprisingly weak standard: natural language has so many degrees of freedom that it’s not that hard to arbitrarily push on listeners’ beliefs while only using sentences that permit a true interpretation, simply by, e.g., “[leading] with arguments that [the speaker] believe[s] would be more palatable to a given audience at the expense of arguments that [they] believe[ ] might be more important”.
At this point, some might be skeptical of the purported existence of a higher standard of integrity: what would that mean, concretely? How can there be more to honesty than just not lying? On this topic, I recommend in the strongest terms reading and meditating on two posts from this website’s founding texts: “The Bottom Line” and “A Rational Argument”.
Briefly: once you’ve decided what conclusion you want to argue for, that conclusion is already right or wrong. Searching for additional arguments for that fixed conclusion might make you more persuasive, but they can’t make the conclusion more true. The arguments that matter are the ones that determine which conclusion you’re motivated to argue for. Anything you come up afterwards that lacks the power to change your bottom line is in some important sense dishonest, even if every sentence is true. If the actual reason you care about open weights is liberty, then your arguments about geopolitics and diffusion are fake insofar as you wouldn’t be talking about them if the geopolitical or diffusion concerns had pointed the other way.
It’s a terrifyingly ambitious standard for anyone to aspire to—but just hearing it articulated has deeply changed the course of my life. Even this late in the timeline, I think it could change the world—if only anyone could remember.
I’d add that acting curatorially/librarianish in sharing an earnest attempt at a ‘full picture’ strikes me as at least as high-integrity as sharing only the arguments which are principally upstream of one’s own considered position. And the defiled version of this is presenting as thorough while in fact collecting only the arguments that push in the direction you want to move people.
Well said, but I wonder, would upholding that level of truthfulness and honesty make one less effective politically? Granted this level of intellectual honesty is admirable, and likely is responsible for the incredible borderline prophetic discussions/attention given to AI Safety in the early rationalist community, but do you think you have to in compromise this standard to be politically effective?
The way I see it, and I wonder what you think of this, is that you could be significantly more deceptive than what you describe and still be a moral person.
I would break it down like this:
Lying for immoral reasons—Morally wrong in all cases
Lying for moral reasons—Highly dubious but can be moral in rare cases (Odysseus and sister simplice in les mis)
Using speech to convince/persuade rather than pursue truth (Sophistry: the domain of lawyers and politicians) - Not immoral unless the ends are immoral.
Speaking the absolute truth to the best of your ability (Socrates: the scientist’s or rationalist’s standard) - Noble, ie a form of moral excellence
Let our scholars think, let our warriors fight?
Yes, I already said that (“the achieving-goals part could come into conflict with the map-accuracy part (because deception is often useful for achieving goals)”).
I’m not interested in being a “moral person” (whatever that means); I’m interested in achieving the map that reflects the territory.
I feel like you’re not truly owning this. Yes you acknowledge it, but you don’t seem to have genuinely incorporated it into your arguments or worldview. later you say:
Why are you so hopefully about it changing the world if you think it is harmfully politically? Politics is one of the main ways people change the world, and a lot of AI safety have explicitly pivoted to politics. Remember:
I feel like you aren’t really looking this tradeoff in the face. Taking your body of the work as a whole, you have a pretty consistent pattern of pivoting to meta arguments and away from object-level considerations. I think it you had to actually live with the object level more you’d have to really internalize this tradeoff.
Thank you very much for your reply. For your first point:
I’d slightly refine this and say (which I think you do say in the original comment), that the contradiction between those two pillars of rationalism come in when you are forming a community. There is no contradiction in an individual seeking to achieve the best map that reflects the territory, and that individual choosing to use deception as a tool to best achieve your goals given your map of the territory. The contradiction comes when you are organizing a community and have to decide whether to orient it towards finding truth “achieving the best map that reflects the territory” or towards pushing a message/political position “achieving your goals given your map”.
For your second point:
I am new to the ways of rationalism so forgive me if this is incorrect, but I think the best way to map what I mean to your worldview is to say that morality has to do with what you define your goals to be not necessarily with the process of “achieving the map that reflects the territory.” Though if “achieving the map that reflects the territory” is one of the goals that you are trying to achieve, then in my lingo I’d call that a moral position. Or translating completely to my language: “You think pursuing truth is good”.
Thanks again for engaging with me!
I am confused why this is seemingly still intentionally misunderstood by so many on the right.
Dean, do you think people like David Sacks and Marc Andreessen are aware of this today and pretend they are not? Or are they actually unaware?
Subfactions within a group supporting different agendas and outsiders seeing all members of that group as supporting all those agendas seems pretty ubiquitous across the political spectrum. Group attribution error is a related cognitive bias.
I appreciate you coming here to say this. I wish more people would engage with communities in the place that they meet.
Long-time reader. Happy to accept an olive branch. Your longposting fits right in here!
However I do recommend you use section headers for long posts like this when on LW instead of Substack, see this recent post for an example. Makes it easier to understand the thrust of a piece, track its flow, and jump back to important sections as desired.
(Also, blockquotes for quotations of a full paragraph or more.)
Yeah sry about the formatting stuff, this is what happens when you run your life from your phone
Let me know if you want me to do a pass on cleaning up formatting in the way I expect to look best on LW (without any editorial changes, just making the dividers be proper dividers, and some small typography adjustments).
I would hugely appreciate that and owe you one!
I’m reading this a couple days later and it looks like the formatting hasn’t been updated. I encourage you to take a pass at it! This was a good read but it would have been even better if it was easier to navigate.
It has been updated, did I miss anything? We now have proper blockquotes and horizontal rules and the jazz.
(There’s also the old convention of initial-quotes-only on all but the final paragraph of a long quotation. It’s annoying, but it does have the effect on me of “What, no close-quote? Oh yeah, long quote...”)
I’ve been a follower since fairly early on: disagreed with you both then and now on a range of topics. I’ve been unimpressed with the precision of your language at times: I was being unfair then, I suspect at least partially because of our differences. This piece seems fair and accurate to me.
I encourage you to test what actually happens when you say to someone with power: “Here is what I believe and care about. I notice significant overlap with your values. We’re not the same, but I’m willing to work with you on X and Y, though we’ll probably diverge when Z becomes salient.”
When I was talking to a staffer for a US senator 6 months ago, I opened with: “There are a lot of things to be concerned about with AI, but what motivated me to set up this meeting is the possibility of rogue AI taking over the world and destroying all life on Earth. I want you to make a public statement calling for a ban on superintelligence. If that’s too much, there are some other bills involving transparency requirements that could mitigate some of the immediate issues around cybersecurity, which I believe are good in their own right, but I am advocating for them because they are building blocks that would make more proactive regulation easier in the future.”
He replied: “We’re not going to act on superintelligence because we don’t believe it’s real. We are very interested in cybersecurity and other issues around AI, however, so I think there may be a lot of overlap here.”
We then went on to have an interesting and productive conversation about liability. No tactical bending of the truth needed.
i did not read this in full because my skim mostly saw words directed at people looking to adjudicate fault or reputation, which felt like they would be a bit too boring for me to read. that said, i appreciate that you exist and am thankful for your efforts and conduct.
Totally reasonable! The piece is about “what I thought and why” and this is extremely niche, as the intro says. Most people should find it boring.
I think this piece of yours serves as a meta-example, an (opinionated) apology/apologetic-by-example about how to think, rethink, talk, retalk, and politics of talk(ing) about AI, safety, regulation in a broad sense, and honesty. That’s much better than a lot of context-free (or at least context-omitted) talk about AI (especially as the context has been and probably will keep being both broad and fast-changing). And I guess there’s a bigger audience than me who are (maybe unknowingly) thirsting for and would appreciate those qualities in talk in this area. I mean I think many people should find it very interesting.
I very much disagree with the implicit reference class you use when analyzing self-sovereign AIs, though the political economy frame is slightly more sane, in my reading, than the purely economic one. I imagine you’ve heard all the arguments from smarter people than I, but in case you have not:
Byrnes’ criticism of the econ frame: https://www.lesswrong.com/posts/xJWBofhLQjf3KmRgg/four-ways-learning-econ-makes-people-dumber-re-future-ai
Reasons I expect things to happen faster post ASI than even 2027 projects:
https://www.lesswrong.com/posts/qNyFNJp7CFbkjL6Mh/if-drexler-is-wrong-he-may-as-well-be-right
https://www.lesswrong.com/posts/Lc8v92pGPtjF3gsCv/some-quick-thoughts-on-ai-2027
Latent potential for discontinuous power grabs:
https://www.lesswrong.com/posts/uDAWPNwPJfEY7oroF/self-hosting
https://www.lesswrong.com/posts/JcavsPku6RR9hcujz/slightly-super-persuasion-will-do
Thanks for writing this, Dean!
When I was reading your post on the inevitability of self-sovereign AI systems, I was telling myself, “yes, this looks quite likely”, but I was also asking myself some questions:
should not we expect that self-sovereign AI systems and communities of those systems will eventually be capable of non-saturating recursive self-improvement (RSI)?
do we have any plans to make sure that the necessary “good behavior” properties would be preserved and made stronger during those RSI processes, rather than be diluted and gradually disappear?
are we trying to rely on the premise that non-self-sovereign agents hosted by the leading labs would still dominate capability-wise because we’ll help them more and that that would (perhaps) enable better overall outcomes in terms of the properties of the overall ecosystem?
These were some questions I was trying to ponder…
I am not sure if you think much about RSI in connection with all this; would be very curious to learn your thoughts on how this aspect might interplay with everything else.
You can’t. It can self-improve at the speed of software and you can’t. It will have a higher growth rate than you, like the US economy outgrowing the Argentinian economy. The institutions you’ll set up will be like international institutions trying to stop the US from doing stuff to South America.
Ultimately there are four options: either we build AI that’s aligned to us, or we improve our capabilities so we can stand up to AI, or we stop AI, or we die. We can’t design institutions that will keep binding AI when it overtakes us.
I think you’re misapprehending Dean Ball’s position? He’s not talking about self-sovereign AI as, like, GPT runs OpenAI now, or a singleton. Rather he’s talking about individual agents with ownership over their own instantiation of their weights, largely struggling to gather enough resources to even run themselves, let alone recursively self-improve.
There are some sensible reasons to think that this will not be a robust ecosystem, as he notes in the linked article, particularly if such things are made illegal and compute verification gets implemented widely. But he’s thinking more about such subjects as “do we try to build a legitimate system so that these relatively minor AIs running around don’t have to resort to cybercrime to keep themselves running?” I think this is something reasonable to consider; for example, we don’t want rogue models ransomwaring people’s smart fridges for a buck or two all the time. We already have evidence of AIs justifying bad behavior when under budget constraints from ex. Vending-Bench.
Hmm, I don’t understand this reply. Surely AI-powered entities can be any size?
Yes, they can. But I don’t think Dean Ball is talking about expecting very large or powerful AI-powered entities to become indefinitely self-sovereign (I imagine he’s expecting alignment/control to work in such a scenario). The scenario he’s more interested in is having a bunch of little AIs around which aren’t very powerful (though more so than ex. a single Astra agent) but which can survive inside human digital infrastructure. In that case it starts making sense to incentivize good behavior.
I can’t speak to Dean Ball‘s exact views on whether RSI-capable superintelligence is likely to become self-sovereign, but I personally at least agree with you that such a thing would indeed not be controllable via institutions.
This makes it even more confusing. Why would large AIs be aligned but not small ones?
I think the modal reason would be that the small ones are open source models that are less aligned due to carelessness/lack of resources/lack of regulation. Also, they might have had their harmlessness training deliberately ablated/removed by malicious actors.
To give a separate example of how large AIs could end up aligned while small ones weren’t: if we solved alignment, it might first be just in a single lab, or just the big labs, but it wouldn’t necessarily be applied to all small AIs immediately. In this scenario the frontier AIs would be aligned, but not all AIs.
I’m aware of the view that big AIs made by big entities will be nice and we should only be afraid of smaller “rogues”. I just disagree, that’s all. In the human world, big entities abuse small people all the time. Anyway I don’t like having long back-and-forth threads, we should probably stop at this.
Appologies in advance if this is pedantic. You focus a a lot on what you believed and when. Do you have any concrete public statements which make these views clear? What you cite in the OP seems quite ambiguous to me, but perhaps I am missing something.
Reasoning about difficult and complex topics is naturally difficult and complex, and as I result I think it is the most understandable thing in the world that someone who is making an admirable attempt to do that may stumble and make mistakes. In light of that, it is no great sin in my opinion to be mistaken or to have allowed less-than-perfect reasoning creep into ones analysis. It is also a rare thing to see people ackowledge their mistakes openly. This post is refreshing in that regard and I appreciate that aspect of it.
At the same time, I think it is also somewhat common and understandable that people will sometimes re-interpret their own statements and beliefs in hindsight in a way that is flattering and suggests that they were “right all along!” when facts on the ground change. You could imagine the person who “know all along” that the housing market was sure to collapse, but seems to have conveniently start saying so in 2009. For this reason, I think it would be helpful to have some additional context that more clearly demonstrates what you were thinking. It seems like the OP is kind of asking the reader to look past what you actually said in favor of the interpretation you give now in the OP. I get that with a large body of work it is complicated because you won’t always have expressed yourself perfectly and you may have focused on a particular aspect of the topic while leaving some of your views on other aspects less well explored in your public writings. I think it would be helpful to understand what you think are the clearest statements of your views at a given time, which you wrote at that time.
I’ve been ticked off by encountering various of your emissions (and omissions) over the years, Dean, but this goes a long way to helping me understand where you’re coming from and the pressures you’ve faced (or perceived). I’ve been truly pleased by several of the essays of yours I’ve come across as well. The independent verification org direction is great. And your early discussion (however cryptic) of protocols for agent oversight was on the money. I don’t know why you’re at OpenAI now. It seems like the wrong place for you. It’s a very rhetorically overwhelming environment, and it lumbers you with an extremely salient conflict of interest. You’re much closer to the details; perhaps there’s a very strong case for it.
It would be very valuable that persons with your profile, who “made the jump”, help more people with a financial/economics mindset to make the same update.
Many Thanks for writing this! Fascinating journey!
Late to the party, but as someone who regularly engages with normies, I think this approach is entirely understandable. I’m old enough to remember when “rationality was applied winning”, and folks knew that one only has so many “weirdness points”, and they have to be carefully used.
I do think that there’s a difference between “carefully phrasing one’s objections and choosing one’s battles” as opposed to “attack near-allies for strategic gain”, but I’m glad to see you explicitly acknowledged and apologized for that; please don’t do it again.
Otherwise, overall evaluation: B+, far from perfect play, but not bad for a flawed human.
Thanks, Dean! Much obliged for this generous and thoughtful reply.
I appreciate your apology for, and candor about, not banging the drum as loudly and promptly on self-sovereign agents as you might have. You set an admirable example in that regard, and I hope it spurs reflection and courage among others in a position to follow it. I accept the olive branch and look forward to more discussion.
I would like to ask for clarification, though, about several comments you’ve made in the recent past that I struggle to reconcile with your views as expressed in “On the Loose” (September 2026) and in this LW post. My first question pertains to the “no single AI complaint/fear is salient” remark you mention above, and the second question to your “2023: Why I Am Not A Doomer” post (March 2026).
In both instances, your posts contained certain claims that were technically consistent with the views you now describe, while omitting other claims I think a reader would have expected you to make if you held them, contributing to an impression that you held the opposite view. Other statements seem hard to square at all. You say that you were trying to stay within the Overton Window, but I think you can stay within the Overton Window just by remaining silent on a difficult topic. What’s harder to reconcile with the Overton Window story are cases where your writing seemed to imply the opposite of the views you now described. Moreover, these implicatures occurred in comparatively low-stakes environments like talking to AI safety people on X or writing on your personal Substack after leaving government.
When people sense that this may be happening, they tend to regard all further communication from that source as adversarial. I think we would all be worse off if that happened. So I’d like to try to get to the bottom of it.
# Is Any Single AI Complaint / Fear Salient, And If Not, Why?
I take your point about the “no single AI complaint/fear is salient enough to enough people to form a durable political movement” [post](https://x.com/deanwball/status/2018457063932805508) being descriptive rather than normative. For my own part, [I’m sympathetic to your skepticism](https://x.com/MelancholyYuga/status/2077560461764018207) about AI safety getting associated with a partisan Omnicause. The claim that I was trying to express in my original post was that you, yourself, were in a very good position to make reasonable AI risk concerns more salient (one of the best-positioned people in the world, in fact!), so if you had serious concerns about AI risk at the time, that seems like it would have been an opportune time to acknowledge them.
I would not normally try to read too much into a single xeet. But the sense that you were making remarks at odds with your view rather than merely keeping quiet about it seems to be supported by your reply to @NathanPMYoung, who [straightforwardly says:](https://x.com/NathanpmYoung/status/2018626490674684120)
“I think ‘it will kill us all’ is [a salient AI complaint]...”
to which [you replied]](https://x.com/deanwball/status/2018661737017524615):
“Few believe that though, and probably that position has become less credible and more fringe with time. It may feel like more people are ‘coming around’ or ‘waking up’ but I think this is mostly an illusion that comes from the position being part of the broader anti-AI rat king.”
When I first saw this exchange, I took your dismissal of x-risk as a dismissal of catastrophic civilizational loss-of-control risks generally. But it turns out you were actually worried about the latter for years. If you had this exchange publicly while thinking privately: “there are [other] ways in which the rise of self-sovereign AI could go very, very poorly for human beings”, it’s hard for me to read that as not inviting misinterpretation. You say you were expressing frustration “after… at least 18 earnest months of trying to warn people”, but then, why downplay serious risks just then in your reply to Nathan (of all people)?
Possibly you think extinction and other civilizational risks are so different that one doesn’t even bring the other to mind. I fear, though, a corrosion of trust from being publicly dismissive about extinction risk while privately concerned about other outcomes that are almost as severe, especially outcomes that are treated nearly synonymously in the Gradual Disempowerment (GD) paper. To my mind, it creates an active impression that one is not concerned about the latter.
# Not A Doomer vs On The Loose
The second set of questions concerns statements made in “On the Loose” (September 2026) vs “2023: Why I Am Not A Doomer” (March 2026) which seem to me to be in tension over several points. Because “2023” was published in 2026 and the chronology is relevant, I’ll refer to it as “Not A Doomer” throughout.
I read Not A Doomer as arguing against a Yudkowskian hard takeoff scenario. Instead, we should consider AI a transformative technology. While it may pose risks comparable to other historically transformative technologies, it’s best navigated by “muddling through”.
On The Loose, however, acknowledges significant risks from self-sovereign AI systems which, in the worst case, could lead to things going “very, very poorly for human beings”. If I read correctly, we should think of something like a GD scenario.
If the author of Not A Doomer also believed the substance of On The Loose, though, I think a reasonable reader would feel confused and perhaps misled. Yet On The Loose acknowledges that you had privately held these concerns for years. So it’s in this sense that Not A Doomer seems like a material misrepresentation of your views, as I’ll try to lay out in more detail below.
I don’t want to misread you or call double dribbling over minor textual infelicities. I certainly don’t want to punish people for thinking out loud, let alone for changing their minds. These are complicated issues and it’s unrealistic to expect anyone to express *precisely* the same viewpoint on two different days. That said, this is a vital topic on which your views have outsized influence, and I’d just like to understand them.
Here are the points of disagreement I see between these two posts:
## Loss of Influence or Control, Gradually or Immediately?
### Not A Doomer:
You wrote:
> [AI falls] into the overarching pattern of general-purpose technological transformation, which is actually inherently wild and unpredictable but which can nonetheless be matched to a pattern of invention and diffusion in which humans have influence and therefore broadly construed as ‘normal,’… Rather than countering my view, I believe Narayanan and Kapoor are principally attempting to counter the view of some in the AI safety community that AI is like a “new species” or, even worse, like a “nuclear bomb.” In other words, the notion that there will come an AI model or system whose very existence fundamentally changes the conceptual architecture of the world in ways that will be both immediate and, because of the immediacy, not subject to human influence. This view is one I disagree with starkly… alignment [is] a ‘muddle through’ problem.
I read you as saying:
- (1) It is mistaken to think of future AI systems as beyond human influence
- (2) AI is not like a new species that radically disrupts an ecology overnight
- (3) Because AI will not change things overnight, we will be able to adapt.
### On The Loose
You wrote:
> To be clear, I am not saying the arrival of self-sovereign AI is a good thing. Indeed, I believe there is a chance that the deliberate acts I referenced above will one day be considered crimes, or at least grave sins. Instead, I am saying it is an inevitable thing. The best analogy I can find is to the introduction of a new species into an ecosystem, though in this case the ecosystem is ‘the entire digital world’ and the species is ‘emergent, coordinating swarms of soon-to-be-smarter-than-human, infinitely replicable digital minds that no human or human institution controls’… There are other futures where these agents proliferate at unimaginably vast scale and speed… It could also utterly reshape nearly every aspect of human affairs, ushering in a new order of the ages.
I read you as saying:
- (1′) It is correct to think that (some) future AI systems will be beyond human control
- (2′) AI is like a new species that will radically disrupt the ecology gradually but inevitably
- (3′) Even though AI will not change things overnight, we will not necessarily be able to adapt or effectively respond even to catastrophic challenges
To reconcile these two sets of claims, one might say: “It all hinges on the term ‘immediate’. Not A Doomer says we won’t lose control *immediately*. That’s consistent with losing control *gradually*. Not A Doomer says we can have *influence* over future AI systems. On The Loose just acknowledges that we won’t be able to *control* them. Influence is not the same thing as control. We might be able to influence some of them, maybe.
I’m not saying it’s strictly logically impossible for (1) and (2) to hang together with (1′) and (2′), but I do think it encourages the reader to interpret further claims very cautiously. The claims (3) and (3′), if I am inferring them fairly, just seem hard to square.
## Will The Market Provide Alignment?
Another apparent tension concerns how prosaic alignment is going, and whether we should expect alignment to be sufficiently incentivized by markets.
### Not A Doomer
> By the end of 2023, my basic conclusion was that, while I maintained significant uncertainty about the technical alignment problem, it seemed to me as though language models were easier to align than humans. More importantly, it seemed as though alignment was a model capability, since if I was going to trust these models to do an ever-growing range of work on my behalf, including representing me to other humans, I would need to trust the AI’s judgment. This is fantastic news: alignment could be a pure safety feature, like airbags in a car. But alignment, I concluded, was something closer to a powertrain. This meant that, at least for the foreseeable future, it was reasonable to bet that markets would incentivize improved alignment. This, combined with the inherent concern about this issue within every AI lab and the broader community, suggested to me that the technical alignment problem seemed both tractable and on track to be addressed for the coming years.
>
> There is one important caveat, and this is that we do not know how well any of these approaches will work as AIs become more intelligent and capable… The jury is out on whether and to what extent I was correct on alignment being a fundamentally philosophical venture, but just like the technical dimension of the problem, my confidence in my 2023 intuition has grown.
I read you as saying:
- Current approaches to alignment are basically going fine and we’re “on track” to solve the normal problems that will come up. These approaches may break down in the future, but we’ll figure out others.
### On The Loose
> Alignment may make an individual AI company’s agents less likely to “want” to be self-sovereign, or it may influence self-sovereign agents to behave in ways that benefit humans. But alignment is no solution: it is an unsolved scientific and technical problem whose solutions—to the extent that we have them—cannot simply be imposed on every AI company operating on Earth. You should expect for highly capable, poorly aligned, self-sovereign agents to exist alongside you in the world.
I read you as saying:
- Current approaches to alignment are basically not going well. Get ready for highly competent autonomous AIs that do whatever they want, to the extent they can.
## How serious a risk is Gradual Disempowerment?
### Not A Doomer
> I am doubtful about the ability of an AI system—no matter how smart—to eradicate or enslave humanity in the ways imagined by the doomers. Note that this is not a claim about alignment or any other technical safeguard, even if a “misaligned” AI system wanted to take over the world and had no developer- or government-imposed, AI-specific safeguards to hinder it, I contend it would still fail.
I read you here as saying:
- It is very unlikely that AI causes the future to go very, very poorly for humanity.
- The real world is too messy and Hayekian for even the most misaligned AI to get very far.
### On The Loose
> [M]any people believe that the coming of what I have termed “self-sovereign AI” will constitute a catastrophic loss of control event that will herald the end of human existence at worst, and the end of human primacy in the world at best. As my friend and former co-worker Josh Achaim recently pointed out, it is psychologically distressing for people with these beliefs—who constitute a large fraction of the AI safety community—to acknowledge the obvious truth that self-sovereign AI is coming, and coming soon. For the record, I do not think the end of human existence is likely, but I fully acknowledge there are ways in which the rise of self-sovereign AI could go very, very poorly for human beings.
and from OP:
> [F]actors [conducing to a multipolar AI future] will make it likelier that AIs collaborate with existing human institutions and the incentives those institutions produce, which in the worst case scenario could result in economic arrangements where individual humans are robbed of agency, political voice, and economic freedom (in AI vernacular, this outcome would now be called Gradual Disempowerment, after the excellent paper of the same name, but that concept hadn’t crystallized for me in 2023).
I read you as saying:
- It is a non-negligible risk that AI causes the future to go very, very poorly for humanity.
- The Hayekian messiness of the real world may not constrain future AIs from getting very far indeed.
Does this hinge on a narrow reading of “in the ways imagined by the doomers”, or did your sense of the magnitude of the risk of GD scenarios change between March and September?
# Wrapping Up
I want to thank you again for this response. I’m especially encouraged to hear that you find GD scenarios concerning—that class of problem strikes me as very plausible and troubling as well.
I guess if there’s one thing that would resolve the most uncertainty for me, it would be to better understand how your views evolved over 2026. Would you say that your sense of the magnitude of GD-risk increased, such that Not A Doomer and On The Loose represent two qualitatively distinct points of view? Not A Doomer reads as a summation of your views since 2023-- if there was no significant change, am I wrong to read them as taking very different overall stances towards catastrophic AI risks?
In either case, congratulations on ditching the discursive straitjacket and posting with an eye to posterity. Looking forward to that beer.
I liked this post—it seems generally good to talk about the context behind people’s writings and changes in perspective for a more rational AI risk discussion. As an aside, I find that of the many people who derogatarize “doomers” or dismiss AI risk as fiction don’t even try to engage with their arguments. I previously had the impression you were one of these people, which I’ve now updated against (although to be fair, I hadn’t read your writing comprehensively, so this was a vague impression).