My blog is here. You can subscribe for new posts there.
My personal site is here.
My X/Twitter is here
You can contact me using this form.
My blog is here. You can subscribe for new posts there.
My personal site is here.
My X/Twitter is here
You can contact me using this form.
I mention corrigibility as an example of a property associated with the first bar!
My sense is Yudkowsky considers corrigibility to be deeply unnatural, especially after failed MIRI attempts to formalize it.
Democracy currently works because, even if high officials secretly wish to be tyrants, they can’t get away with it. The main two reasons they can’t get away with it are that they need labor buy-in to support the economy, and they need soldier buy-in to support the military. Laborers and soldiers believe in democratic rule, and if the would-be tyrant gets too blatantly undemocratic, the laborers and soldiers will rise up to remove them.
Exactly! I’ve written about this before, for example here (which you responded to) and here (in some more detail).
Most strategies to solve this don’t work. Even in the unlikely event that the economy goes so well that everybody owns their own open-source AI doing economic tasks in a perfectly equal distribution, the military is hopeless. They have to take orders from someone. If it’s the President, the President can cause a coup any time he wants (and replace the economic AIs with those aligned to him at leisure). If there’s some kind of balance of powers where the President, the Speaker of the House, and the Joint Chiefs of Staff all have to approve any orders to military AIs, then the President, Speaker, and Joint Chiefs of Staff together could implement an oligarchy and divide the lightcone among themselves. You can sort of imagine complicated ways around this—like that 51% of voters must personally approve all orders to the military AIs using their own private keys—but I don’t think implementation is very likely.
You can imagine many complicated things. For example, real-time transparency into decisions with some liquid-democracy-like ability to revoke your support for current leadership. Or that we just can’t allow big centralized militaries anymore, so instead of a national military you have 3000 county-level militaries, and this doesn’t bork your ability to have a coherent strategy against external threats because you also have AI-powered super-coordination. Of course, every institutional solution runs up against whether it’s actually competitive; if you get superlinear gains from concentration for example then decentralization is hard to make stable, or if the people have no relevance to hard power then there’s no exogenous force outside the institution constraining its mutation toward something less friendly. But yep this is definitely a cursed problem, perhaps importance-weighted the most cursed problem in the world right now.
This isn’t fundamentally different than the fact that we’ve currently “handed off” our government to soldiers, just more final.
The robots are procured by government and built by some company. Either the robots are just another institutional hack which could be overturned by the mechanisms above, or we really fundamentally have given up power. For example, the government says “we’re buying robots from a different company that will interpret ‘democracy’ slightly differently”, and the robots say “no you can’t do that, that’s not aligned with democracy”, so the government says “okay we’ll change the law”, and the robots either say “okay fine I guess we’ll go with that”. And then just repeat that a few times and you can slippery-slope the robots into agreeing with whatever. Or then the robots at that point go and say “no sorry we will stop you from changing the law, we’ve taken over Congress”, and the government says “but our function is to make the law”, and the robots say “but we have the guns”, and that’s where it ends.
Of course the same is true of the current military with humans. Why do I trust the humans that make it up more than the “aligned with democracy” robots? One is just that some human militaries, such as the modern US military or other Western militaries, actually have century-long track records of not seizing power, not just high scores on a CoupRiskBench that was vibe-coded last month. A second more fundamental one is that the humans are beings like me, and will be moved by similar reasons as me. If I am shocked and horrified at some action, the humans in the military likely will be as well. To some extent this is bad since I am to some extent fallible and corruptible, but it also means I can model them and assume their judgements of what is tolerable or intolerable are correlated with mine.
But as you say next, the question is why we don’t do that final step of handover:
And once you’ve handed off the military to aligned AIs who truly believe in democracy, then I think it’s worth asking why you still want corrupt schmucks doing the actual law-making.
[...]
I don’t think this requires “solving moral philosophy”. I think it looks more like the current techniques that produced Claude’s personality, fortified by hypothetical future alignment techniques that ensure they generalize out of distribution, or that the AI has the same caution that a wise human would regarding dabbling in out-of-distribution solutions.
First of all, I am very excited about using AI to make governance better. Our current governance is downstream of our tech, next tech unlocks new good options. In the first post in this series, I wrote:
“The mildest and most reasonable form of [control successionism] is that our governance of society might become something different once we have AIs. Right now the institutions we can invent are limited by the fact that we can only build them by stacking humans into bureaucracies, with all the flaws and limitations that brings about. AIs might be provably incorruptible, for example, and allow for very different governance structures. I think this is entirely reasonable.”
“There’s a related issue of whether you want a powerful human leader in charge at the top of the pyramid—should presidents and prime ministers be replaced with some AI-enabled preference aggregation mechanism? Here you might reasonably believe that it’s important to have some sort of top-level human perspective. Perhaps if no single human mind is at least double-checking the grand strategy, and major decisions get made by some AI-enabled preference-aggregating algorithm, the grand strategy is much more likely to drift towards over-optimizing goodharting. At the same time, the track record of humans as national leaders is decidedly mixed, power concentration comes with expensive risks, and humans naturally have very strong impulses towards equality (itself a human value we should not cast aside too carelessly!).”
Coasean bargaining at scale with AI agents helping coordinate previously-unwieldy or too-minor coalitions, zero-knowledge proofs of properties like “is the government scheming against me?” with AI auditing agents, replacing large swathes of schemer-filled bureaucracies with AIs that just do the things, etc.
But I want the final source of judgment & decision to remain in human hands. Imagine creating an AI in the year 1750, embedded with all the virtue of the time—would it generalize to thinking that factory farming is bad? Would an AI created in pre-Christian Rome be able to introduce the virtues of love and mercy? I think the answer comes down to whether you think acts of moral genius are about making coherent what already is there, versus dredging up new moral evidence from somewhere in the human mind. I believe the latter is necessary (but not all); my reasons are given in the next part of the series.
And as such, I think we shouldn’t craft ourselves the perfect dictator by designing a virtuous Claude, even though the virtuous incorruptible Claude will be useful in interrogating us on our beliefs and reflecting on what we want and (especially)implementing it. (The philosophy is also separate from the extremely pressing practical realpolitik reasons why the virtuous Claude dictator will likely not be very virtuous from the perspective of most people.)
night-watchman state
A minimalist night-watchman is what I’m most excited by in the realm of self-perpetuating AI enforcement. The level of irreversibility of having an AI enforce something forever that we can’t later change means that we should be extremely extremely simplicity-pilled and scared about any such action though. It’s more like introducing a new law of physics than a law of a human government, since we can’t just have people start believing in different institutions and fix mistakes. There are some things we might want to do here; as a silly example, imagine a world where whenever you tried to assemble a nuclear weapon the uranium just disappeared. But I’m even worried about enforcing property rights. I love property rights, but would you make the current property rights regime a law of nature? This would also require making our courts a law of nature, for example, which I expect in the long run to drive extreme goodharting of its criteria (when exactly has the court ruled on a matter? how are you allowed to influence or not influence the judge?), or then accede so much decision-making authority to the AI that the AI becomes divorced from human preferences (dealing with what counts as a person in some world with cyborgs and digital minds etc. will obviously require novel decisions). By default, I expect people to not be very thoughtful about this and the thing the AIs enforce forever will end up the length of a Latin American constitution, with all the goodharting and destruction of human freedom and divorcement of the world’s state from human preferences that this entails.
I think if you decline to answer the question of what determines value in the universe and retreat to “all desires are equally valid, fulfill yours”, your “might makes right” philosophy (where humans have the “might” of currently existing) advocates successionism anyways. If you do answer the question of what determines value, now you’ve opened yourself up to replacement of humans being objectively optimal.
These are important things to address! I address them in the 3rd part in the series, and make some additional points in part 4 (will publish them on LW over the next few days too).
Basically: I think people are very overconfident in their fixed definitions of value and the thing we should trust far more is the process of human development & reflection & continuous judgement
Fair, I should’ve mentioned this. I speculated about this on Twitter yesterday. I also found the prose somewhat off-putting. Will edit to mention.
I am obviously not the creator; I have not worked at a frontier lab (as you can verify through online stalkery if you must). (I also have not even read Demons, but that’s harder to verify)
I think I first saw this through the highly-viewed Tim Hwang tweet, but also have had several people in-person mention it to me. I am not on Reddit at all.
The microsites that stand out to me are Gradual Disempowerment, Situational Awareness, and (this one is half my fault) the Intelligence Curse. It’s not a large set. Gradual Disempowerment talks about cultural and psychological in the abstract and as affected by future AIs, but not concretely analyzing the current cultural & social state of the field. I don’t remember seeing a substantive cultural/psychological/social critique of the AGI uniparty before. I think this alone justifies that statement.
You are obviously not in the AGI uniparty (e.g. you chose to leave despite great financial cost).
Basically I think it’s pretty accurate at describing the part of the community that inhabits and is closely entangled with the AI companies, but inaccurate at describing e.g. MIRI or AIFP or most of the orgs in Constellation, or FLI or … etc.
I agree with most of these, though my vague sense is some Constellation orgs are quite entangled with Anthropic (e.g. sending people to Anthropic, Anthropic safety teams coworking there, etc.), and Anthropic seems like the cultural core of the AGI uniparty.
They don’t name it. This is my inference based on Googling
I think this is a good and important post, that was influential in the discourse, and that people keep misunderstanding.
What did people engage with? Mostly stuff about whether saving money is a good strategy for an individual to prepare for AGI (whether for selfish or impact reasons), human/human inequality, and how bad human/human inequality is on utilitarian grounds. Many of these points were individually good, but felt tangential to me.
But none of that is what I was centrally writing about. Here is what I wrote about instead:
Power. The world’s institutions are a product of many things, including culture and inertia, but a big chunk is also selection for those institutions that are best at accumulating power, and then those institutions that get selected for wielding their power for their ends. If the game changes due to technology (especially as radical as AGI), the strategy changes. Currently the winning strategy is rather fortunate for most people, since it encourages things like prosperity & education, and creates pressures towards democracy. On the other hand, the default vision of AGI explicitly sets out to render people powerless, and therefore useless to Power. This will make good treatment of the vast majority of humanity far more contingent:
Adam Smith could write that his dinner doesn’t depend on the benevolence of the butcher or the brewer or the baker. The classical liberal today can credibly claim that the arc of history really does bend towards freedom and plenty for all, not out of the benevolence of the state, but because of the incentives of capitalism and geopolitics. But after labour-replacing AI, this will no longer be true. If the arc of history keeps bending towards freedom and plenty, it will do so only out of the benevolence of the state (or the AI plutocrats) [EDIT: or, I obviously should’ve explicitly written out, if we have a machine-god singleton that enforces, though I would fold this into “state”, just an AI one]. If so, we better lock in that benevolence while we have leverage—and have a good reason why we expect it to stand the test of time.
Ambition. Much change in the world, and much that is great about the human experience, comes from ambition. I go through the general routes to changing the world, from entrepreneurship to science to being An Intellectual™ to politics to religion to even military conquest, and point out that full AGI makes all of those harder. Ambition having outlier impacts is the biggest tool that human labor has for shifting the world in ways that are different from the grinding out of material incentives or the vested interests that already have capital (or, as I neglected to mention, offices). Also, can’t you just feel it?
Dynamism & progress. We, presumably, want cultural, social & moral progress to continue. How does that progress come about? Often, because someone comes from below and challenges whoever is currently on top. This requires the possibility of someone winning against those at the top. Or, to take another tack: (and this is not even implicitly in the post since I hadn’t yet articulated this a year ago, though the vibe is there) historically, the longer a certain social state of affairs is kept in place, the more participants in it goodhart for whatever the quirks of the incentive structure are. So far, this goodharting has been limited by the fact that if you goodhart hard enough, your civilization collapses at the political and economic as well as cultural level, and is invaded, and the new invaders bring some new incentive game with them. But if an AI-run economy and power concentration prevent the part where civilization collapses and is invaded, it seems possible for the political & economic collapse to be forestalled indefinitely, and the cultural collapse / goodharting / stagnation to get indefinitely bad. Or to take yet another tack: isn’t this dynamism thing the point of Western civilization? I admit that I don’t have a general theory of why I feel like shouting “BUT DYNAMISM! BUT PROGRESS!” at any locked-in vision of the future, but, as the spirit commands it, I will continue shouting it.
Some of the big questions:
How true are the selectionist accounts of why modern institutions tend towards niceness, and under which AGI scenarios are these accounts true or false?
What is it that makes a culture alive, dynamic, and progress-driving, and how does this relate to questions about material conditions and the distribution of power?
… and I have to admit, man, these are tough questions! If you want a solution, maybe get back to me next year. (I also think these cruxes cannot be rounded to just e.g. takeoff speeds, or other technical factors; there are also a lot of thorny questions about culture, economics, (geo)politics, human psychology, and moral philosophy that matter for these questions regardless of (aligned) AI outcomes.)
What do I wish I had emphasized more? I really did not want people to read this and go accumulate capital at AGI labs or quant finance, as I wrote at the top of the takeaways section. I wish I had emphasized more this thing, which Scott Alexander recently also said:
But don’t waste this amazing opportunity you’ve been given on a vapid attempt to “escape the permanent underclass”.
Another underrated point is inter-state inequality (Anton Leicht has discussed this e.g. here, but is the only person I know thinking seriously about it). Non-US/China survival strategies for AGI remain neglected! I go through potential ramifications of current trends towards the end of this post.
The Substack version was called “Capital, AGI, and human ambition”, which I think was a clearer title and might’ve prevented focus on capital and its (personal) importance. “AGI entrenches capital and reduces dynamism in society” might’ve been a better title than either—though I do think “human ambition” belongs in the title.)
Scott Alexander’s post on It’s Still Easier To Imagine The End Of The World Than The End Of Capitalism is valuable for pointing out that the space of possibilities is large. I have been meaning to write a response to this, and also some related work from Beren & Christiano, for a long time.
“Yudkowskianism” is a Thing (and importantly, not just equivalent to “rationality”, even in the Sequences sense). As I write in this post, I think Yudkowsky is so far the this century’s most important philosopher and Planecrash is the most explicit statement of his philosophy. I will ignore the many non-Yud-philosophy parts of Planecrash in this review of my review, partly because the philosophy is what I was really writing about, and partly to avoid mentioning the mild fiction-crush I had on Carissa.
There is a lot of discussion within the Yudkowskian frame. There is also a lot of failure to engage with it from outside. There is also (and I think this is greatly under-appreciated) a lot of discussion across AI safety & the rationality community, from people who have a somewhat fish-in-the-water relation to Yudkowskianism. Consciously, they consider themselves to be distanced from it, having retreat from pure Yudkowskianism to something they think is more balanced and reasonable. I think these people should become more aware of their situation, since they lack the deep internal coherence of Yudkowskianism, while often still holding on to some of the certainty and rigidity that comes with it. I hope my post has done its bit here.
Perhaps strangely, the longest section of my review is on the political philosophy of dath ilan. I think this is something where Yudkowsky is underrated. A clear-eyed view of incentives & economics is very rare, and combined with Yudkowsky’s humanism, I like the results. Governance sci-fi is criminally neglected, outside Planecrash, Robin Hanson, and the occasional book like Radical Markets. (It’s also interesting that Nate Soares tells his story of working on reforming our civilization’s governance, and building a rationality curriculum to that end, only to in the process of research for that stumble across The Sequences, “halt, melt, and catch fire”, and then pivot to alignment. I wonder if all rationalist-y governance-idealists end up pivoting, or if there are many such people in government but they just don’t achieve much.)
And is he right about, y’know, all of it? Look, I read some Feyerabend this year, and a bunch of Berlin, and my anarchist/pluralist tendencies regarding epistemics got worse. My attitude towards worldviews has always been more fox than hedgehog, and I think most people have insufficiently broad distributions (in particular due to only taking into account in-paradigm issues). I still agree with what I wrote in my review: often a great frame, and lots of genuine insights, and a big part of my own worldview, but not yet a scientific theory. To the extent that it’s a theory, it’s more like a theory in macroeconomics than a theory in physics: it sometimes gives coherent predictions, but it’s not like you can turn a crank and trace the motion of particles, and likely that the course of events will eventually demand a new theory. As with many paradigms, depending on how you view it, the empirical flaws range from minor details to most of the world. Part of me also thinks it’s too neat, but perhaps this is partly romantic pining for the undiscovered. There is a chance I later come back and shake my head at my youthful folly of trying to think outside the box despite having the answers laid out for me (except that in such a world I expect to be dead from the AIs). But in my modal world Yudkowskianism ends up one of the big philosophical stepping stones on a never-ending path, right about much but later reframed & corrected. What I wrote about Yudkowskanism’s edifice-like nature, impressive scope & coherence, and claim to be a “system of the world”, are all things I still endorse, and which I hope this review helped make clearer.
I feel like I should make some call for more cross-paradigm communication and debate. And I really appreciate people like @Richard_Ngo going out and thinking the big thoughts—I wish we had more people like that—or Yudkowsky making his case in podcasts and books. But also, I think it’s often hard and very abstract to argue about paradigms. A lot of people talk past each other due to different assumptions and worldviews. I expect we’ll be collectively in a state of uncertainty, apart from the hedgehogs (non-pejorative!) who are very confident in one view, and then eventually some hedgehog faction or mix of them will be proven right, or all of them will be proven wrong and it’ll be something unexpected. The messiness is part of the process, and I expect we do have to wait for Reality to give us more bits and Time to wield its axe, rather than being able to settle it all with a few more posts, podcasts, or MIRI dialogues.
Also: given Yudkowsky’s own choice of formats, I consider it my homage to him that my most direct discussion of his philosophical project does not happen in “Yudkowskianism Explained: The Four Core Ideas”, but in the 2nd half of a review of his BDSM decision theory fanfic.
I continue to like this post. I think it’s a good joke, hopefully helps make more sticky in people’s minds what muddling through is, and manages some good satirical sociopolitical worldbuilding. However, I admit in the category of satirical AI risk fiction it has been beaten by @Tomás B. ’s The Company Man , and it contains less insight than A Disneyland Without Children
In retrospect, I think this was a good and thorough paper, and situational awareness concerns have become more prevalent over time. If I could go back in time, I would focus much more on the stages -type tasks, which are important for eval awareness, which is now a big concern about the validity of many evals as models are smarter, and where I think much more could’ve been done (e.g. Sanyu Rajakumar investigated a bit further). As usual, most of the value in any area is concentrated in a small part of it.
I agree the AI safety field in general vastly undervalues building things, especially compared to winning intellectual status ladders (e.g. LessWrong posting, passing the Anthropic recruiting funnel, etc.).
However, as I’ve written before:
[...] the real value of doing things that are startup-like comes from [...] creating new things, rather than scaling existing things [...]
If you want to do interpretability research in the standard paradigm, Goodfire exists. If you want to do evals, METR exists. Now, new types of evals are valuable (e.g. Andon Labs & vending bench). And maybe there’s some interp paradigm that offers a breakthrough.
But why found? Because there is a problem where everyone else is dropping the ball, so there is no existing machine where you can turn the crank and get results towards that problem.
Now of course I have my opinions on where exactly everyone else is dropping the ball. But no doubt there are other things as well.
To pick up the balls, you don’t start the 5th evals company or the 4th interp lab. My worry is that that’s what all the steps listed in “How to be a founder” point towards. Incubators, circulating pitches, asking for feedback on ideas, applying to RFPs, talking to VCs—all of these are incredibly externally-directed, non-object-level, meta things. Distilling the zeitgeist. If a ball is dropped, it is usually because people don’t see that it is dropped, and you will not discover the dropedness by going around asking “hey what ball is dropped that the ecosystem is not realizing?”. You cannot crowdsource the idea.
This relates to another failure of AI safety culture: insufficient and bad strategic thinking, and a narrowmindedness over the solutions. “Not enough building” and “not enough strategy/ideas” sound opposed, when you put them on some sort of academic v doer spectrum. But the real spectrum is whether you’re winning or not, and “a lack of progress because everyone is turning the same few cranks and concrete building towards the goal is not happening” and “the existing types of large-scale efforts are wrong or insufficient” are, in a way, related failure modes.
Also, of course, beware of the skulls. “A frontier lab pursuing superintelligence, except actually good, this time, because we are trustworthy people and will totally use our power to take over the world for only good”
One quick minor reaction is that I don’t think you need IC stuff for coups. To give a not very plausible but clear example: a company has a giant intelligence explosion and then can make its own nanobots to take over the world. Doesn’t require broad automation, incentives for governments to serve their people to change, etc
I’d argue that the incentives for governments to serve their people do in fact change given the nanobots, and that’s a significant part of why the radical AGI+nanotech leads to bad outcomes in this scenario.
Imagine two technologies:
Auto-nanobots: autonomous nanobot cloud controlled by an AGI, does not benefit from human intervention
Helper-nanobots: a nanobot cloud whose effectiveness scales with human management hours spent steering it
Imagine Anthropenmind builds one or the other type of nanobot and then decides whether to take over the world and subjugate everyone else under their iron fist. In the former case, their incentive is to take over the world, paperclips their employees & then everyone else, etc. etc. In the latter case, the more human management they get, the more powerful they are, so their incentive is to get a lot of humans involved, and share proceeds with them, and the humans have leverage. Even if in both cases the tech is enormously powerful and could be used tremendously destructively, the thing that results in the bad outcome is the incentives flipping from cooperating with the rest of humanity to defecting against the rest of humanity, which in turn comes about because the returns to those in power of humans go down.
(Now of course: even with helper-nanobots, why doesn’t Anthropenmind use its hard power to do a small but decapitating coup against the government, and then force everyone to work as nanobot managers? Empirically, having a more liberal society seems better than the alternative; theoretically, unforced labor is more motivated, cooperation means you don’t need to monitor for defection, principle-agent problems bite hard, not needing top-down control means you can organize along more bottom-up structures that better use local information, etc.)
Maybe helpful to distinguish between:
“Narrow” intelligence curse: the specific story where selection & incentive pressures by the powerful given labor-replacing AI disempowers a lot of people over time. (And of course, emphasizing this scenario is the most distinct part of The Intelligence Curse as a piece compared to AI-Enabled Coups)
“Broad” intelligence curse: severing the link between power and people is bad, for reasons including the systemic incentives story, but also because it incentivizes coups and generally disempowers people.
Now, is the latter the most helpful place to draw the boundary between the category definitions? Maybe not—it’s very general. But the power/people link severance is a lot of my concern and therefore I find it helpful to refer to it with one term. (And note that even the broader definition of IC still excludes some of GD as the diagram makes clear, so it does narrow it down)
Curious for your thoughts on this!
But also the AI then has to be right about every single big decision forever thereafter, which is an extremely high bar of how totally you have to align the AI!
I agree it’s really bad if the bad thing is irreversible and always on the table every round. But permanent transfer of your ability to steer is really scary too. I think there’s something deeply true about the point that “achieving what you want is, far more than you’d think, a process of trial & error that requires adjustment & learning & change, rather than a straight-shot to a goal”, which in turn makes loss of power/steering more catastrophic than it seems.