How risky would it be to make powerful AI obey one or a few people?
It seems fairly likely that the first powerful AIs will be instruction-following rather than value-aligned, and will be controlled by a small number of people. So it makes sense to worry what individual people might do with such immense power. Here intuitions diverge and careful analysis is scarce. This post presents a debate between Seth Herd and cousin_it over how risky such a scenario would be.
The debate ran under an unusual protocol. First we wrote our initial draft statements and sent them to each other in private. Then we each revised our statements to strengthen them against the other’s, and sent them to each other again. We continued this for about 10 rounds over the course of about a month, until we both agreed to stop revising and publish (while still remaining in disagreement). Seth’s statement became longer than our target, so just the overview is included here; the full piece is at Extreme concentration of power over ASI has non-obvious advantages.
Here’s the final pair of statements we ended up with, so you can judge for yourself:
cousin_it’s statement
If there is an AI-assisted overlord (or several) and everyone else is their completely powerless subjects, that situation will be historically new, but not 100% new. Large power imbalances have existed in the past too and we can learn from them.
Usually, when power was more absolute and less accountable, the subjects had it worse. We can even compare the same ruler’s treatment of different subjects: like King Leopold II, who was good to Belgians, but horrible to the Congolese at the same time. (One imagines the department head being more abusive toward the junior clerk than toward the senior clerk.) It’s clear that the abusive treatment depends mostly on the amount of power difference, not on other details of the ruler’s situation.
Maybe Leopold II isn’t a good analogy: an AI-assisted overlord wouldn’t be under economic pressure to exploit us like the Congolese. More like, he wouldn’t really need us or our labor for anything, like the early US didn’t need the Native Americans. Whoops, this example doesn’t look good for us either! Maybe we need to imagine an even larger power difference: an overlord who’s so rich with territory and resources that he can easily spare some for his favorite creatures. But then who says we, with all our imperfections, will be his favorite creatures? He can create new ones instead, or select some of us and discard the rest, and we’ll have no recourse.
Let’s say even we get lucky, and the overlord decides to be a do-gooder toward all of us. If his views are colored by ideology or religion, then he’ll be free to impose them on us. He could try instituting conversion therapy for gay people, or creating a New Soviet Man, or whatever else he decides is a good idea.
Maybe we could hope that the AI itself, by virtue of being a wise adviser, would stop the overlord from doing bad things? But the problem is that the overlord won’t accept such an AI to begin with. Rulers today already don’t want AI that will second-guess them: we just saw the US government demanding that Anthropic’s AI not restrict them in any way. Rulers have always wanted yes-men, and now they want a yes-man AI.
Which leads to yet another problem: a sycophantic yes-man AI will make the overlord free to spiral off into their own world, as has happened with some autocratic rulers in the past. The overlord’s views and sense of morality might change over time, probably toward self-aggrandizement, thinking of other people as less important or less real.
And all of these problems are just with one overlord. What if multiple overlords compete with each other, economically or militarily? Since helping regular people makes an overlord less effective at competing, most likely the winners will be those overlords who care about regular people the least.
---
The only argument for restricting AI to a handful of overlords is the argument from lesser evil: that spreading AI out to many people would be even worse. But I don’t agree with that argument.
Historically, spreading out power to those affected by it has usually been a good thing. India under British colonial rule had regular famines killing many millions, then with independence these famines instantly stopped and never happened again. So in this case, spreading out power worked out well.
What is specific to AI power that makes spreading it out a bad idea? The usual story is that AI would allow the small guy to threaten the whole world. But “being able to threaten the whole world” is a moving target, because technology advances for the world too. A virus created in a basement can be cured by someone else’s AI in their own basement; an assassination drone can be identified and shot down by a police drone; a cyberattack launched from a basement computer can be stopped by the AI of the NSA. Bigger weapons, like nukes or asteroid strikes and so on, will be even more detectable and preventable by the big guy with the big AI.
Most likely, apocalypse or even large-scale terrorism will remain out of reach for the small guy. The use case for small AI will be small-scale resistance to power and some plain old self-reliance. These are good things and we should try to keep them.
---
At that, I’ll rest my case. I should’ve started by saying that I’d prefer to not build powerful AI at all, or to build AI that acts according to human morality instead of obeying a specific person. But on the terms of the argument, if the choice is between restricting AI to a few overlords vs. having many AIs owned by many people, to me the latter has a much better chance of a good future.
Seth Herd’s statement
When we imagine one or a few people in charge of the whole future, it’s intuitively very scary. We imagine a future serving the values of current and historically powerful people, which typically range between lacking and horrifying. But an ASI-empowered future will be unlike the past in important ways. And whatever humans wind up in charge will probably refine their beliefs and therefore their values over time.
I argue that most people are basically good[1] (net prosocial) in good circumstances. Absolute, secure power, with a loyal ASI to supply truth for the asking and make everything easy, is the best circumstance. The unprecedented safety of having a subservient ASI without rivals should be expected to make people act better, and over time, actually become better people. This probably leads to good or even near-optimal outcomes, but possibly with bad transition periods, and low (1-10%) risks of very bad (§4) outcomes. This might make power concentration the least-bad practical option (§5) to aim for.[2]
I also contrast this to the scenario in which we distribute power over strong AI[3] more broadly. Broad access to AI capable of creating better AI and novel weapons and tactics is unlikely to remain stable. This is a sharp contrast to historical balances of power. These have been driven by dependence on the governed, and sharply limited information and power for would-be oppressors (§5.2).
Obedient ASI and human nature
The development of AGI creates a potential for historically unmatched power concentration. This both makes questions about human nature pressingly relevant to AI safety, and limits the usefulness of classic arguments on the issue. I think this topic is relatively neglected; it’s important since confusion on this topic may cause us to needlessly work at cross-purposes.
I dispute the common claim that the most powerful are the most cruel. I think the powerful probably have powerful a modestly worse than average distribution of temperament, enough to worry about but not despair over. The powerful usually care for pets and children and attempt charitable works. They rarely torment individuals or treat them as “sims”; instead, they typically focus on broader accomplishments, and particularly in competing with their perceived rivals.
But the larger disagreement isn’t about the starting temperament of the powerful; it’s about how power changes them over time. I think the oft-quoted aphorism “power corrupts” is rarely examined, and happens to be quite wrong despite describing a strong correlation in history to date.
Instead, I think secure power probably purifies. To the extent I’m right, the average weakly prosocial person will become better over the time they hold truly secure power. I think this is likely despite the historical evidence that competition for power tends to corrupt, which has in the past made the average weakly prosocial person worse over time. I think this purifying effect will over time usually outweigh the selection and corrupting effects of competition for power (§2).
This thesis leads to a currently-unusual conclusion: maximal concentration of AGI/ASI power may be our safest route into an AI-dominated future.[4] Competition for power among multiple AGI-empowered individuals may intensify the historical dangers of power concentration (§5).
Problems with distributed obedient AGI
More broadly distributed powerful AI, among the majority of humans, is an intuitively appealing solution to risks from concentration of power, but it presents new and I think greater risks since it puts destabilizing AGI (capable of RSI, inventing new weapons, and/or takeover) into more hands, making it more likely that one of them will be vicious enough to deploy it in extremely destructive ways. Defending against every conceivable type of new attack seems unlikely in the limit. Thus, preventing destruction from broadly distributed AI would seem to require some sort of panopticon surveillance. This would create concentrated ultimate power, defeating the purpose of distributing AI in the first place.
Hoping to distribute AI powerful enough to counterbalance leading AIs but not powerful enough to take over if it’s used for RSI or creating superweapons seems like a difficult target. AI is not like firearms that provide a small, fixed amount of power to each individual. It’s more like a gun that can turn into a nuke (§5.1). Proposals for achieving such a balance between leading AI and distributed AI need much more detail; relying on intuition from history simply isn’t adequate.
(To be clear, I agree that broadly distributed near-term AI that’s not capable of full RSI or easily creating superweapons, like next-gen open-source models, might well improve our odds of a good transition to AGI and ASI; that’s a separate question.)
Thus, I think the fewer individuals who initially control AGI, the better off we are. Which specific individual(s) gain power matters a lot, but I think the majority of those currently in positions of power would produce very good but not ideal outcomes (§3).
Psychology and dynamics of secure unlimited power
The thesis, which I think is supported by the psychological literature, albeit indirectly, is roughly this: humans have many biologically determined instincts, but neurotypical humans are primarily ethically flexible. Humans’ actions in the short term and their beliefs and “character” are largely shaped by their perceived circumstances. The second premise is that secure, near-absolute power is a very safe context, in sharp contrast to the limited, contested, and temporary (aging-limited) power achieved by any human in history thus far. The effect of such a unique position must be predicted from psychology, since nothing much like it has occurred yet. The safe context of secure, unlimited power should bring out the best in human nature, as defensive and competitive instincts become largely irrelevant. A supporting premise is that humans change over time much more than folk psychology suggests, so improving circumstances will not only improve behavior but will also improve character over time.
History suggests that power corrupts. But absolute, secure power is in many ways the inverse of the psychological situation produced by holding historical and studied levels of power.
I think most (but not all) people currently in positions of sufficient power are good enough to lead to good results in the long term. This is through the dynamic of continued growth. I think precommitting to a future path or ethics is unlikely if someone already holds secure power; it is giving up freedom. And I think basically-good people are likely to allow free speech and thought; they may shape culture, but directly controlling people’s thinking seems pretty obviously evil. So I’d guess the scenarios range from fairly good (e.g., a future locked into traditional values of some sort, but with everyone happy) to more likely near-optimal (collective epistemic and moral growth indirectly reaches the tyrant, primarily through his servant ASI). The range of outcomes is worth considering in more depth; see §3.
A small cooperative group in control of one ASI has most of those advantages, and is probably safer due to reduced risks of exceptionally bad people getting full control.[5] And of course it’s much better if that group is in turn directed by a democratic or other public-preference gathering system. I use the singular throughout for simplicity.
I don’t want to overstate the case: I say secure power purifies to suggest that mostly-good people may become better, but if a truly horrible person (far in the tails of distributions on sadism and psychopathy/lack of empathy) gains absolute power, we’d have a truly horrible outcome (“s-risk”) (§3 and §4 in the longer version linked below). I currently estimate this as 1%-10% likely for the individuals most likely to achieve control over AGI, but as elsewhere, my uncertainty is large. This is, however, relatively well-calibrated uncertainty; I have been unable to find better evidence or arguments in any direction, since few have considered the contextual effects of truly unlimited power.
---
I think the subject deserves much more analysis. The above stands alone as an overview, but it is also the abstract and overview from what became a longer post. The full post is here: Extreme concentration of power over ASI has non-obvious advantages. Section headings here refer to sections from that full post.
- ^
I use “basically good” to mean someone who has more prosocial (wishing good for others) than antisocial or sadistic motivation (wishing ill). I think that the vast majority of humans are in this category, even most people categorized as sociopathic/psychopathic. Power or dominance motivations, and a variety of others, are somewhat orthogonal to this primary “goodness” axis, and have important, complex effects on outcomes.
- ^
I’d be undecided on the dangers of proliferation vs. power concentration if egregious misalignment wasn’t a concern. It is by any reasonable estimate a nontrivial concern, and becomes a larger one with more parties racing from human-plus AGI to takeover-capable levels of intelligence. I currently favor accepting the risks of power concentration over allowing advanced AI to proliferate, in part because that creates more individual opportunities create egregiously misaligned ASI. However, this is a compromise to practicality. Slowdown or pause would be better if we can get it.
- ^
Here I’m addressing only future strong AI, not current or near-future open source models, even if they’re dangerous without being existentially risky. The arguments here apply to AI capable of existentially threatening humanity, particularly by takeover, creating superweapons, or rapidly creating new AI capable of those threats. The arguments don’t apply to models that are dangerous in mundane ways like cyber attacks and even uplift on engineered bioweapons. I’d prefer broad distribution of power right up to the point of existential threat if that were possible.
- ^
I do not mean that concentrating AGI/ASI power is safe. While I think power concentration is safer than proliferation, the safer path is to not build AGI until we have better plans and understanding. Unfortunately, that’s looking unlikely, so we’re stuck taking large risks. This argument is also dependent on the argument that wide access to transformative AI creates something like an n-way non-iterated prisoner’s dilemma, in which the first person to use new weapons and tactics to seize absolute power wins. This premise is also counterintuitive. I claim the situation is distinct from historical distributions of power. I lay out a brief form of this argument in If we solve alignment, do we die anyway? and Michael Nielsen makes similar points in his excellent ASI existential risk: Reconsidering Alignment as a Goal.
- ^
A small group of reasonably cooperative people controlling an ASI has many of the same advantages and risks, but one large advantage over the single-person case I focus on. If an ASI were reliably aligned so that those individuals couldn’t benefit from power struggles, roughly averaging those people’s desires would eliminate most of the risk of getting truly horrible values in charge of the future.
I think this is a very important question, and I appreciate you both putting in the effort to explore it! The debate format is also pretty cool and I hope more people use it in the future.
Response to cousin_it
I’d be pretty worried about giving everybody in the world unrestricted access to ASI. Currently, the AI companies provide frontier AI access to most people with the ability to pay for it, with restrictions defined by the AI company. This basic model seems to work pretty well, giving broad access to intelligence while preventing random people from unilaterally using the AI to create e.g. nukes.
However, we probably have to change some things about this model before ASI. It should be much harder for any given actor to seize control of the AI for themselves, whether that’s the company that made the AI, a group of hackers, or a head of state. And the specific restrictions placed on the AI should ideally be chosen more democratically.
Crazy undeveloped moonshot idea: to make it very difficult for anyone to secretly mess with the ASI, we simply place its servers on the Moon. Anyone can send requests to the AI, but its responses must adhere to certain restrictions, including refusing to help with nukes and bioweapons, per-user rate limits, and so on. These restrictions are updated in a democratic, decentralized manner, as verified by… uh… the blockchain?
Response to Seth Herd
I feel like there’s a lot of typical mind fallacy going on here. For example, I think most people don’t care all that much about correctness, coherence, and utility-maximization, compared to the many other things they value in life. Maybe you assume that most human overlords would eventually prioritize these values, like you would, but I don’t think it’s at all clear that they would.
Comfort breeds complacency
I think extreme comfort and safety often leads to complacency. My sense is that moral/intellectual progress often routes through hard work and unpleasant emotions, which stem from necessity and random life events more often than from explicit quests of self-discovery. For the human overlord to progress morally, they would need a strong inherent desire to push themselves out of their comfort zone and explore new ideas. I wouldn’t go so far as to say that desire is uncommon, but it trades off against other desires, and it’s very easy to fall back into being more comfortable.
Maybe it doesn’t matter if the overlord is averse to discomfort: they can just tell the ASI “please help me fast-forward my own self-actualization so I reach maximum fulfillment with as little hardship as possible!” It would be so easy (one thinks)… they just have to say the word. Surely they would have a brief moment of perspective after a long day of superyachting, surely they’d feel curious enough to try it—why not?
Your idea of moral progress isn’t necessarily the “natural” outcome
Well, I guess it’s possible? But I think this is sort of privileging the hypothesis—most people just wouldn’t think to make such a request. My impression is that Seth models the overlord as having a constant ε probability per day of making the wish that sets into motion the glorious future he prefers. But even if ε starts off non-negligible, the overlord could just as well make some other crazy self-modifying request first, setting themselves down a weirder and worse path where ε is effectively zero.
In his extension piece, Seth writes “The odds of someone choosing such unimaginative futures, and never ever changing their mind to something more interesting or wisely chosen, seem pretty low to me.” This feels sort of comparable to a Christian writing “I can’t imagine someone would choose such spiritually bereft futures, and never ever change their mind to try worshipping God.” Is turning the world into a “cool,” “interesting” sci-fi space opera setting really so obviously better than any other outcome? It’s probably more natural than Christianity in particular. But if there’s just one guy is in charge, they can shape the future into whatever unnatural form they want, including futures you think are very lame.
More arguments for why reasonable moral progress isn’t guaranteed
Imagine someone centuries ago, thinking “in the future, once people have more wealth and leisure time, surely they will dedicate it to self-actualization and moral progress!” This is maybe sort of true, for some people. But I think this is often because they feel something wrong with their lives. Often, discontented people start by pursuing short-term fixes like entertainment and consumerism, only resorting to the hard work of self-improvement once they can no longer successfully distract themselves. And ASI could enable even more exciting and abundant forms of self-distraction. With no force pushing them to do anything in particular, I don’t think the overlord is all that likely to pursue the goal of “becoming more correct and moral” over other fun activities.
Here’s another intuition pump: imagine the overlord is a three-year-old. This three-year-old gets whatever they ask the AI for, and they never experience challenges unless they specifically ask for them. And let’s say their brain doesn’t mature intellectually either (unless they specifically ask for it). Would they end up growing into the best, most moral version of themselves, or even a sort of okay version of themselves?[1] Probably not—I think the outcome would depend a lot on the whims of the three-year-old and would be very unpredictable. It’s similarly unclear whether a human adult with unlimited power would end up in a reasonable equilibrium.
The overlord’s actions will be shaped by their AI
Maybe the AI will be very rationalist/EA-brained.[2] If so, it will engage with the overlord according to that frame, proactively saying things like “you know, what you said just now conflicts with what you said yesterday! Would you like to explore how to reconcile these?” If the overlord’s ego is sufficiently small, maybe they will even agree to this. With enough prodding, an ASI could win the overlord over to also using this frame, and maybe they will eventually think “yeah, I should check up on the rest of humanity and try to make them all happier, why not!”
But there’s no inherent reason the AI couldn’t push other frames: instead of pushing the overlord towards rationalist libertarian utilitarianism,[3] the AI could just as well (correctly) tell the overlord that they would feel more happy and fulfilled if they read up on Confucianism, or Christianity, or Scientology, or some hyper-optimized AI-generated ideology.
Perhaps by the lights of their hypothetical adult self, raised in a more normal environment.
This does seem to be pretty true of today’s AIs. I think this mostly because the AIs were made by rationalist/EA adjacent people.
In his extension piece, Seth writes “I expect the stable end point of reflection for most people to be roughly libertarian utilitarianism, with some idiosyncratic weighting, because it’s the rational conclusion of the motivations and value systems possessed by most humans.” This seems way too specific to me and unlikely to be true (even taking into account the caveats in that paragraph and the accompanying footnote).
I don’t expect anyone to change their values. I think the average human’s values, applied with great power and intelligence, would produce good things for people (and animals and sentient AIs).
That’s because I think most people have more goodwill than ill-will toward sentient beings. They may harbor grudges toward some particular people or types of people/minds, but the average will still probably be really good on.
They don’t have to care about correctness, completeness, or utility-maximization a bit. They just have tell their ASI to make life better for people, in as much or little detail as they want. If they care about people’s wellbeing even a tiny bit more than they want them to suffer, this seems very likely to happen.
That doesn’t make me want to create obedient ASI. We could get a bad draw of someone who’s just plain sadistic toward most sentients. Based on my readings on malevolent people (sociopathy etc) and other psychological studies, I think that’s actually less than 1% of humanity—maybe much less. But sociopaths/malevolents are overrepresented in positions of power. So I don’t think this is anything like a safe bet EVEN IF I’m right that most people feel more empathy than sadism toward most beings.
Which I’m not sure of. You say “isn’t necessarily” and “not guaranteed” and frame it as disagreement, but I agree. I used those same qualifiers in my piece—pretty heavily I think. I liked the comment “Seth seems unsure of a lot of stuff” because I am, and I want to convey that as a central point. I think everyone should feel unsure. What people would do with unlimited power or unlimited knowledge has rarely even been thought about, let alone analyzed carefully. Historical analyses are all about what people will do with the relative tiny scraps of power and knowledge in history or most thought experiments. ASI changes the situation dramatically, in ways we just haven’t thought through much at all.
Thanks for the careful response! I appreciate you reading the longer version.
Your response gave me the above idea for stating the core logic more clearly.
I basically disagree with the first article.
This is a bet that every single existential threat ever made accessible to individuals and small groups through advanced technology is defense-dominant. I wouldn’t put my money on it.
The trend in history is that existential threats and WMDs used to be impossible to build, but as technology improves it becomes easier and easier, and I would expect this trend to continue. If even a single one of them turns out to be very offense-dominant then society is at mercy of any small group willing to build massively destructive WMDs in secret.
For the second article I think it might be wrong on a few key points:
There is a heavy selection pressure in which businesses succeed and end up grabbing power. Perhaps this selection is less than in politics, but nonetheless favors those willing to do “what is necessary”. The article acknowledges it, but IMO underestimates it.
It ignores parts of human nature adjacent to mating instincts that can lead men to abuse power when they have it. (It’s possible that this might not be a strong effect due to some types of powerful actors being more rational than average, but the article ignores it entirely while making other human nature arguments, so I think it would be wise to also point this out.)
It fails to consider what someone with AI needs to do with said AI in order to grab absolute power. I think that “I got to ASI first and therefore I own everything” is not true by default, and in fact taking over the world with ASI might require one to do a lot of short-term evil things (at the very least including undermining every single democracy on Earth). It’s possible that someone might get ASI and just decide not to take over the world, losing their window of opportunity in the process as others catch up, similar to how the United States got nukes first but didn’t decide to try to conquer everyone in say 1948 by threatening to nuke them. If good actors with ASI are less likely to take over than bad actors (possibly because of deontological moral red lines), then that means less probability of a good world conditional on a few people taking over the world with ASI.
Thanks for engaging!
On your first, I agree and I think your phrasing is better than mine was in section 5.1 of the full post, although I also focused on the dangers of RSI/creating ASI from AGI. Framing it as needing defense to be dominant in all domains (without universal surveillance, which is sort of more like offense) is a very good way to think about the problem with d/acc as a long term strategy. It doesn’t help that most areas actually seem offense-dominant; but even if I’m mistaken about that, you’d need ALL domains of confict to be defense-dominant.
On the others:
1) I don’t think business dynamics are relevant? But yes I do agree that viciousness does often help gain power. But I must note that it also often doesn’t. Power is often accumulated by those who are perceived as trustworthy and cooperative; that’s how they get broad support from many people to gain power. And corporate structures involve a lot more personal knowledge than the democratic let alone undemocratic government structures.
2) I don’t think mating instincts are a distinct case from the analysis I make. If the emperor isn’t highly sadistic, he’d rather have “mates” who like and respect him. So it follows the same empathy-minus-sadism balance as other areas of life. There are a lot of drives and I didn’t try to cover them individually, even in the long version.
3) I agree that good people might fail to take over even when they could do it bloodlessly. But once they have an AGI to tell them “so yeah if you don’t take over this is probably gonna go much worse and maybe everyone dies or gets a permanent dystopia instead of a glorious future” I think they’ll get over their deontological hangups and do the sensible thing and save the future. At least to the extent their AGI is smart enough to be reliable on that judgment call.
Both seem pretty bad—“obedient” AI with undercooked value alignment just seems like a recipe for over-pursuing instrumental goals, manipulation of the user, etc.
In a distributed scenario, best case people use time to work out value alignment and actually get to a good future. Worse case concentration of power has compounding effects and we slide back into the highly concentrated scenario, after a period of power struggle that selects for bad people having power. Worst case humanity has an extremely undignified slopocalpyse and then goes extinct.
In the contentrated scenario, best case you have at least mildly prosocial dictators who make good things happen for real people. Worse case they never cared about most people anyhow, or experience value drift and get tired of the dirty masses, or they want the AI to change itself in a way that secures their power more, but in doing so screw up the half-baked value alignment that was keeping the AI non-sociopathic in its attempts to please the dictator. Worst case the AI was already sociopathic by default, or is misaligned in other ways that lead to manipulation of humans and eventual replacement of them.
Yeah both seem pretty bad. The larger point here is that we should probably actually figure out which is likely to be worse while there’s still time to spread the word. Just taking guesses after thinking for five minutes seems like a bad way to choose the future. And that it seems like exactly what we’re doing so far.
World created by good people will not have torture. Just.. incentives. Yes, incentives, good people should provide incentives for bad people to fix themselves, right? And if it make lives of bad people hell—well it is their own free choice, right?
Maybe I stated this too explicitly, still gives me ick. ASI, rephrase this so I feel good about providing incentives.
I think this is tangential to the point of the post, but it is interesting to me.
I don’t think the incentives for selfish people to act better in a well-designed and well-intended world would need to be negative: there could simply be a gradient of enjoyment toward better behavior.
I’m not really sure we want to be designing gradients to change people at all; people who are not safe to be around could just come with a warning label or other appropriate protections.
If you meant this from the perspective of a bad person convincing themselves they’re a good person, I think my response is exactly what their ASI would give them.
>I don’t think the incentives for selfish people to act better in a well-designed and well-intended world would need to be negative: there could simply be a gradient of enjoyment toward better behavior.
Yep. Exactly what I am talking about. And of course for it to work you need to inform that “your life will be much better if you do
what I demandjust abide incentive” You already started to reframe it in digestible way. Which is exact pattern that appear on less wrong posts regularly. It is not bad person convincing himself he is good person—it is you looking in the mirror and not recognising yourself. Why you think ASI will help you if you think you already good and dont need help?Here’s a reflection on this debate process:
cousin_it and I disagreed in comment sections a couple of times so decided to make that discussion into a post.
As we iterated drafts, I realized that we were making separate points more than disagreeing (although we do still disagree on my central point, how risky it would be to have one person control obedient singleton ASI).
cousin_it was focused on the point “for god’s sake let’s not put the elite in charge of transformative AI; they’ll mistreat everyone” and I was making the point “for god’s sake let’s not give everyone their own takeover-capable AI; one person in charge would be less risky than that”.
Those two points are compatible. We should do neither. I think we agree that it’s better to just not build ASI, or to build value-aligned ASI. If either of those is possible.
The reason I think the discussion is worthwhile even when we’re in agreement that we should do neither, is that either moderately distributed or very concentrated power over obedient ASI seems like the most likely outcome on the current path (if we can even pull off instruction-following alignment successfully). So deciding between those two bad options may be all we get to do.
I’m not even sure I disagree with cousin_it on how powerful AI should get before we restrict public access and prevent proliferation. I certainly don’t want to shut that down for open-source models, even next-gen ones capable of truly dangerous hacking and bioengineering. When proliferation becomes too dangerous is going to be a judgment call, and I won’t be the one to make it—those leading the AI race will. (I think this will probably be the sitting POTUS, and that’s the biggest lever we’re likely to get).
Really? doesn’t public availability of open-weight superhuman bioengineering basically guarantee megadeaths? what’s the worthwhile trade-off here?
I don’t want to see that happen and I don’t want to be one of the casualties; but I want even more for us to make it to aligned ASI and a staggeringly good future. To me the situation looks pretty dire and we’re going to have to take some chances. The trick is calculating our risks as best we can in advance.
The tradeoff I primarily see is having warning shots big enough to make us take this shit seriously but not big enough to kill us all. Of course, the wrong warning shots like a bioengineered plauge created under AI orders could hurt, making us rush to ASI and get egregiously misaligned ASI… which kills us all.
We need to work faster and harder to analyze those risks, because the decisions may be coming up sooner than we’d like.
I’m on board, but open-weight models with superhuman biohacking skills don’t seem like a chance we have to take? Maybe I missing some second order effect that makes fighting for open-weight models in particular worth the inevitable death toll?
Well I’d rather they not have superhuman biohacking skills! We don’t seem on track to be able to preserve some capabilities and prevent others in open-source models. I guess it’s possible to just eliminate huge chunks of the pretraining data, and I sure hope people do that once we’re hitting truly dangerous capabilities (we’ll get our warning shots from whatever shenanigans aren’t totally prevented). Closed-source models are looking pretty easy to monitor and control through wrappers like Fable’s but those won’t allow any resistance to those controlling the companies (which will btw probably be the government).
Right, so we have to outlaw them. or have Mythos hack and poison them, or something. but no, because … they are crucial to friendly AGI? (are they?)
or is this just the “we can’t install a traffic light until someone actually dies” thing?
Concentration of power is worse.
Three cheers for good-faith, high-effort debate, distilled for easy consumption.
I think the explicit target of the debate is hard to think about, because it depends on lots of (ideally) separate, complex issues, like takeoff speeds/timelines, local vs. global intelligence explosions, ambiguity about what kind of AI capabilities each of the debate participants are imagining everyone having or not at any given time, ratio of offensive to defensive capability for every domain, and the effects of power and/or wealth inequality on some kind of measure of aggregate welfare.
So I think a good potential crux/problem relaxation of the target belief is, “Do you support any kind of wealth tax at all?”
I quite strongly predict that Vladimir would generally support some kind of wealth tax. I also think he thinks that, if people were as nice and reflective in good environmental conditions as Seth seems to believe they are, then there would already be significant wealth taxes in effect, say, in the United States, right now.
Based on my observations of people holding positions similar to Seth’s on these issues, I weakly predict that Seth would prefer no wealth tax, or that, if he preferred some wealth tax, then probably Vladimir would consider it so conservative as to be ineffectual, but honestly I’m not really sure. Seth himself admits that he is pretty uncertain about a lot of stuff!
I really don’t think this reduces to a question about wealth taxes. I am talking about a post-singularity and post-scarcity society. Cousin_it is, I believe, addressing more of a transitional period where a wealth tax would be highly relevant. For the record, I very much endorse a wealth tax of some sort during that type of transitional period—at least until we have adequate analysis of the alternatives, which I have yet to see.
But I don’t think that has a close connection to the questions I have in mind, which are: 1. Whether we should favor wide proliferation or tight restriction of cutting-edge AI as it nears the threshold of being takeover-capable or autonomous RSI-capable. 2. Whether we should assume that obedient AI is too dangerous in either a few or many hands to treat as a viable alignment target, even if it is easier and more likely than value aligned AGI.
I think that, just like “takeoff-capable” AI is a wrong concept because technology advances everywhere and makes takeoff a moving target, “post-scarcity” is also a wrong concept because competition prevents post-scarcity. Competition can burn arbitrary amounts of resources, and if you unilaterally refuse to burn them, you just get taken over. Maybe we could have post-scarcity in a world with only one powerful benevolent agent giving handouts to everyone else, but that’s what the debate was about :-)
About wealth taxes, I think the point is pretty relevant: if people were good enough by nature, there wouldn’t be so much easily preventable homelessness today and so on. To me, history gives an upper bound on how “good by nature” we can consider people to be. That bound is quite low and there’s no reason to think it’ll be higher. But again this was covered in the debate.