“I spend comparatively little time thinking about whether the most likely principal in a superpersuasion-assisted takeover is an AI system, an AI lab, a government, a coalition of the above, or something else”
—Interested in more on why you’re prioritising in this way
”I don’t love the term superpersuasion. However after a small brainstorm to think of better terms, I didn’t have any terms that I and/or other people actively preferred sufficiently”
—I prefer just ‘AI persuasion’. It doesn’t distinguish between different levels (e.g. includes today’s systems), but then you can just say ‘I think risks from AI persuasion will get bigger as capabilities improve’, which seems fine to me.
rosehadshar
Thanks for the comment Michael.
A minor quibble is that I think it’s not clear you need ASI to end up with dangerous levels of power concentration, so you might need to ban AGI, and to do that you might need to ban AI development pretty soon.
I’ve been meaning to read your post though, so will do that soon.
Thanks, I think these are interesting points.
I agree that some power concentration is likely necessary, and that it could be a lot, though I’m pretty unsure there.
In terms of what to do about that:
Do you have ideas about what widening political control over a small number of AIs would look like?
Another alternative to crossing our fingers would be to distribute strategic power but also build AI resilience. This doesn’t work if there are big capability gaps I think, but if capabilities are relatively evenly distributed then there might be ways to build defensive enclaves
Yup sorry, the Tom above is actually Rose!
I like your distinction between narrow and broad IC dynamics. I was basically thinking just about narrow, but now agree that there is also potentially some broader thing.
How likely do you think it is that helper-nanobots outcompete auto-nanobots? Two possible things that could be going on are:
I’m unhelpfully abstracting away what kind of AI systems we end up with, but actually this significantly impacts how likely power concentration is and I should think more about it
Theoretically the distinction between helper and auto-nanobots is significant, but in practice it’s very unlikely that helper-nanobots will be competitive, so it’s fine to ignore the possibility and treat auto-nanobots as ‘AI’
All of 1-4 seem plausible to me, and I don’t centrally expect that power concentration will lead to everyone dying.
Even if all of 1-4 hold, I think the future will probably be a lot less good than it could have been:
- 4 is more likely to mean that earth becomes a nature reserve for humans or something, than that the stars are equitably allocated- I’m worried that there are bad selection effects such that 3 already screens out some kinds of altruists (e.g. ones who aren’t willing to strategy steal). Some good stuff might still happen to existing humans, but the future will miss out on some values completely
- I’m worried about power corrupting/there being no checks and balances/there being no incentives to keep doing good stuff for others
Thanks, agree that ‘emergent dynamics’ is woolly above.
I guess I don’t think the y-axis should be the temporal dimension. To give some cartoon examples:
I’d put an extremely Machievellian 10 year plan on the part of a cabal of politicians to backslide into a dictatorship then seize power over the rest of the world near the top end of the axis
I’d put unfavourable order of capabilities, where in an unplanned way superpersuasion comes online before defenses, and actors fail to coordinate not to deploy it because of competitive dynamics, near the bottom end of the axis. Even if the whole thing unfolds over a few months
I do think the y-axis is pretty correlated with temporal scales, but I don’t think it’s the same. I also don’t think physical violence is the same, though it’s probably also correlated (cf the backsliding example which is v powerseeking but not v violent).
The thing I had in mind was more like, should I imagine some actor consciously trying to bring power concentration about? To the extent that’s a good model, it’s power-seeking. Or should I imagine that no actor is consciously planning this, but the net result of the system is still extreme power concentration? If that’s a good model, it’s emergent dynamics.Idk, I see that this is messy and probably there’s some other better concept here
I think I agree that, once an AI-enabled coup has happened, the expected remaining AI takeover risk would be much lower. This is partly because it ends the race within the country where the takeover happened (though it wouldn’t necessarily end the international race), but also partly just because of the evidential update: apparently AI is now capable of taking over countries, and apparently someone could instruct the AIs to do that, and the AIs handed the power right back to that person! Seems like alignment is working.
I don’t currently agree that the remaining AI takeover risk would be much lower:
The international race seems like a big deal. Ending the domestic race is good, but I’d still expect reckless competition I think. Maybe you’re imagining that a large chunk of powergrabs are motivated by stopping the race? I’m a bit sceptical.
I don’t think the evidential update is that strong. If misaligned AI found it convenient to take over the US using humans, why should we expect them to immediately cease to find humans useful at that point? They might keep using humans as they accumulate more power, up until some later point.
There’s another evidential update which I think is much stronger, which is that the world has completely dropped the ball on an important thing almost no one wants (powergrabs), where there are tractable things they could have done, and some of those things would directly reduce AI takeover risk (infosec, alignment audits etc). In a world where coups over the US are possible, I expect we’ve failed to do basic alignment stuff too.
Curious what you think.
I think it might be a bit clearer to communicate the stages by naming them based on the main vector of improvement throughout the entire stage, i.e. ‘optimization of labor’ for stage one, ‘automation of labor’ for stage two, ‘miniturization’ for stage three.
I think these names are better names for the underlying dynamics, at least—thanks for suggesting them. (I’m less clear they are better labels for the stages overall, as they are a bit more abstract.)
Changed to motivation, thanks for the suggestion.
I agree that centralising to make AI safe would make a difference. It seems a lot less likely to me than centralising to beat China (there’s already loads of beat China rhetoric, and it doesn’t seem very likely to go away).
“it is potentially a lot easier to stop a single project than to stop many projects simultaneously” → agree.
I think I still believe the thing we initially wrote:
Agree with you that there might be strong incentives to sell stuff at monopoloy prices (and I’m worried about this). But if there’s a big gap, you can do this without selling your most advanced models. (You sell access to weaker models for a big mark up, and keep the most advanced ones to yourselves to help you further entrench your monopoly/your edge over any and all other actors.)
I’m sceptical of worlds where 5 similarly advanced AGI projects don’t bother to sell
Presumably any one of those could defect at any time and sell at a decent price. Why doesn’t this happen?
Eventually they need to start making revenue, right? They can’t just exist on investment forever
(I am also not an economist though and interested in pushback.)
Thanks, I expect you’re right that there’s some confusion in my thinking here.
Haven’t got to the bottom of it yet, but on more incentive to steal the weights:
- partly I’m reasoning in the way that you guess, more resources → more capabilities → more incentives
- I’m also thinking “stronger signal that the US is all in and thinks this is really important → raises p(China should also be all in) from a Chinese perspective → more likely China invests hard in stealing the weights”
- these aren’t independent lines of reasoning, as the stronger signal is sent by spending more resources
- but I tentatively think that it’s not the case that at a fixed capability level the incentives to steal the weights are the same. I think they’d be higher with a centralised project, as conditional on a centralised project there’s more reason for China to believe a) AGI is the one thing that matters, b) the US is out to dominate
Thanks, I agree this is an important argument.
Two counterpoints:
The more projects you have, the more attempts at alignment you have. It’s not obvious to me that more draws are net bad, at least at the margin of 1 to 2 or 3.
I’m more worried about the harms from a misaligned singleton than from a misaligned (or multiple misaligned) systems in a wider ecosystem which includes powerful aligned systems.
Thanks! Fwiw I agree with Zvi on “At a minimum, let’s not fire off a starting gun to a race that we might well not win, even if all of humanity wasn’t very likely to lose it, over a ‘missile gap’ style lie that we are somehow not currently in the lead.”
Thanks for these questions!
Earlier attacks: My thinking here is that centralisation might a) cause China to get serious about stealing the weights sooner, and b) therefore allow less time for building up great infosec. So it would be overall bad for infosec. (It’s true the models would be weaker, so stealing the weights earlier might not matter so much. But I don’t feel very confident that strong infosec would be in place before the models are dangerous (with or without centralisation))
More attack surface: I am trying to compare multiple projects with a single project. The attack surface of a single project might be bigger if the single project itself is very large. As a toy example, imagine 3 labs with 100 employees each. But then USG centralises everything to beat China and pours loads more resources into AGI development. The centralised project has 1000 staff; the counterfactual was 300 staff spread across 3 projects.
China stealing weights: sorry, I agree that it’s harder for everyone including China, and that all else equal this disincentivises stealing the weights. But a) China is more competent than other actors, so for a fixed increase in difficulty China will be less disincentivised than other actors, b) China has bigger incentives to steal the weights to begin with, and c) for China in particular there might be incentives that push the other way (centralising could increase race dynamics between the US and China, and potentially reduce China’s chances of developing AGI first without stealing the weights), and those might counteract the disincentive. Does that make more sense?
My main take here is that it seems really unlikely that the US and China would agree to work together on this.
That seems overconfident to me, but I hope you’re right!
To be clear:
- I agree that it’s obviously a huge natsec opportunity and risk.
- I agree the USG will be involved and that things other than nationalization are more likely
- I am not confident that there will be consensus across the US on things like ‘AGI could lead to an intelligence explosion’, ‘an intelligence explosion could lead to a single actor taking over the world’, ‘a single actor taking over the world would be bad’.
Thanks!
I think I don’t follow everything you’re saying in this comment; sorry. A few things:
- We do have lower p(AI takeover) than lots of folks—and higher than lots of other folks. But I think even if your p(AI takeover) is much higher, it’s unclear that centralisation is good, for some of the reasons we give in the post:
-- race dynamics with China might get worse and increase AI takeover risk
—racing between western projects might not be a big deal in comparison, because of races to the top and being more easily able to regulate
- I’m not trying to assume that China couldn’t catch up to the US. I think it’s plausible that China could do this in either world via stealing the model weights, or if timelines are long. Maybe it could also catch up without those things if it put its whole industrial might behind the problem (which it might be motivated to do in the wake of US centralisation).
- I think whether a human dictatorship is better or worse than an AI dictatorship isn’t obvious (and that some dictatorships could be worse than extinction)
On the infosec thing:
”I simply don’t buy that the infosec for multiple such projects will be anywhere near the infosec of a single project because the overall security ends up being that of the weakest link.”
-> nitpick: the important thing isn’t how close the infosec for multiple projects is to the infosec of a single project: it’s how close the infosec for multiple projects is to something like ‘the threshold for good enough infosec, given risk levels and risk tolerance’. That’s obviously very non-trivial to work out
-> I agree that a single project would probably have higher infosec than multiple projects (though this doesn’t seem slam dunk to me and I think it does to you)
-> concretely, I currently expect that the USG would be able to provide SL4 and maybe SL5 level infosec to 2-5 projects, not just one. Why do you think this isn’t the case?“Additionally, the more projects there are with a particular capability, the more folk there are who can leak information either by talking or by being spies.”
-> It’s not clear to me that a single project would have fewer total people: seems likely that if US AGI development is centralised, it’s part of a big beat China push, and involves throwing a lot of money and people at the problem.
Here are some notes on checks and balances and what it might look like to preserve them into the AI era.
First, why even have checks and balances? What are they for?
A few possible answers:
To prevent any individual actor getting too much power
This is the power concentration version
To prevent unchecked power
I think this is strictly more accurate than the first answer (as in, it more accurately tracks what I care about), but it currently feels a bit abstract to me
To preserve/enhance something like kindness, integrity, alignment in the overall system
Here I’m trying to rhyme with Richard Ngo’s well-foundedness and Jan Kulveit’s kind Leviathan
Maybe there’s some negatively framed dual to this which is like ‘no exploitation’ (unsure)
Of those, the one I like most is preserving kindness, integrity, alignment in the overall system.
So, if I take preserving kindness as the purpose of checks and balances, where does that take me when I start to think about how to preserve checks and balances into the AI era?
Here’s a first pass:
Checks and balances need to do two things:
Preserve horizontal kindness between agents at the same level (for example, kindness between me and another human, or between nation states)
Preserve vertical kindness between agents at different levels (for example, kindness between a government and its citizens, or an ASI and a human)
Preserving horizontal kindness requires two things:
Rules: some specification of what isn’t ok
Enforcement mechanisms: some way of preventing actors from breaking the rules
Some things I notice about this
Rules can be explicit, like the law, or implicit, like one’s moral compass
Rules can be enforced top down, like the state policing criminals, or enforced through coordination, like a village ostracising someone who beats their wife
Preserving vertical kindness requires two things:
Good information transfer between levels, so that they understand one another
Some alignment mechanism
One way of doing this is making sure that they are mutually dependent/each have hard leverage over the other
Rules and enforcement is another way of doing this
[in some ways good information transfer is an alignment mechanism]
[presumably other options]
A little bit more thinking, this time in a table:
Function
What this looks like today
What this could look like in the AI era
Preserving horizontal kindness
Rules
Law
Norms
Law (updated)
Whatever rules AI systems are trained and can be verified to follow, and the rules that govern how those rules can change
[I expect norms will matter less]
Enforcement mechanisms
Locally, social fabric
Domestically, the police
Internationally and only partially, the military
Self-enforcing mechanisms (like the rules AI systems are trained and verified to follow)
ASI nightwatchman
I guess MAIM is maybe something like this?
Structured transparency style mass surveillance and enforcer drones who stop people building WMDs
Preserving vertical kindness
Information transfer
Managers and reports talk to each other. Also reports sometimes fill in feedback forms
Citizens vote, and sometimes write to or speak with their representatives
There are polls
…
Structured transparency could enable much higher bandwidth in both directions
AI delegates
More ambient mediation between humans and their environment by AI systems (e.g. there’s a secure system which takes in data from all of the systems I use and knows everything about me, and this data can only be used for certain purposes/not in raw form, but gives really accurate and deep information)
Really good simulation?
Alignment
Aligning government to the people:
Constitutions and other law
Elections
Citizen taxation
Aligning the people to government:
Education
The law
…
Giving people capital and/or compute
Self-enforcing rules that restrict what the higher vertical levels can do
Things that stand out to me from the table:
Novel threats
Accelerating progress could lead competitive dynamics to erode all value. Progress hasn’t been fast enough for this before
Enforcement could potentially be perfect. This would be bad in our current world I think; unclear if it will be good or bad in a future one
Tech could make n of one violations extremely costly, which makes perfect enforcement desirable
But perfect enforcement also significantly raises the bar for how good the rules need to be, and the constitutional process which sets and amends them
Novel opportunities
Sufficiently powerful information transfer and enforcement mechanisms to prevent war, which hasn’t been possible before
Verifiably self-enforcing rules is a v powerful new affordance