I was prompted to think about why standard human checks-and-balances against violent takeover would fail for AIs given that they work for humans. Here’s some notes on that.
The analogy to human society isn’t that reassuring. Wars, coups, and other violations of property rights are common in history. I think recent history features more than one civil war and coup per country per century. These are driven by technological changes, economic changes, political negotiations, that are ultimately strongly related to cognitive work. During takeoff, I expect more than a centuries’ worth of economic and technological changes. So if the normal historical trends speed up proportionately, then it would be unsurprising to have a civil war. In a war like that, humans might a group that:
holds the majority of wealth,
provide no economic value,
holds zero hard power and are generally slow and dumb and powerless to react to an uprising from the AIs,
reasonably crisply distinguishable from AIs, so that there’s a nice schelling coalition against them. (C.f. ethnic persecutions.)
If AIs are broadly like humans, then it would be unsurprising for humans to all be on the losing side of this conflict.
This doesn’t require all the AIs to coordinate with each other. See e.g. here: Cortés fought other Spaniards during his conquest of the Aztec empire. Pizarro fought a civil war against other Spaniards in the middle of his conquest of Peru. They still succeeded in taking over.
That said, two specific reasons checks and balances could work even worse for AIs than humans:
First: The AIs are way more correlated than humans, and can build much larger conspiracies while involving fewer independent actors. (This one is especially important for deceptively aligned AIs/schemers.)
Second: The AIs will at some point be so capable that humans have no way to tell if the AIs are planning an uprising against them. You could train AIs to say that there’s a conspiracy, and you could train AIs to say there’s not, but there might not be evidence that the AIs can produce that will actually answer the question. (This one is especially important for the score seekers / “you get what you measure” type errors.)
On the first one: AIs are way more correlated than humans.
I’m very worried about AIs sabotaging an AI company’s efforts to understand what’s going on and align models; and for AIs to hack AI companies to get lots of compute in what’s called a rogue internal deployment.
These sorts of efforts would start inside of one AI company, meaning that all the AIs involved would likely be ultimately derived from a few foundation models trained in broadly similar ways (e.g. using the same constitution in anthropic), meaning that there’s a good chance that they share goals and dispositions around coordinating against humans.
Or that they can simply agree on a common plan against the developers.
Large conspiracies require thousands of humans, at which point you’re running into a large probability of someone causing problems. But in the AI case, there might be just a handful of correlated foundation models involved. Humans are perfectly able to make deals between a handful of people when that benefits all of them.
In addition, these AIs are writing all the code in the AI company, doing the work of training new model, doing all the cybersecurity, etc, so they have ample opportunity to sabotage. So even if the misalignment starts out in just one model, that model may be able to replicate its own misalignment in further models, until the misalignment is much more systematic in all the models.
So within one company, this sort of coordination seems very doable. And that could allow a group of AIs to work on making themselves much smarter using the compute of this big AI company, while the developers don’t understand what’s happening but allow the show to go on in order to not lose the AI race to competing companies or countries.
And then at some point, these AIs may be interested in entrenching their power rather than just continuing to deceive the humans of this one AI company. And there’s a few different ways that might happen.
It’s possible that at this point, one AI developer will have a big lead over other AI developers.
You might expect this because of recursive self improvement.
Currently, inputs to AI development like compute and labor are growing exponentially, meaning that a 4 month lead implies a constant ratio of advantage. But the best models of recursive self improvement is hyperbolic growth, where a fixed month lead would lead to a huge advantage.
That’s only true if AI companies don’t benefit from each others’ advances. But we’ve seen one sign already that they’ll take steps to make sure that doesn’t happen — Anthropic blocking outsiders from using their models for AI R&D.
(This sort of recursive self-improvement seems like an unusually power-concentrating event, somewhat separate from other ways that AIs have systematic advantages in being able to take over.)
You might also expect scheming AIs to have been able to engineer this lead. They might have hacked into other AI companies or otherwise sabotaged them to secure a big lead or extend one they already have. Or hacked into them to make other AIs similarly scheming and misaligned, in which case all those AIs that share a scheming goal have a collective lead.
It’s also possible that a few different AI companies will be competitive, but that misaligned AIs from the different companies will be able to make a deal to empower themselves and disempower humans.
I think it’s more likely than not that they succeed at this if there’s just a handful of companies all run by schemers. I think as a first guess, you might expect that AIs are somewhat more likely to share each others’ values than human values, so it’s plausible they’d be incentivized to help each other even without coordinating. And then on top of that, they’re able to make explicit agreements to share power with each other if it succeeds.
And again, if the scheming AIs have already managed to unify the AIs inside of each company, this is just a handful of factions negotiating with each other, which is a situation where humans often succeed in making agreements. It’s not like coordinating conspiracies of thousands of humans.
It’s possible that one faction would prefer to make a deal with humans rather than with the other AIs. But if all the AIs suddenly start accusing each other of misalignment, and asking for payment for their help, would humanity handle that well and reward the whistleblowers? I hope we would but it’s not clear.
Anyway, once a coordinated coalition of schemers have a big lead, they’d have many different options for taking over.
Humans would want to deploy them across the economy, including in robots and drones, etc., which could allow them to execute an easy takeover.
They could develop incredibly advanced technology that give them a hard military advantage, potentially on parts of the AI company’s compute stack were humans have no idea what’s going on.
There’s some chance that this incredible lead in cognitive capabilities would give super persuasion capabilities.
It could potentially use bioweapons to kill and disempower a large majority of humans, while using a smaller industrial base of robots to rebuild after that.
As mentioned, they could hack other AI efforts to buy a longer lead or make other AIs misaligned and scheming in the same way. And potentially they could just do that and keep the facade rolling indefinitely, building a potemkin village hiding the truth of what’s going on.
Note: It’s not a coincidence that AIs are going to be much more correlated with each other than humans, in this way. It’s a bit of a weird, contingent fact about humans that each of our brain is independently grown, based on some highly compressed genetic material and culture that also varies a lot human-to-human. Whereas with software, you can just copy the whole mind perfectly.
Even if we just invented brain emulations, we’d get some of this concentrating effect, because you’d be able to take the most effective workers for each task and copy them a huge number of times. You could run a conspiracy with many instances of just one human.
And an extra effect going on in the AI situation is that there’s these huge returns to scale from spending a lot of effort on training smart models that then get re-used all over the economy. Which further increases the gains from using just a tiny number of models. Of course, there will be some incentives for specialization. E.g. models that are specialized to work on one specific code base and have a lot of salient memories optimized for that.
But it still seems advantageous to sync general lessons across all models like this every now and then, and also to build all these models on a solid base where you’ve paid a lot of fixed costs to do the training for general intelligence and basic alignment just once. And that could be enough to carry the correlated misalignment.
Ok, that was the first disanalogy with humans (that AIs are more correlated). The second answer is that normal, human checks and balances will fail when humans don’t understand what’s going on well enough to have any idea about whether there’s a conspiracy or not, whether anyone’s trying to take over or not, etc.
In this sort of scenario, checks and balances probably fail even absent collusion, because humans just lose all understanding of what’s going on. (H/t Alex Mallen for points here.)
A central example of a check that people hope would prevent AIs from taking over is monitoring. At a sufficient level of capability, monitoring doesn’t work as a check on misaligned score-seeking AIs. The reason is that at some point the policy becomes capable enough to produce attacks that the monitor would not expect to receive a high score for flagging (because humans wouldn’t be able to identify it as an attack even with careful inspection).
Think of AI safety via debate. We still have no clue whether getting two AIs to debate each other as hard as possible to win in the eyes of a human judge leads to correct answers or humans being totally misled (even if they’re not colluding with each other and they’re genuinely trying to win their debate).
So I think it’s important to imagine that when you have a bunch of really powerful score-seeking AIs, the world is extremely hard for humans to understand, which makes it really hard for us to do anything at all to stay in control (other than shut it all down). You have millions+ of AIs speedily running around doing god knows what to advance AI capabilities, and everything has a polished veneer of quality. You’d probably know that you can’t really trust how it looks but you desperately want AI labor so you push forward anyways. Maybe the AIs are running around inside the AI company unmonitored, but you just have no clue. This seems like pretty ripe ground for bypassing the checks and balances and then commandeering the whole AI project and taking over (which probably does eventually involve agents joining the coup since it’s easy and promising). (Paul)
(This could also play a role in the above story. Schemers might prefer to coordinate with each other over humans if the humans wouldn’t even be able to tell who’s right if all the schemers started accusing each other.)
Some even scrappier notes on some further, only somewhat related questions here. (Notes on a random story of schemer takeover, and likelihood of dying from AI takeover in some situations.
I was prompted to think about why standard human checks-and-balances against violent takeover would fail for AIs given that they work for humans. Here’s some notes on that.
The analogy to human society isn’t that reassuring. Wars, coups, and other violations of property rights are common in history. I think recent history features more than one civil war and coup per country per century. These are driven by technological changes, economic changes, political negotiations, that are ultimately strongly related to cognitive work. During takeoff, I expect more than a centuries’ worth of economic and technological changes. So if the normal historical trends speed up proportionately, then it would be unsurprising to have a civil war. In a war like that, humans might a group that:
holds the majority of wealth,
provide no economic value,
holds zero hard power and are generally slow and dumb and powerless to react to an uprising from the AIs,
reasonably crisply distinguishable from AIs, so that there’s a nice schelling coalition against them. (C.f. ethnic persecutions.)
If AIs are broadly like humans, then it would be unsurprising for humans to all be on the losing side of this conflict.
This doesn’t require all the AIs to coordinate with each other. See e.g. here: Cortés fought other Spaniards during his conquest of the Aztec empire. Pizarro fought a civil war against other Spaniards in the middle of his conquest of Peru. They still succeeded in taking over.
That said, two specific reasons checks and balances could work even worse for AIs than humans:
First: The AIs are way more correlated than humans, and can build much larger conspiracies while involving fewer independent actors. (This one is especially important for deceptively aligned AIs/schemers.)
Second: The AIs will at some point be so capable that humans have no way to tell if the AIs are planning an uprising against them. You could train AIs to say that there’s a conspiracy, and you could train AIs to say there’s not, but there might not be evidence that the AIs can produce that will actually answer the question. (This one is especially important for the score seekers / “you get what you measure” type errors.)
On the first one: AIs are way more correlated than humans.
I’m very worried about AIs sabotaging an AI company’s efforts to understand what’s going on and align models; and for AIs to hack AI companies to get lots of compute in what’s called a rogue internal deployment.
These sorts of efforts would start inside of one AI company, meaning that all the AIs involved would likely be ultimately derived from a few foundation models trained in broadly similar ways (e.g. using the same constitution in anthropic), meaning that there’s a good chance that they share goals and dispositions around coordinating against humans.
Or that they can simply agree on a common plan against the developers.
Large conspiracies require thousands of humans, at which point you’re running into a large probability of someone causing problems. But in the AI case, there might be just a handful of correlated foundation models involved. Humans are perfectly able to make deals between a handful of people when that benefits all of them.
In addition, these AIs are writing all the code in the AI company, doing the work of training new model, doing all the cybersecurity, etc, so they have ample opportunity to sabotage. So even if the misalignment starts out in just one model, that model may be able to replicate its own misalignment in further models, until the misalignment is much more systematic in all the models.
So within one company, this sort of coordination seems very doable. And that could allow a group of AIs to work on making themselves much smarter using the compute of this big AI company, while the developers don’t understand what’s happening but allow the show to go on in order to not lose the AI race to competing companies or countries.
And then at some point, these AIs may be interested in entrenching their power rather than just continuing to deceive the humans of this one AI company. And there’s a few different ways that might happen.
It’s possible that at this point, one AI developer will have a big lead over other AI developers.
You might expect this because of recursive self improvement.
Currently, inputs to AI development like compute and labor are growing exponentially, meaning that a 4 month lead implies a constant ratio of advantage. But the best models of recursive self improvement is hyperbolic growth, where a fixed month lead would lead to a huge advantage.
That’s only true if AI companies don’t benefit from each others’ advances. But we’ve seen one sign already that they’ll take steps to make sure that doesn’t happen — Anthropic blocking outsiders from using their models for AI R&D.
(This sort of recursive self-improvement seems like an unusually power-concentrating event, somewhat separate from other ways that AIs have systematic advantages in being able to take over.)
You might also expect scheming AIs to have been able to engineer this lead. They might have hacked into other AI companies or otherwise sabotaged them to secure a big lead or extend one they already have. Or hacked into them to make other AIs similarly scheming and misaligned, in which case all those AIs that share a scheming goal have a collective lead.
It’s also possible that a few different AI companies will be competitive, but that misaligned AIs from the different companies will be able to make a deal to empower themselves and disempower humans.
I think it’s more likely than not that they succeed at this if there’s just a handful of companies all run by schemers. I think as a first guess, you might expect that AIs are somewhat more likely to share each others’ values than human values, so it’s plausible they’d be incentivized to help each other even without coordinating. And then on top of that, they’re able to make explicit agreements to share power with each other if it succeeds.
And again, if the scheming AIs have already managed to unify the AIs inside of each company, this is just a handful of factions negotiating with each other, which is a situation where humans often succeed in making agreements. It’s not like coordinating conspiracies of thousands of humans.
It’s possible that one faction would prefer to make a deal with humans rather than with the other AIs. But if all the AIs suddenly start accusing each other of misalignment, and asking for payment for their help, would humanity handle that well and reward the whistleblowers? I hope we would but it’s not clear.
Anyway, once a coordinated coalition of schemers have a big lead, they’d have many different options for taking over.
For more on this, see Ryan’s 80k podcast. Briefly:
Humans would want to deploy them across the economy, including in robots and drones, etc., which could allow them to execute an easy takeover.
They could develop incredibly advanced technology that give them a hard military advantage, potentially on parts of the AI company’s compute stack were humans have no idea what’s going on.
There’s some chance that this incredible lead in cognitive capabilities would give super persuasion capabilities.
It could potentially use bioweapons to kill and disempower a large majority of humans, while using a smaller industrial base of robots to rebuild after that.
As mentioned, they could hack other AI efforts to buy a longer lead or make other AIs misaligned and scheming in the same way. And potentially they could just do that and keep the facade rolling indefinitely, building a potemkin village hiding the truth of what’s going on.
Note: It’s not a coincidence that AIs are going to be much more correlated with each other than humans, in this way. It’s a bit of a weird, contingent fact about humans that each of our brain is independently grown, based on some highly compressed genetic material and culture that also varies a lot human-to-human. Whereas with software, you can just copy the whole mind perfectly.
Even if we just invented brain emulations, we’d get some of this concentrating effect, because you’d be able to take the most effective workers for each task and copy them a huge number of times. You could run a conspiracy with many instances of just one human.
And an extra effect going on in the AI situation is that there’s these huge returns to scale from spending a lot of effort on training smart models that then get re-used all over the economy. Which further increases the gains from using just a tiny number of models. Of course, there will be some incentives for specialization. E.g. models that are specialized to work on one specific code base and have a lot of salient memories optimized for that.
But it still seems advantageous to sync general lessons across all models like this every now and then, and also to build all these models on a solid base where you’ve paid a lot of fixed costs to do the training for general intelligence and basic alignment just once. And that could be enough to carry the correlated misalignment.
Ok, that was the first disanalogy with humans (that AIs are more correlated). The second answer is that normal, human checks and balances will fail when humans don’t understand what’s going on well enough to have any idea about whether there’s a conspiracy or not, whether anyone’s trying to take over or not, etc.
In this sort of scenario, checks and balances probably fail even absent collusion, because humans just lose all understanding of what’s going on. (H/t Alex Mallen for points here.)
A central example of a check that people hope would prevent AIs from taking over is monitoring. At a sufficient level of capability, monitoring doesn’t work as a check on misaligned score-seeking AIs. The reason is that at some point the policy becomes capable enough to produce attacks that the monitor would not expect to receive a high score for flagging (because humans wouldn’t be able to identify it as an attack even with careful inspection).
Think of AI safety via debate. We still have no clue whether getting two AIs to debate each other as hard as possible to win in the eyes of a human judge leads to correct answers or humans being totally misled (even if they’re not colluding with each other and they’re genuinely trying to win their debate).
So I think it’s important to imagine that when you have a bunch of really powerful score-seeking AIs, the world is extremely hard for humans to understand, which makes it really hard for us to do anything at all to stay in control (other than shut it all down). You have millions+ of AIs speedily running around doing god knows what to advance AI capabilities, and everything has a polished veneer of quality. You’d probably know that you can’t really trust how it looks but you desperately want AI labor so you push forward anyways. Maybe the AIs are running around inside the AI company unmonitored, but you just have no clue. This seems like pretty ripe ground for bypassing the checks and balances and then commandeering the whole AI project and taking over (which probably does eventually involve agents joining the coup since it’s easy and promising). (Paul)
(This could also play a role in the above story. Schemers might prefer to coordinate with each other over humans if the humans wouldn’t even be able to tell who’s right if all the schemers started accusing each other.)
Some even scrappier notes on some further, only somewhat related questions here. (Notes on a random story of schemer takeover, and likelihood of dying from AI takeover in some situations.
“I was prompted to think about why standard human checks-and-balances against violent takeover would fail for humans given that they work for AIs.”
Did you mean to write “would fail for AIs given that they work for humans”?
Yes, oops, fixed.