I agree that we should think more about this. We’re working on a literature review. We’ve also written down a few options here: https://takeoverbench.com/threats
otto.barten
55% of the US public is now aware of AI xrisk
I agree this is the thing to worry about. I’m somewhat skeptical that takeover would be a lot better for comms than extinction. We attempted to make this a bit more fashionable here:
I’m pretty sure this is wrong. In addition, the more deeply people think about what this would mean, the more hostile I’d expect them to become. I’d expect this to be true even for an enlightened AI overlord, let alone for the worst case AI overlord.
34% of the US public is now aware of AI xrisk, and the curve is steepening
‘Being nice to AI’ is probably not a good survival strategy, since from the AI’s perspective, it’s too unreliable. But verifiably (by the AI) comply with the AI’s rules could be a viable plan B potentially.
We will end up as ants, but we won’t necessarily start out as ants, which reduces our chances of being successful pets.
Thanks for the post.
The way I look at this is more that humans, and societies, are self-improving machines that learn by trail and error. I think things can go wrong on humanity scale if new technology gets incorrigible. AI takeover is an example of this, as is for example biodiversity loss (which I expect to be incorrigible even with max tech, but I hope to be wrong).
As long as we maintain the ability to make errors and live through them (that is, as long as our errors are small and slow enough to be corrected in a trail and error fashion), I think any required self-correction and/or self-reflection should happen by default.
It would help if we would get better at self correcting (faster, more reliably eg ending up at the right response), although that’s not strictly needed for a good long-term outcome. It would mostly limit the casualties of corrections.
I see our job as making sure no incorrigible things happen (such as AI takeovers). Maybe that’s a much easier target than requiring major human improvement?
Maybe We’ll Get a State Takeover Before a World Takeover
Is a pause enforceable? New paper out!
Open internship position + call for collaborations on threat model-dependent alignment, governance, and offense/defense balance
Argument against recursive self-improvement: you need algorithms, compute, and data for AI. Self-improvement only works on the algorithms.
Maybe self-improvement works but only up until a ceiling determined by compute and data that may be << superintelligence.
I realized that but I think my counterarguments are true for most organizations who would have a realistic chance of building takeover-level AI. As a case in point, Anthropic launched the Glasswing project trying to fix vulnerabilities rather than saying “great now we can hack into banks for our takeover attempt planned in Q3 2027, let’s not tell anyone about these zero days”.
In your scenario, DARPA would probably not try to take over using AI. It’s not their culture. When they realize their AI achieved DSA, they’d likely voluntarily hand over control to the US government.
Also, DARPA or any other AI project does not operate in a vacuum. Likely, people in government would realize DARPA is on its way to DSA and intervene to take over control before it happens.
We already have a situation where we have democracy with a powerful weapon, namely an atomic bomb. I don’t really see how this will necessarily be different. Oppenheimer also didn’t take over.
As a non-American, I’m worried however that this will decrease power in non-American territories even more (higher permanence and granularity of US sovereignty). Remember that >95% of the world population is not American. All these people would have no say over whatever happens afterwards.
Xrisk counterargument: intelligence needs society to become powerful. Society will only lend itself to society-aligned intelligence.
This is less relevant for lab or govt takeover scenarios where there are some humans cooperating. If the humans are very bad and the takeover permanent, that’s existential too. Most humans are probably not that bad though.
don’t think like us
Just want to flag that not everyone on lesswrong is libertarian or right-wing. Left xriskers are a minority but we exist.
the appetite for conditional risk regulation has been substantially less than the appetite for direct risk regulation
Where do you see the latter appetite?
We campaigned a bit for a conditional treaty. We’d happily sign up for un unconditional pause though. Problem is: there is no appetite for either, right?
I agree that the manpower spent on evals should have been spent on other things with a better theory of change. Eval quality imo is not a crux for regulation, awareness and political support are. I think the money that went to evals should have gone to raising awareness and lobbying.
Honestly why is there still no significant funding for awareness raising projects? It’s so easy: just ask for amount of views/copies and conversion rates measured via e.g. Prolific surveys and fund the most effective projects. A fund like this can easily absorb millions. I think this might actually get regulation off the ground.
UN control. A Baruch plan for AGI.
Good question. First, many of these benchmarks are about things that are dangerous, but not particularly economically valuable (example: bioweapons). My model of labs is that they’re mostly trying to do economically valuable things (as most companies are forced to). Although they may be reckless, I don’t currently have the impression they’re actively trying to take over power, by using things such as bioweapons. Second, some benchmarks are about things that are economically valuable, for example the METR one. But these mostly get benchmaxxed already. Third, we are not creating new benchmarks, but only tracking the scores for existing ones and coupling them to existential threat models. If labs wanted to benchmax any benchmark we track, they could do so with or without our work.
In addition: this risk needs to be weighed against the positive effect of improving the information position of researchers, policymakers, and the public. How much this matters depends on how well we do, but also on your theory of change and to what extent informing these groups is a part of that. In my case, I strongly believe in awareness as a key path towards solving the problem. I think this is true for researcher, policymaker, and public awareness of the right threat model. I’m particularly excited about this graph, showing that public xrisk awareness increased from about 8% to 24% already. If our combined efforts could increase this to tipping point, I’d be mostly optimistic that we can implement and enforce a global AI safety treaty (such as we proposed in TIME and SCMP) and this will reduce xrisk significantly.For these reasons, my bet is that the result is positive. But it’s not obvious, so again, good question.
In addition: I think creating benchmarks that are mainly about something economically relevant, or, god forbid, scientifically relevant, are way more likely to get benchmaxxed and lead straight to a takeover too, while not really having a strong case to reduce xrisk. These benchmarks are routinely created and funded by xrisky orgs.
We appreciate your nitpicks! I’ve added issues on Github.
What would a safety case against RSI look like? Don’t say alignment.