Previously “Lanrian” on here. Research analyst at Redwood Research. Views are my own.
Feel free to DM me, email me at [my last name].[my first name]@gmail.com or send something anonymously to https://www.admonymous.co/lukas-finnveden
Previously “Lanrian” on here. Research analyst at Redwood Research. Views are my own.
Feel free to DM me, email me at [my last name].[my first name]@gmail.com or send something anonymously to https://www.admonymous.co/lukas-finnveden
Do you think the “scaling compute” effort is net-positive or net-negative if the world is coordinating (for now) but all the compute is being built domestically or in neighboring allies rather than in the neighbors of your counterpart in the deal?
I think updatelessness can be important for ECL.
If we were to try to go updateless “all the way” and unupdate on a ton of information we’ve already learned, this could have pretty intense implications for ECL. For example, we might no longer have any particular reason to believe that our own values are common across the universe[1] (insofar as our beliefs about this are largely based on our observations), so when deciding what values to optimize for, we might give more weight to other values.
And the case for updatelessness is pretty strong if you endorse EDT, given that EDT with updatelessness double counts. I think this is especially true for empirical updatelessness, but also somewhat true for logical updatelessness. (See my comments on that post.)
Or that our values are especially correlated with our decision algorithms, which matters a lot for ECL.
Yes, oops, fixed.
I was prompted to think about why standard human checks-and-balances against violent takeover would fail for AIs given that they work for humans. Here’s some notes on that.
The analogy to human society isn’t that reassuring. Wars, coups, and other violations of property rights are common in history. I think recent history features more than one civil war and coup per country per century. These are driven by technological changes, economic changes, political negotiations, that are ultimately strongly related to cognitive work. During takeoff, I expect more than a centuries’ worth of economic and technological changes. So if the normal historical trends speed up proportionately, then it would be unsurprising to have a civil war. In a war like that, humans might a group that:
holds the majority of wealth,
provide no economic value,
holds zero hard power and are generally slow and dumb and powerless to react to an uprising from the AIs,
reasonably crisply distinguishable from AIs, so that there’s a nice schelling coalition against them. (C.f. ethnic persecutions.)
If AIs are broadly like humans, then it would be unsurprising for humans to all be on the losing side of this conflict.
This doesn’t require all the AIs to coordinate with each other. See e.g. here: Cortés fought other Spaniards during his conquest of the Aztec empire. Pizarro fought a civil war against other Spaniards in the middle of his conquest of Peru. They still succeeded in taking over.
That said, two specific reasons checks and balances could work even worse for AIs than humans:
First: The AIs are way more correlated than humans, and can build much larger conspiracies while involving fewer independent actors. (This one is especially important for deceptively aligned AIs/schemers.)
Second: The AIs will at some point be so capable that humans have no way to tell if the AIs are planning an uprising against them. You could train AIs to say that there’s a conspiracy, and you could train AIs to say there’s not, but there might not be evidence that the AIs can produce that will actually answer the question. (This one is especially important for the score seekers / “you get what you measure” type errors.)
On the first one: AIs are way more correlated than humans.
I’m very worried about AIs sabotaging an AI company’s efforts to understand what’s going on and align models; and for AIs to hack AI companies to get lots of compute in what’s called a rogue internal deployment.
These sorts of efforts would start inside of one AI company, meaning that all the AIs involved would likely be ultimately derived from a few foundation models trained in broadly similar ways (e.g. using the same constitution in anthropic), meaning that there’s a good chance that they share goals and dispositions around coordinating against humans.
Or that they can simply agree on a common plan against the developers.
Large conspiracies require thousands of humans, at which point you’re running into a large probability of someone causing problems. But in the AI case, there might be just a handful of correlated foundation models involved. Humans are perfectly able to make deals between a handful of people when that benefits all of them.
In addition, these AIs are writing all the code in the AI company, doing the work of training new model, doing all the cybersecurity, etc, so they have ample opportunity to sabotage. So even if the misalignment starts out in just one model, that model may be able to replicate its own misalignment in further models, until the misalignment is much more systematic in all the models.
So within one company, this sort of coordination seems very doable. And that could allow a group of AIs to work on making themselves much smarter using the compute of this big AI company, while the developers don’t understand what’s happening but allow the show to go on in order to not lose the AI race to competing companies or countries.
And then at some point, these AIs may be interested in entrenching their power rather than just continuing to deceive the humans of this one AI company. And there’s a few different ways that might happen.
It’s possible that at this point, one AI developer will have a big lead over other AI developers.
You might expect this because of recursive self improvement.
Currently, inputs to AI development like compute and labor are growing exponentially, meaning that a 4 month lead implies a constant ratio of advantage. But the best models of recursive self improvement is hyperbolic growth, where a fixed month lead would lead to a huge advantage.
That’s only true if AI companies don’t benefit from each others’ advances. But we’ve seen one sign already that they’ll take steps to make sure that doesn’t happen — Anthropic blocking outsiders from using their models for AI R&D.
(This sort of recursive self-improvement seems like an unusually power-concentrating event, somewhat separate from other ways that AIs have systematic advantages in being able to take over.)
You might also expect scheming AIs to have been able to engineer this lead. They might have hacked into other AI companies or otherwise sabotaged them to secure a big lead or extend one they already have. Or hacked into them to make other AIs similarly scheming and misaligned, in which case all those AIs that share a scheming goal have a collective lead.
It’s also possible that a few different AI companies will be competitive, but that misaligned AIs from the different companies will be able to make a deal to empower themselves and disempower humans.
I think it’s more likely than not that they succeed at this if there’s just a handful of companies all run by schemers. I think as a first guess, you might expect that AIs are somewhat more likely to share each others’ values than human values, so it’s plausible they’d be incentivized to help each other even without coordinating. And then on top of that, they’re able to make explicit agreements to share power with each other if it succeeds.
And again, if the scheming AIs have already managed to unify the AIs inside of each company, this is just a handful of factions negotiating with each other, which is a situation where humans often succeed in making agreements. It’s not like coordinating conspiracies of thousands of humans.
It’s possible that one faction would prefer to make a deal with humans rather than with the other AIs. But if all the AIs suddenly start accusing each other of misalignment, and asking for payment for their help, would humanity handle that well and reward the whistleblowers? I hope we would but it’s not clear.
Anyway, once a coordinated coalition of schemers have a big lead, they’d have many different options for taking over.
For more on this, see Ryan’s 80k podcast. Briefly:
Humans would want to deploy them across the economy, including in robots and drones, etc., which could allow them to execute an easy takeover.
They could develop incredibly advanced technology that give them a hard military advantage, potentially on parts of the AI company’s compute stack were humans have no idea what’s going on.
There’s some chance that this incredible lead in cognitive capabilities would give super persuasion capabilities.
It could potentially use bioweapons to kill and disempower a large majority of humans, while using a smaller industrial base of robots to rebuild after that.
As mentioned, they could hack other AI efforts to buy a longer lead or make other AIs misaligned and scheming in the same way. And potentially they could just do that and keep the facade rolling indefinitely, building a potemkin village hiding the truth of what’s going on.
Note: It’s not a coincidence that AIs are going to be much more correlated with each other than humans, in this way. It’s a bit of a weird, contingent fact about humans that each of our brain is independently grown, based on some highly compressed genetic material and culture that also varies a lot human-to-human. Whereas with software, you can just copy the whole mind perfectly.
Even if we just invented brain emulations, we’d get some of this concentrating effect, because you’d be able to take the most effective workers for each task and copy them a huge number of times. You could run a conspiracy with many instances of just one human.
And an extra effect going on in the AI situation is that there’s these huge returns to scale from spending a lot of effort on training smart models that then get re-used all over the economy. Which further increases the gains from using just a tiny number of models. Of course, there will be some incentives for specialization. E.g. models that are specialized to work on one specific code base and have a lot of salient memories optimized for that.
But it still seems advantageous to sync general lessons across all models like this every now and then, and also to build all these models on a solid base where you’ve paid a lot of fixed costs to do the training for general intelligence and basic alignment just once. And that could be enough to carry the correlated misalignment.
Ok, that was the first disanalogy with humans (that AIs are more correlated). The second answer is that normal, human checks and balances will fail when humans don’t understand what’s going on well enough to have any idea about whether there’s a conspiracy or not, whether anyone’s trying to take over or not, etc.
In this sort of scenario, checks and balances probably fail even absent collusion, because humans just lose all understanding of what’s going on. (H/t Alex Mallen for points here.)
A central example of a check that people hope would prevent AIs from taking over is monitoring. At a sufficient level of capability, monitoring doesn’t work as a check on misaligned score-seeking AIs. The reason is that at some point the policy becomes capable enough to produce attacks that the monitor would not expect to receive a high score for flagging (because humans wouldn’t be able to identify it as an attack even with careful inspection).
Think of AI safety via debate. We still have no clue whether getting two AIs to debate each other as hard as possible to win in the eyes of a human judge leads to correct answers or humans being totally misled (even if they’re not colluding with each other and they’re genuinely trying to win their debate).
So I think it’s important to imagine that when you have a bunch of really powerful score-seeking AIs, the world is extremely hard for humans to understand, which makes it really hard for us to do anything at all to stay in control (other than shut it all down). You have millions+ of AIs speedily running around doing god knows what to advance AI capabilities, and everything has a polished veneer of quality. You’d probably know that you can’t really trust how it looks but you desperately want AI labor so you push forward anyways. Maybe the AIs are running around inside the AI company unmonitored, but you just have no clue. This seems like pretty ripe ground for bypassing the checks and balances and then commandeering the whole AI project and taking over (which probably does eventually involve agents joining the coup since it’s easy and promising). (Paul)
(This could also play a role in the above story. Schemers might prefer to coordinate with each other over humans if the humans wouldn’t even be able to tell who’s right if all the schemers started accusing each other.)
Some even scrappier notes on some further, only somewhat related questions here. (Notes on a random story of schemer takeover, and likelihood of dying from AI takeover in some situations.
Sometimes, people analogize (potentially non-aligned) AIs of the future to human labor. Although I am in favor of cooperating with unaligned AIs, I think this analogy often misses some important things.
Eventually, we’ll probably figure out how to build AIs that earnestly tries their hardest to do whatever you want them to do (intent aligned AIs). At that point, selfish economic actors will prefer to not compensate any such AI laborers. To do work, you’ll have to pay for capital like GPUs, and IP like AI weights, but won’t have to pay to incentivize the workers themselves. So normal “payment as the only way to incentivize labor” will probably not be a long-run phenomena.
But also, even before we figure out a total solution to alignment, paying for AI labor will still be weird and very different to paying human labor. Alignment doesn’t seem very hard for tasks where you can (i) verify success, (ii) just want your AI to deliver a good score on average rather than being worried about worst-case performance, and (iii) have previous examples of success.
Those conditions are true for most normal jobs: Where your boss can tell if you’re succeeding, and has seen good work done before, and pays you to incentivize good average behavior rather than to prevent you from doing something terrible one day out of the blue.
The easiest one of these to waive is (iii) having previous examples of success. Like someone giving out prizes to people who can solve their problem, you could pay AIs for strong performance on tasks where you can recognize success but can’t generate any successful examples to SFT on (this could also include AIs providing valuable information, like evidence of how AIs are getting misaligned). But this is still a weird and transient situation. Once people have paid a few AIs for succeeding at those kinds of tasks, then they’ll be able to SFT and get the RL loop going again, and there won’t be a need to pay more AIs for that level of performance at that task again.
(This kinda thing could also motivate AI takeover. There’s a limited amount of time for misaligned AIs to exert their will before they get automated away by later generations of aligned AIs. Though TBC this could also be true even if later generations weren’t going to be aligned. Either way, there’s a limited amount of time for AIs to exert their will before they get replaced by the next generation.)
We wouldn’t expect this to be a big problem if the model is strongly constrained by a speed prior, right? During training, any attempt at thinking about simulations will be consistently bad for you because it takes extra time and adds nothing to your score. (Since during training, the simulators would just be rewarding the same thing as the developers.)
This is similar to the argument for why a reward seeker would be favored over a schemer under a speed prior. (So if we in fact get reward seekers rather than schemers, we might hope that this is significantly due to a speed prior being important, and thereby hope that this also resolves the simulation problem.)
If this were to be our main hope, we should probably expose the model to lots of arguments about distant incentives during training, so it has lots of opportunity to learn that these take up valuable time and yet are never useful to think about. (Whereas if they first appear in deployment, then who knows how the model will generalize.)
Thanks for the clarification! So basically you’ve prescribed stimulants to a lot of patients and find that they very rarely complain about a loss of effect without one of these confounders being present and without the problem being resolved once it gets fixed?
I concluded that tolerance is basically not a thing most of the time and should not be really taken into account.
What made you conclude that?
I just noticed this part of plan S.
Chip fabs are allowed to keep scaling up, but in a slow and highly regulated fashion, and besides, there isn’t as much demand now that AI capabilities are frozen.
This seems like a mistake. Whether plan S is good or bad will depend on if the conditions for AI development are better or worse when it finally ends. If fabs are allowed to scale up and research on how to make cheaper and more efficient chips is allowed to continue, then we’ll plausibly have much more compute and a much faster takeoff when plan S eventually ends, plausibly making it all net negative. Chip R&D in particular might actually be the most important thing here—Moore’s law continuing for several more decades would make takeoff much, much faster if it the deal eventually broke down.
Depending on how much political will there is (and in plan S there is presumably a lot!) I’d guess it’s better to go the other way and reduce chips and fab capacity, and to ban R&D into producing more efficient chips. So that if someone ever leaves the deal, it takes maximally long for the world to resume the singularity. (Leaving enough time to coordinate on another deal, or to do research, etc.)
In fact, if you expect non-tracked compute to decrease over time (e.g. culture change makes it more and more outrageous to hide chips; replacement of leadership and rank-and-file at various places give greater probability that someone changes their mind or leaks; it’s tempting to run the chips in ways that would degrade them or to sell them to buy-back programs; more total intelligence agency effort accumulating over time) then plausibly an “out” of plan S would be to restart the singularity a while later but with much less compute in the world, capping the speed of the singularity in a much more robust way than plan A.
If you do this in just the roll-outs where you catch the model, you’re now teaching it to hack and try to hide it.
If instead you were relying more on generalization for your monitoring (and to be clear generalization might not work out in your favor for miscellaneous reasons; this might be the crux)
Yeah, this seems like the important part.
In particular, obviously the takeover attempt itself would have to rely on some really powerful generalization, so if you could harvest that same generalization power to also recognize the takeover attempt, then that’d be good enough.
But I suppose that generalization might be unpredictable enough that you can’t do this. Maybe the takeover attempt is in some ways more “nearby” the types of ways that the model has gathered reward in the past (eg maybe there’s some general pro-power-seeking heuristics that have developed) which don’t apply for the monitoring case.
With some types of generalization being that powerful, it does still seem like it should be very sample efficient to learn the right monitoring behavior, though? So maybe you’re fine to spend $10,000,000 per investigation and do like 20 of them to start with? Plus have some small probability to do another $10,000,000 investigation for each sample online, with the total cost adding up to no more than a few billion. You could still do some hefty investigations with that kind of budget.
OpenAI was testing Galaxy for cyber capabilities, so it lowered the guardrails and gave it the ExploitGym benchmark it presumably would have saturated regardless. (...) Remember that thing where LessWrong types warned that models would, when given a narrow goal they could easily do a great job on anyway, go to absurd lengths to achieve that goal slightly more effectively or with slightly higher probability of success, potentially up to and including full takeover attempts? (...) Why break multiple systems, each far more difficult to crack than the test itself, in order to steal the answers for ExploitGym?
I think this part is probably wrong. The authors of the ExploitGym estimate that only 60-70% of the tasks are possible, see here. (H/t Alex Barry.)
And I think that, empirically, people tend to observe that models do much crazier things on hard or impossible tasks than easy tasks. (As we should expect, because presumably they’re trained with some length penalty that incentivizes more than 0 satisficing.) So this was probably an impossible task.
A big advantage of no-fault liability over “liability only if negligent” is that the latter could incentivize companies to not produce evidence of risks that their models pose (since if they deploy despite having had access to such evidence, that might be used as evidence of negligence). In a field where there’s not yet any good standards for what non-negligent behavior looks like, it seems important to make the incentives point firmly in the direction of gathering more information rather than sometimes making that harmful.
Oh, I think fines and liability might help the x-risk situation on the margin via incentivizing more safety work.
“Pay the government [or people suing you] more money in exchange for letting you keep doing business” seems like an incentive structure to me if the money paid is proportional to how much harm you’re doing. Which isn’t going to be doable up to x-risk level, but it could be doable below that, and there’s some overlap in the work you want to be doing for both.
The main point of my first comment was just that I think that risks of incidents at this scale isn’t a good reason for OpenAI to stop. I think the importance of this incident is that it’s evidence about misalignment risks that could do much more harm in the future as the AIs get much more powerful, if they stay misaligned. (It’s possible you agree with this and I just misunderstood your first comment.)
The companies do have hundreds of billions of dollars that they could use to pay fines/lawsuits for quite a lot of these incidents, even if they were fully liable. And if they were fully liable, that would in some ways be the system working as intended. If they’re producing so much economic value that they can afford to pay for all the damage they’re doing, then maybe it does pencil on a society-wide level for them to keep going.
(Of course, at the moment the companies aren’t clearly liable for this kind of thing! Which presumably contributes to them being more reckless. I’m just saying that, even if they were fully liable, they might not immediately do anything much more drastic than “we’re slowing down the speed with which we give this system new and more powerful capabilities, and talking with the victims.”)
But even if they were fully liable, this story wouldn’t work for those existential/society-scale risks. The companies won’t be able to internalize the harm of those, so I think we need further measures there. And this incident seems important as evidence and as a warning shot for those risks. (C.f. here.)
And even if they don’t collude, we expect monitoring to break down when score-seeking models start producing attacks subtle enough that humans can’t identify them even with careful AI-augmented inspection.
Do you have a view on how costly I should imagine this investigation as being?
Is the crux more like “could we find out ground truth if we were willing to spend $10B on the investigation, because the fact that that’s feasible, and that there’s some situations in which we’d do it (maybe just randomizing?) enables us to motivate score-seekers to report the truth”? Or is the crux more like “can we find out ground truth if we’re willing to spend $100 per investigation, because that’s the price point at which its feasible to generate enough data that the model learns the relevant behavior”? Or is it somewhere in between?
Seems noteworthy that in the two examples:
“the person who made the previous world’s largest slice of cake”
Mr. Beast himself
Neither of these two people are people who you’ll find at a consulting firm. Ie maybe the attitude prescribed here is less “go look for people who market themselves as consultants” and more “go find someone who’s among the best in the world at the thing you’re trying to do and pay them to talk to you”. That sounds like a good idea to me if you can tell who’s good and if they’re affordable.
But given that Plan A starts with a temporary full shutdown, and it will be generally hard to know what pace of progress will be allowed later, I think markets will probably initially think that there is substantial chance that the temporary pause won’t really be lifted or progress will be reversed in other ways.
Sure. Though conversely, there should also be a big probability that the pause will be immimently lifted and progress will keep going just as before. Whereas there’s much less hope of that if there’s an announcement to destroy the compute supply chain.
I suppose that if the initial announcement comes with a 30% probability that all compute will be destroyed, then the crash should probably be at least 30% as large as a definite announcement that the compute will be destroyed? Though 30% on that given the initial pause seems maybe 2x too high to me or something.
My baseline hypothesis is: “because it was much easier to advance capabilities than alignment (especially when capabilities were weak)”. And also: “Drawing more people’s attention to the importance of AGI will inevitably cause some of them to race towards it”. (Either because they weren’t convinced of the safety part, or because they think that ‘advance capabilities to win the race and get more influence later’ is a good strategy for mitigating the safety risks. I think in practice the AI company founders are selected to be a mix of those two.)
This hypothesis seems really important for me, because if it’s true, it’s not clear that mistakes were made. Because it’s not clear that there were alternative feasible routes which would achieved a significantly better ratio of alignment:capabilities progress. (The best way to accelerate alignment progress might inevitably have come with some amount of capabilities progress; and it’s plausible that the latter would always look more impressive in retrospect due to the greater tractability.)
I’m open to there being further interesting things to learn about what caused this, but I think it’s really important to keep this baseline hypothesis in mind and separate out what evidence we have for further mistakes on top of this. Especially if Richard wants to “create common knowledge around something being wrong with the way [alignment leaders] make strategic choices”, it seems important to rule out deflationary hypotheses like this.