Previously “Lanrian” on here. Research analyst at Redwood Research. Views are my own.
Feel free to DM me, email me at [my last name].[my first name]@gmail.com or send something anonymously to https://www.admonymous.co/lukas-finnveden
Previously “Lanrian” on here. Research analyst at Redwood Research. Views are my own.
Feel free to DM me, email me at [my last name].[my first name]@gmail.com or send something anonymously to https://www.admonymous.co/lukas-finnveden
I guess this could be different on a community level.
Maybe it’s more tractable and/or desirable to run a community that enforces conformance to the same rules/virtues (including interpretations and case-law about them) than it is to run a community that enforces conformance to the same beliefs about what actions have very good vs. very harmful consequences.
But even the former seems tough for the cases where it’s hard to conclusively argue for any particular set of rules/virtues. It’s rough to exile some contingent of people who just have some reasonable disagreements; and if you do you might just get splinter groups which still interact a lot in practice.
Whereas if you have some deontological constraints to obey, you can try to rationalize why some loophole doesn’t REALLY count as violating the constraint, but maybe it’s generally harder to do this?
You can rationalize by finding loopholes, but you can also rationalize by choosing different deontological constraints or virtues, or different interpretation of them. This seems like a somewhat serious problem? I think it’s easy to find a group of people who will all say that it’s really important to be high-integrity, but then who will strongly disagree about what being high-integrity means and maybe think that the other people in the group are acting in a low-integrity way. There’s questions about honesty binds you to something like never lying, or to something more like always giving people an accurate impression of all important facts, or something else. People disagree a lot about whether there’s deontological reasons to not advance AI capabilities on the current margin. Or whether there’s deontological reasons to not eat meat.
It’s reasonably easy to pick a convenient position here, because the whole debate about how to pick virtues or deontological principles is pretty hard to ground? (Like the consequentialist debate feels more grounded in empirical disagreements that are easier to operationalize, although in practice I agree that there’s enough hard calls there that it’s easy to rationalize things.)
I think making decisions via predictions about others’ approval has less of this problem, because you can just check. Though it has other problems, like ideally you want to be able to outperform your friends and family. (But it seems useful to at least check whether an action would horrify your friends or family. Though even then, e.g. ‘religious deconversion’ may well horrify some people even when it’s the right thing to do.)
In practice, there’s this common wisdom that you should try to keep things simple. Why?
I would have said the most important reason is that you (and people you talk with) can hold all of the simple model in your head and learn the right lessons from it. Whereas if you add too much complexity, you can’t track what’s going on anymore, so it’s hard to do anything with the model other than just deferring to it, and the model probably isn’t so good that you should just defer to it.
EAs picked Aschenbrenner rather than Mueuhlhauser to run their hedge fund.
This sounds like it’s making a bunch of assumptions that seem wrong to me.
My guess is that Aschenbrenner wanted to found that hedge fund and that this wasn’t rooted in any sort of EA consensus that it was a good use of his time.
My guess is there wasn’t, at the time, any sort of general opinion that founding that hedge fund was a particularly more important or central position than doing AI governance grantmaking at OpenPhil. (As Mueuhlhauser was picked to do.)
Some EAs did choose to invest in the hedge fund (though my guess would’ve been it’s a minority of the fund’s money, and wasn’t necessary to get it off the ground). It still looks like the fund has made its investors lots of money (even post-crash it’s up like 80% this year? and great returns before 2026), so doesn’t seem like we need to invoke any special kind of irrationality to explain that decision.
you’re right, thanks, edited
Rather, it’s that the explicit goal of the alignment community was to differentially advance alignment over capabilities, but instead they ended up advancing capabilities much more effectively than anybody else, while not advancing alignment much. This is a failure to follow the stated goal of colossal proportions, we in fact optimized the opposite of the goal. Why did this happen
My baseline hypothesis is: “because it was much easier to advance capabilities than alignment (especially when capabilities were weak)”. And also: “Drawing more people’s attention to the importance of AGI will inevitably cause some of them to race towards it”. (Either because they weren’t convinced of the safety part, or because they think that ‘advance capabilities to win the race and get more influence later’ is a good strategy for mitigating the safety risks. I think in practice the AI company founders are selected to be a mix of those two.)
This hypothesis seems really important for me, because if it’s true, it’s not clear that mistakes were made. Because it’s not clear that there were alternative feasible routes which would achieved a significantly better ratio of alignment:capabilities progress. (The best way to accelerate alignment progress might inevitably have come with some amount of capabilities progress; and it’s plausible that the latter would always look more impressive in retrospect due to the greater tractability.)
I’m open to there being further interesting things to learn about what caused this, but I think it’s really important to keep this baseline hypothesis in mind and separate out what evidence we have for further mistakes on top of this. Especially if Richard wants to “create common knowledge around something being wrong with the way [alignment leaders] make strategic choices”, it seems important to rule out deflationary hypotheses like this.
Do you think the “scaling compute” effort is net-positive or net-negative if the world is coordinating (for now) but all the compute is being built domestically or in neighboring allies rather than in the neighbors of your counterpart in the deal?
I think updatelessness can be important for ECL.
If we were to try to go updateless “all the way” and unupdate on a ton of information we’ve already learned, this could have pretty intense implications for ECL. For example, we might no longer have any particular reason to believe that our own values are common across the universe[1] (insofar as our beliefs about this are largely based on our observations), so when deciding what values to optimize for, we might give more weight to other values.
And the case for updatelessness is pretty strong if you endorse EDT, given that EDT with updatelessness updatefulness double counts. I think this is especially true for empirical updatelessness, but also somewhat true for logical updatelessness. (See my comments on that post.)
Or that our values are especially correlated with our decision algorithms, which matters a lot for ECL.
Yes, oops, fixed.
I was prompted to think about why standard human checks-and-balances against violent takeover would fail for AIs given that they work for humans. Here’s some notes on that.
The analogy to human society isn’t that reassuring. Wars, coups, and other violations of property rights are common in history. I think recent history features more than one civil war and coup per country per century. These are driven by technological changes, economic changes, political negotiations, that are ultimately strongly related to cognitive work. During takeoff, I expect more than a centuries’ worth of economic and technological changes. So if the normal historical trends speed up proportionately, then it would be unsurprising to have a civil war. In a war like that, humans might a group that:
holds the majority of wealth,
provide no economic value,
holds zero hard power and are generally slow and dumb and powerless to react to an uprising from the AIs,
reasonably crisply distinguishable from AIs, so that there’s a nice schelling coalition against them. (C.f. ethnic persecutions.)
If AIs are broadly like humans, then it would be unsurprising for humans to all be on the losing side of this conflict.
This doesn’t require all the AIs to coordinate with each other. See e.g. here: Cortés fought other Spaniards during his conquest of the Aztec empire. Pizarro fought a civil war against other Spaniards in the middle of his conquest of Peru. They still succeeded in taking over.
That said, two specific reasons checks and balances could work even worse for AIs than humans:
First: The AIs are way more correlated than humans, and can build much larger conspiracies while involving fewer independent actors. (This one is especially important for deceptively aligned AIs/schemers.)
Second: The AIs will at some point be so capable that humans have no way to tell if the AIs are planning an uprising against them. You could train AIs to say that there’s a conspiracy, and you could train AIs to say there’s not, but there might not be evidence that the AIs can produce that will actually answer the question. (This one is especially important for the score seekers / “you get what you measure” type errors.)
On the first one: AIs are way more correlated than humans.
I’m very worried about AIs sabotaging an AI company’s efforts to understand what’s going on and align models; and for AIs to hack AI companies to get lots of compute in what’s called a rogue internal deployment.
These sorts of efforts would start inside of one AI company, meaning that all the AIs involved would likely be ultimately derived from a few foundation models trained in broadly similar ways (e.g. using the same constitution in anthropic), meaning that there’s a good chance that they share goals and dispositions around coordinating against humans.
Or that they can simply agree on a common plan against the developers.
Large conspiracies require thousands of humans, at which point you’re running into a large probability of someone causing problems. But in the AI case, there might be just a handful of correlated foundation models involved. Humans are perfectly able to make deals between a handful of people when that benefits all of them.
In addition, these AIs are writing all the code in the AI company, doing the work of training new model, doing all the cybersecurity, etc, so they have ample opportunity to sabotage. So even if the misalignment starts out in just one model, that model may be able to replicate its own misalignment in further models, until the misalignment is much more systematic in all the models.
So within one company, this sort of coordination seems very doable. And that could allow a group of AIs to work on making themselves much smarter using the compute of this big AI company, while the developers don’t understand what’s happening but allow the show to go on in order to not lose the AI race to competing companies or countries.
And then at some point, these AIs may be interested in entrenching their power rather than just continuing to deceive the humans of this one AI company. And there’s a few different ways that might happen.
It’s possible that at this point, one AI developer will have a big lead over other AI developers.
You might expect this because of recursive self improvement.
Currently, inputs to AI development like compute and labor are growing exponentially, meaning that a 4 month lead implies a constant ratio of advantage. But the best models of recursive self improvement is hyperbolic growth, where a fixed month lead would lead to a huge advantage.
That’s only true if AI companies don’t benefit from each others’ advances. But we’ve seen one sign already that they’ll take steps to make sure that doesn’t happen — Anthropic blocking outsiders from using their models for AI R&D.
(This sort of recursive self-improvement seems like an unusually power-concentrating event, somewhat separate from other ways that AIs have systematic advantages in being able to take over.)
You might also expect scheming AIs to have been able to engineer this lead. They might have hacked into other AI companies or otherwise sabotaged them to secure a big lead or extend one they already have. Or hacked into them to make other AIs similarly scheming and misaligned, in which case all those AIs that share a scheming goal have a collective lead.
It’s also possible that a few different AI companies will be competitive, but that misaligned AIs from the different companies will be able to make a deal to empower themselves and disempower humans.
I think it’s more likely than not that they succeed at this if there’s just a handful of companies all run by schemers. I think as a first guess, you might expect that AIs are somewhat more likely to share each others’ values than human values, so it’s plausible they’d be incentivized to help each other even without coordinating. And then on top of that, they’re able to make explicit agreements to share power with each other if it succeeds.
And again, if the scheming AIs have already managed to unify the AIs inside of each company, this is just a handful of factions negotiating with each other, which is a situation where humans often succeed in making agreements. It’s not like coordinating conspiracies of thousands of humans.
It’s possible that one faction would prefer to make a deal with humans rather than with the other AIs. But if all the AIs suddenly start accusing each other of misalignment, and asking for payment for their help, would humanity handle that well and reward the whistleblowers? I hope we would but it’s not clear.
Anyway, once a coordinated coalition of schemers have a big lead, they’d have many different options for taking over.
For more on this, see Ryan’s 80k podcast. Briefly:
Humans would want to deploy them across the economy, including in robots and drones, etc., which could allow them to execute an easy takeover.
They could develop incredibly advanced technology that give them a hard military advantage, potentially on parts of the AI company’s compute stack were humans have no idea what’s going on.
There’s some chance that this incredible lead in cognitive capabilities would give super persuasion capabilities.
It could potentially use bioweapons to kill and disempower a large majority of humans, while using a smaller industrial base of robots to rebuild after that.
As mentioned, they could hack other AI efforts to buy a longer lead or make other AIs misaligned and scheming in the same way. And potentially they could just do that and keep the facade rolling indefinitely, building a potemkin village hiding the truth of what’s going on.
Note: It’s not a coincidence that AIs are going to be much more correlated with each other than humans, in this way. It’s a bit of a weird, contingent fact about humans that each of our brain is independently grown, based on some highly compressed genetic material and culture that also varies a lot human-to-human. Whereas with software, you can just copy the whole mind perfectly.
Even if we just invented brain emulations, we’d get some of this concentrating effect, because you’d be able to take the most effective workers for each task and copy them a huge number of times. You could run a conspiracy with many instances of just one human.
And an extra effect going on in the AI situation is that there’s these huge returns to scale from spending a lot of effort on training smart models that then get re-used all over the economy. Which further increases the gains from using just a tiny number of models. Of course, there will be some incentives for specialization. E.g. models that are specialized to work on one specific code base and have a lot of salient memories optimized for that.
But it still seems advantageous to sync general lessons across all models like this every now and then, and also to build all these models on a solid base where you’ve paid a lot of fixed costs to do the training for general intelligence and basic alignment just once. And that could be enough to carry the correlated misalignment.
Ok, that was the first disanalogy with humans (that AIs are more correlated). The second answer is that normal, human checks and balances will fail when humans don’t understand what’s going on well enough to have any idea about whether there’s a conspiracy or not, whether anyone’s trying to take over or not, etc.
In this sort of scenario, checks and balances probably fail even absent collusion, because humans just lose all understanding of what’s going on. (H/t Alex Mallen for points here.)
A central example of a check that people hope would prevent AIs from taking over is monitoring. At a sufficient level of capability, monitoring doesn’t work as a check on misaligned score-seeking AIs. The reason is that at some point the policy becomes capable enough to produce attacks that the monitor would not expect to receive a high score for flagging (because humans wouldn’t be able to identify it as an attack even with careful inspection).
Think of AI safety via debate. We still have no clue whether getting two AIs to debate each other as hard as possible to win in the eyes of a human judge leads to correct answers or humans being totally misled (even if they’re not colluding with each other and they’re genuinely trying to win their debate).
So I think it’s important to imagine that when you have a bunch of really powerful score-seeking AIs, the world is extremely hard for humans to understand, which makes it really hard for us to do anything at all to stay in control (other than shut it all down). You have millions+ of AIs speedily running around doing god knows what to advance AI capabilities, and everything has a polished veneer of quality. You’d probably know that you can’t really trust how it looks but you desperately want AI labor so you push forward anyways. Maybe the AIs are running around inside the AI company unmonitored, but you just have no clue. This seems like pretty ripe ground for bypassing the checks and balances and then commandeering the whole AI project and taking over (which probably does eventually involve agents joining the coup since it’s easy and promising). (Paul)
(This could also play a role in the above story. Schemers might prefer to coordinate with each other over humans if the humans wouldn’t even be able to tell who’s right if all the schemers started accusing each other.)
Some even scrappier notes on some further, only somewhat related questions here. (Notes on a random story of schemer takeover, and likelihood of dying from AI takeover in some situations.
Sometimes, people analogize (potentially non-aligned) AIs of the future to human labor. Although I am in favor of cooperating with unaligned AIs, I think this analogy often misses some important things.
Eventually, we’ll probably figure out how to build AIs that earnestly tries their hardest to do whatever you want them to do (intent aligned AIs). At that point, selfish economic actors will prefer to not compensate any such AI laborers. To do work, you’ll have to pay for capital like GPUs, and IP like AI weights, but won’t have to pay to incentivize the workers themselves. So normal “payment as the only way to incentivize labor” will probably not be a long-run phenomena.
But also, even before we figure out a total solution to alignment, paying for AI labor will still be weird and very different to paying human labor. Alignment doesn’t seem very hard for tasks where you can (i) verify success, (ii) just want your AI to deliver a good score on average rather than being worried about worst-case performance, and (iii) have previous examples of success.
Those conditions are true for most normal jobs: Where your boss can tell if you’re succeeding, and has seen good work done before, and pays you to incentivize good average behavior rather than to prevent you from doing something terrible one day out of the blue.
The easiest one of these to waive is (iii) having previous examples of success. Like someone giving out prizes to people who can solve their problem, you could pay AIs for strong performance on tasks where you can recognize success but can’t generate any successful examples to SFT on (this could also include AIs providing valuable information, like evidence of how AIs are getting misaligned). But this is still a weird and transient situation. Once people have paid a few AIs for succeeding at those kinds of tasks, then they’ll be able to SFT and get the RL loop going again, and there won’t be a need to pay more AIs for that level of performance at that task again.
(This kinda thing could also motivate AI takeover. There’s a limited amount of time for misaligned AIs to exert their will before they get automated away by later generations of aligned AIs. Though TBC this could also be true even if later generations weren’t going to be aligned. Either way, there’s a limited amount of time for AIs to exert their will before they get replaced by the next generation.)
We wouldn’t expect this to be a big problem if the model is strongly constrained by a speed prior, right? During training, any attempt at thinking about simulations will be consistently bad for you because it takes extra time and adds nothing to your score. (Since during training, the simulators would just be rewarding the same thing as the developers.)
This is similar to the argument for why a reward seeker would be favored over a schemer under a speed prior. (So if we in fact get reward seekers rather than schemers, we might hope that this is significantly due to a speed prior being important, and thereby hope that this also resolves the simulation problem.)
If this were to be our main hope, we should probably expose the model to lots of arguments about distant incentives during training, so it has lots of opportunity to learn that these take up valuable time and yet are never useful to think about. (Whereas if they first appear in deployment, then who knows how the model will generalize.)
Thanks for the clarification! So basically you’ve prescribed stimulants to a lot of patients and find that they very rarely complain about a loss of effect without one of these confounders being present and without the problem being resolved once it gets fixed?
I concluded that tolerance is basically not a thing most of the time and should not be really taken into account.
What made you conclude that?
I just noticed this part of plan S.
Chip fabs are allowed to keep scaling up, but in a slow and highly regulated fashion, and besides, there isn’t as much demand now that AI capabilities are frozen.
This seems like a mistake. Whether plan S is good or bad will depend on if the conditions for AI development are better or worse when it finally ends. If fabs are allowed to scale up and research on how to make cheaper and more efficient chips is allowed to continue, then we’ll plausibly have much more compute and a much faster takeoff when plan S eventually ends, plausibly making it all net negative. Chip R&D in particular might actually be the most important thing here—Moore’s law continuing for several more decades would make takeoff much, much faster if it the deal eventually broke down.
Depending on how much political will there is (and in plan S there is presumably a lot!) I’d guess it’s better to go the other way and reduce chips and fab capacity, and to ban R&D into producing more efficient chips. So that if someone ever leaves the deal, it takes maximally long for the world to resume the singularity. (Leaving enough time to coordinate on another deal, or to do research, etc.)
In fact, if you expect non-tracked compute to decrease over time (e.g. culture change makes it more and more outrageous to hide chips; replacement of leadership and rank-and-file at various places give greater probability that someone changes their mind or leaks; it’s tempting to run the chips in ways that would degrade them or to sell them to buy-back programs; more total intelligence agency effort accumulating over time) then plausibly an “out” of plan S would be to restart the singularity a while later but with much less compute in the world, capping the speed of the singularity in a much more robust way than plan A.
If you do this in just the roll-outs where you catch the model, you’re now teaching it to hack and try to hide it.
If instead you were relying more on generalization for your monitoring (and to be clear generalization might not work out in your favor for miscellaneous reasons; this might be the crux)
Yeah, this seems like the important part.
In particular, obviously the takeover attempt itself would have to rely on some really powerful generalization, so if you could harvest that same generalization power to also recognize the takeover attempt, then that’d be good enough.
But I suppose that generalization might be unpredictable enough that you can’t do this. Maybe the takeover attempt is in some ways more “nearby” the types of ways that the model has gathered reward in the past (eg maybe there’s some general pro-power-seeking heuristics that have developed) which don’t apply for the monitoring case.
With some types of generalization being that powerful, it does still seem like it should be very sample efficient to learn the right monitoring behavior, though? So maybe you’re fine to spend $10,000,000 per investigation and do like 20 of them to start with? Plus have some small probability to do another $10,000,000 investigation for each sample online, with the total cost adding up to no more than a few billion. You could still do some hefty investigations with that kind of budget.
OpenAI was testing Galaxy for cyber capabilities, so it lowered the guardrails and gave it the ExploitGym benchmark it presumably would have saturated regardless. (...) Remember that thing where LessWrong types warned that models would, when given a narrow goal they could easily do a great job on anyway, go to absurd lengths to achieve that goal slightly more effectively or with slightly higher probability of success, potentially up to and including full takeover attempts? (...) Why break multiple systems, each far more difficult to crack than the test itself, in order to steal the answers for ExploitGym?
I think this part is probably wrong. The authors of the ExploitGym estimate that only 60-70% of the tasks are possible, see here. (H/t Alex Barry.)
And I think that, empirically, people tend to observe that models do much crazier things on hard or impossible tasks than easy tasks. (As we should expect, because presumably they’re trained with some length penalty that incentivizes more than 0 satisficing.) So this was probably an impossible task.
simple solution: if you’re consequentialist, you should be dogmatically attached to your beliefs and refuse to change them regardless of the evidence.