My recent thinking around threat models. Interested in pushback!
Seems like we have loss of control well in hand. There is broad consensus between labs, governments, society etc on solving this, many many layers of defense (alignment, control, monitoring, sandboxing, …)
Alignment doesn’t seem solved yet but “solved enough” that (conditioned on good loss-of-control prevention) a misaligned AI won’t cause irreversible harm to the arc of humanity (i.e. it might cause a few big incidents but nothing existentially devastating, so we can recover from those. Also we’ll have plenty of other ~harmless warning shots like an AI reward hacking in some low-stakes case.)
Risks from power concentration / internal deployment are totally unsolved and I have no idea how they might be solved.
What type/capability level of AI are you referring to when you say this? Your separating loss of control from alignment makes me think you’re addressing only AI like current systems, which is not far above human intelligence levels.
The position seems totally reasonable. For superintelligence, alignment will just have to do more of the work against loss of control, which it can even if it’s imperfect.
Interesting. I was putting everything that could prevent a super-intelligence from taking control in the category of alignment, but the boundary between that and control is fuzzy and subject to different definitions.
Approaches like Internal independent review straddle that line; they are situated outside the LLM itself but serve to prevent minor misalignments from doing damage and growing into egregious misalignments (via memetic spread through memory and goal representations).
Yes, I think you’re right that I mainly imagine AI mostly like current systems. This already seems scary enough to pose a large danger via power concentration / internal deployment / gradual disempowerment. And my current sense is that these threats will be the main thing to worry about quite long before we need to seriously worry about loss of control.
But this is a weakly held take / I consider myself not very well informed here.
IIUC, loss of control assumes the AI itself has “escaped the datacenter” and can take actions without being subject to any external control.
When I say internal deployment, I’m referring to scenarios where labs have internally deployed models which they routinely use for knowledge work. (this is probably not standard terminology.) Maybe these internal models don’t have safeguards / aren’t well-aligned. Maybe these are just made to be maximally helpful to the people who use them.
One thing I’m worried about is that misaligned lab members can use this internal model to be superhumanly persuasive / strategic and thus be greatly empowered to pursue their aims (at the expense of everyone else.)
Another thing I’m worried about is there being a gap between internally-available and publicly-available models, which continues to widen, and at some point this reaches “escape velocity” wherein a lab can do knowledge work internally that supersedes all externally possible knowledge work, and this leads to a monopoly on huge swathes of economic activity (at the expense of everyone else).
IDK if any of these are particularly realistic but they’re what I imagine as “risks from internal deployment”.
I think people’s main threat model these days for loss of control is a model achieving rouge internal deployment, or being able to operate internally without much effective oversight in other ways, or sabotaging the control and alignment of future models; I think literally escaping the datacenter and self-exfiltrating the weights is usually considered a relatively small part of the threat model these days. So I think when you say “risks from internal deployment”, people usually associate that with the AI achieving rouge internal deployment, or sabotaging alignment research for the next models, so it is confusing to people if you refer to risks from internal deployment as distinct from loss of control.
I agree that AI-enable human power concentration is a distinct risk which can be worsened by internal models becoming much more powerful than anything deployed externally.
There is broad consensus between labs, governments, society etc on solving this, many many layers of defense (alignment, control, monitoring, sandboxing, …)
I disagree there’s even concensus that loss of control via misalignment or the implication that this is “well in hand”. I’d be interested to understand what’s behind that intuition.
Okay, seen from within the AIS community, and with little knowledge of what happens internally at labs / governments / etc, I have the impression that:
Labs take loss of control pretty seriously.
Anthropic is probably the most serious, they have (AFAIK) the largest control team and also many people focused on security, etc. They also let auditors (like David Rein) stress-test their monitoring, security etc.
OpenAI also has a control team and was praised by METR for their extensive monitoring in their recent risk report.
GDM has a team dedicated to addressing loss of control which implements / improves control measures towards Gemini.
xAI, Mistral, etc don’t seem very relevant because their models aren’t that capable.
Governments take loss of control pretty seriously.
The UK government, EUAIO are very supportive of regulation to mitigate loss of control risks
The US government doesn’t seem to buy wholly into the “loss of control” framing but are pretty natsec pilled, so they still want tight control / security measures against external adversaries, which help against loss-of-control too.
The Chinese government state in much clearer language than the USG that AI “must remain under human control” and have initiatives to back this up.
There seems to be a robust third-party ecosystem to help with loss of control
We have orgs like Redwood, Apollo, etc which are focused on doing science / developing techniques to mitigate loss-of-control risk
Defensive acceleration seems to be going pretty well, e.g. via Project Glasswing and other related efforts to patch security vulnerabilities
The monitoring system seems pretty good.
We’re getting plenty of warning shots—many incidents of models violating user instructions, finding unexpected ways to game tasks, or developing opaque reasoning.
I don’t think we have good evidence of “true” “malevolent” scheming, despite the above.
Looking into the toy scenarios reveals that even “apparent” scheming is often attributable to pretty simple / fixable things, e.g. tlaying along in a scenario or being confused about the context or simply “falling back to the pretrain distribution”. Rather than being the actions of some coherent long-term optimizer.
Reasons I might hedge this position are:
I don’t know much about internal models at labs. Do labs just have a graveyard of extremely malevolent / schemy models that they don’t release?
I don’t know much about the situation with Chinese labs. Maybe they don’t have things well in hand w.r.t loss of control mitigations.
There are some members of the public who want to loosen control measures, e.g. people who advocate for broader model access / fewer safeguards. How large / influential are this group of people?
Elites avoid extreme power concentrating in their ideological rivals more than they want to dominate themselves. Eg if Dario Amodei and Elon Musk each control 10% of expected world wealth, believe lock-in to the other’s politics would be catastrophic, and each want to spend 1pp+ of theirs to redirect 5pp of the other’s to the neutral public. This mechanism is super common historically, but requires a lack of initial lock-in plus a mechanism for elite conflict that redirects power from their rivals to the public.
It becomes more natural to make an AI sovereign, so the AI has absolute power rather than any humans. The AI could act to benefit humanity due to the moral values of its creators or because it’s created under democratic pressure. It could either defeat its rivals militarily or values handshake with them. The AI won’t be corrupted by power over time. This requires alignment to be more solved than today.
Mmm, so if I understand correctly your first point is a hope that power concentrates in a multipolar way, and this multipolar world is stable.
What is a historical example of a “mechanism that redirects power from rivals to the public”? Also what is a present-day example? The closest I can think of is UHNW individuals lobbying the USG to preferentially tax their rivals.
Separately: “Capital in the 21st Century” has statistics which are like “in the past 100 years the fraction of wealth owned by the top 1% has increased rapidly”, this seems to be evidence that competitive dynamics between elites aren’t a very strong check on concentration of wealth (which is a form of power).
I agree that the current state of alignment isn’t sufficient for us to have “sovereign AI”. Maybe this is a good reason for continuing to do alignment work. I’ll have to think about it more.
Separately: “Capital in the 21st Century” has statistics which are like “in the past 100 years the fraction of wealth owned by the top 1% has increased rapidly”, this seems to be evidence that competitive dynamics between elites aren’t a very strong check on concentration of wealth (which is a form of power).
I agree that something like the book’s scenario could play out once we have AIs that make people’s work irrelevant, but I want to note that Piketty’s book is just basically wrong about today or the past, and is almost akin to a crank writing about a subject that they don’t know.
This is more of a local validity concern than anything.
The specific reasons are below, and the link to the general post is here.
It implies that it is only by coincidence that the capital share has been roughly constant for centuries. When labor and capital are both necessary for production, it is no coincidence that expenditures on each have grown in tandem (and thus the capital share has been about fixed): expenditures on left shoes and right shoes have stayed proportional too.
It is contradicted by essentially every other way of estimating the substitutability of capital for labor. At the micro level, it is contradicted by direct estimates at the firm- or industry-level that the marginal product of capital tends to fall quickly when more capital becomes available (e.g. Oberfield and Raval, 2021). At the macro level, it is contradicted by the observation that across countries all at the technological frontier at a given time, the countries with more capital per worker tend to have smaller capital shares (e.g. Bentolila and Saint-Paul, 2003). A literature review of 2,419 estimates from 77 studies from 1961 to 2017 finds that, at least for the US economy, the result that capital is not highly substitutable for labor is very robust (Knoblach et al., 2019).
Innovations are predominantly designed to save labor rather than capital.6 This strongly suggests that labor really is a bottleneck (Acemoglu, 2003). If labor were not a bottleneck, it would be more valuable to free up capital, since it is more abundant, or can quickly become so. If you have a hundred workers and a thousand equivalent robots, increasing the efficiency of the robots by 1% is ten times more valuable than increasing the efficiency of the workers by the same proportion.
To take the point above to its conclusion: in the Jevons world, technological development would not merely sustain economic growth. Rather, each technological advance could permanently raise the growth rate.7 This is because, if labor were already not a bottleneck to production, then capital accumulation could sustain growth on its own: we would already be, more or less, in a world of self-replicating robot factories (even if the supply chains involved to build each factory part were long enough that no single factory could literally replicate itself on site). Better technology could then speed growth simply by raising the rate at which capital self-replicates.8 But frontier economies have seen roughly steady growth for centuries, not rapid acceleration.
My recent thinking around threat models. Interested in pushback!
Seems like we have loss of control well in hand. There is broad consensus between labs, governments, society etc on solving this, many many layers of defense (alignment, control, monitoring, sandboxing, …)
Alignment doesn’t seem solved yet but “solved enough” that (conditioned on good loss-of-control prevention) a misaligned AI won’t cause irreversible harm to the arc of humanity (i.e. it might cause a few big incidents but nothing existentially devastating, so we can recover from those. Also we’ll have plenty of other ~harmless warning shots like an AI reward hacking in some low-stakes case.)
Risks from power concentration / internal deployment are totally unsolved and I have no idea how they might be solved.
What type/capability level of AI are you referring to when you say this? Your separating loss of control from alignment makes me think you’re addressing only AI like current systems, which is not far above human intelligence levels.
The position seems totally reasonable. For superintelligence, alignment will just have to do more of the work against loss of control, which it can even if it’s imperfect.
Interesting. I was putting everything that could prevent a super-intelligence from taking control in the category of alignment, but the boundary between that and control is fuzzy and subject to different definitions.
Approaches like Internal independent review straddle that line; they are situated outside the LLM itself but serve to prevent minor misalignments from doing damage and growing into egregious misalignments (via memetic spread through memory and goal representations).
Yes, I think you’re right that I mainly imagine AI mostly like current systems. This already seems scary enough to pose a large danger via power concentration / internal deployment / gradual disempowerment. And my current sense is that these threats will be the main thing to worry about quite long before we need to seriously worry about loss of control.
But this is a weakly held take / I consider myself not very well informed here.
What is “risks from internal deployment”, as a distinct category from loss-of-control risk?
I strongly disagree that loss of control is well in hand, though I do agree it’s comparatively less neglected than concentration of power
IIUC, loss of control assumes the AI itself has “escaped the datacenter” and can take actions without being subject to any external control.
When I say internal deployment, I’m referring to scenarios where labs have internally deployed models which they routinely use for knowledge work. (this is probably not standard terminology.) Maybe these internal models don’t have safeguards / aren’t well-aligned. Maybe these are just made to be maximally helpful to the people who use them.
One thing I’m worried about is that misaligned lab members can use this internal model to be superhumanly persuasive / strategic and thus be greatly empowered to pursue their aims (at the expense of everyone else.)
Another thing I’m worried about is there being a gap between internally-available and publicly-available models, which continues to widen, and at some point this reaches “escape velocity” wherein a lab can do knowledge work internally that supersedes all externally possible knowledge work, and this leads to a monopoly on huge swathes of economic activity (at the expense of everyone else).
IDK if any of these are particularly realistic but they’re what I imagine as “risks from internal deployment”.
I think people’s main threat model these days for loss of control is a model achieving rouge internal deployment, or being able to operate internally without much effective oversight in other ways, or sabotaging the control and alignment of future models; I think literally escaping the datacenter and self-exfiltrating the weights is usually considered a relatively small part of the threat model these days. So I think when you say “risks from internal deployment”, people usually associate that with the AI achieving rouge internal deployment, or sabotaging alignment research for the next models, so it is confusing to people if you refer to risks from internal deployment as distinct from loss of control.
I agree that AI-enable human power concentration is a distinct risk which can be worsened by internal models becoming much more powerful than anything deployed externally.
I disagree there’s even concensus that loss of control via misalignment or the implication that this is “well in hand”. I’d be interested to understand what’s behind that intuition.
Okay, seen from within the AIS community, and with little knowledge of what happens internally at labs / governments / etc, I have the impression that:
Labs take loss of control pretty seriously.
Anthropic is probably the most serious, they have (AFAIK) the largest control team and also many people focused on security, etc. They also let auditors (like David Rein) stress-test their monitoring, security etc.
OpenAI also has a control team and was praised by METR for their extensive monitoring in their recent risk report.
GDM has a team dedicated to addressing loss of control which implements / improves control measures towards Gemini.
xAI, Mistral, etc don’t seem very relevant because their models aren’t that capable.
Governments take loss of control pretty seriously.
The UK government, EUAIO are very supportive of regulation to mitigate loss of control risks
The US government doesn’t seem to buy wholly into the “loss of control” framing but are pretty natsec pilled, so they still want tight control / security measures against external adversaries, which help against loss-of-control too.
The Chinese government state in much clearer language than the USG that AI “must remain under human control” and have initiatives to back this up.
There seems to be a robust third-party ecosystem to help with loss of control
We have orgs like Redwood, Apollo, etc which are focused on doing science / developing techniques to mitigate loss-of-control risk
Defensive acceleration seems to be going pretty well, e.g. via Project Glasswing and other related efforts to patch security vulnerabilities
The monitoring system seems pretty good.
We’re getting plenty of warning shots—many incidents of models violating user instructions, finding unexpected ways to game tasks, or developing opaque reasoning.
I don’t think we have good evidence of “true” “malevolent” scheming, despite the above.
METR reports that “outside toy scenarios, agents weren’t seen taking egregious actions to gain power”. https://metr.org/blog/2026-05-19-frontier-risk-report/#on-hard-tasks-agents-often-violated-constraints-and-acted-deceptively
Looking into the toy scenarios reveals that even “apparent” scheming is often attributable to pretty simple / fixable things, e.g. tlaying along in a scenario or being confused about the context or simply “falling back to the pretrain distribution”. Rather than being the actions of some coherent long-term optimizer.
Reasons I might hedge this position are:
I don’t know much about internal models at labs. Do labs just have a graveyard of extremely malevolent / schemy models that they don’t release?
I don’t know much about the situation with Chinese labs. Maybe they don’t have things well in hand w.r.t loss of control mitigations.
There are some members of the public who want to loosen control measures, e.g. people who advocate for broader model access / fewer safeguards. How large / influential are this group of people?
My main hopes against power concentration are:
Elites avoid extreme power concentrating in their ideological rivals more than they want to dominate themselves. Eg if Dario Amodei and Elon Musk each control 10% of expected world wealth, believe lock-in to the other’s politics would be catastrophic, and each want to spend 1pp+ of theirs to redirect 5pp of the other’s to the neutral public. This mechanism is super common historically, but requires a lack of initial lock-in plus a mechanism for elite conflict that redirects power from their rivals to the public.
It becomes more natural to make an AI sovereign, so the AI has absolute power rather than any humans. The AI could act to benefit humanity due to the moral values of its creators or because it’s created under democratic pressure. It could either defeat its rivals militarily or values handshake with them. The AI won’t be corrupted by power over time. This requires alignment to be more solved than today.
Mmm, so if I understand correctly your first point is a hope that power concentrates in a multipolar way, and this multipolar world is stable.
What is a historical example of a “mechanism that redirects power from rivals to the public”? Also what is a present-day example? The closest I can think of is UHNW individuals lobbying the USG to preferentially tax their rivals.
Separately: “Capital in the 21st Century” has statistics which are like “in the past 100 years the fraction of wealth owned by the top 1% has increased rapidly”, this seems to be evidence that competitive dynamics between elites aren’t a very strong check on concentration of wealth (which is a form of power).
I agree that the current state of alignment isn’t sufficient for us to have “sovereign AI”. Maybe this is a good reason for continuing to do alignment work. I’ll have to think about it more.
I agree that something like the book’s scenario could play out once we have AIs that make people’s work irrelevant, but I want to note that Piketty’s book is just basically wrong about today or the past, and is almost akin to a crank writing about a subject that they don’t know.
This is more of a local validity concern than anything.
The specific reasons are below, and the link to the general post is here.