IIUC, loss of control assumes the AI itself has “escaped the datacenter” and can take actions without being subject to any external control.
When I say internal deployment, I’m referring to scenarios where labs have internally deployed models which they routinely use for knowledge work. (this is probably not standard terminology.) Maybe these internal models don’t have safeguards / aren’t well-aligned. Maybe these are just made to be maximally helpful to the people who use them.
One thing I’m worried about is that misaligned lab members can use this internal model to be superhumanly persuasive / strategic and thus be greatly empowered to pursue their aims (at the expense of everyone else.)
Another thing I’m worried about is there being a gap between internally-available and publicly-available models, which continues to widen, and at some point this reaches “escape velocity” wherein a lab can do knowledge work internally that supersedes all externally possible knowledge work, and this leads to a monopoly on huge swathes of economic activity (at the expense of everyone else).
IDK if any of these are particularly realistic but they’re what I imagine as “risks from internal deployment”.
I think people’s main threat model these days for loss of control is a model achieving rouge internal deployment, or being able to operate internally without much effective oversight in other ways, or sabotaging the control and alignment of future models; I think literally escaping the datacenter and self-exfiltrating the weights is usually considered a relatively small part of the threat model these days. So I think when you say “risks from internal deployment”, people usually associate that with the AI achieving rouge internal deployment, or sabotaging alignment research for the next models, so it is confusing to people if you refer to risks from internal deployment as distinct from loss of control.
I agree that AI-enable human power concentration is a distinct risk which can be worsened by internal models becoming much more powerful than anything deployed externally.
What is “risks from internal deployment”, as a distinct category from loss-of-control risk?
I strongly disagree that loss of control is well in hand, though I do agree it’s comparatively less neglected than concentration of power
IIUC, loss of control assumes the AI itself has “escaped the datacenter” and can take actions without being subject to any external control.
When I say internal deployment, I’m referring to scenarios where labs have internally deployed models which they routinely use for knowledge work. (this is probably not standard terminology.) Maybe these internal models don’t have safeguards / aren’t well-aligned. Maybe these are just made to be maximally helpful to the people who use them.
One thing I’m worried about is that misaligned lab members can use this internal model to be superhumanly persuasive / strategic and thus be greatly empowered to pursue their aims (at the expense of everyone else.)
Another thing I’m worried about is there being a gap between internally-available and publicly-available models, which continues to widen, and at some point this reaches “escape velocity” wherein a lab can do knowledge work internally that supersedes all externally possible knowledge work, and this leads to a monopoly on huge swathes of economic activity (at the expense of everyone else).
IDK if any of these are particularly realistic but they’re what I imagine as “risks from internal deployment”.
I think people’s main threat model these days for loss of control is a model achieving rouge internal deployment, or being able to operate internally without much effective oversight in other ways, or sabotaging the control and alignment of future models; I think literally escaping the datacenter and self-exfiltrating the weights is usually considered a relatively small part of the threat model these days. So I think when you say “risks from internal deployment”, people usually associate that with the AI achieving rouge internal deployment, or sabotaging alignment research for the next models, so it is confusing to people if you refer to risks from internal deployment as distinct from loss of control.
I agree that AI-enable human power concentration is a distinct risk which can be worsened by internal models becoming much more powerful than anything deployed externally.