Okay, seen from within the AIS community, and with little knowledge of what happens internally at labs / governments / etc, I have the impression that:
Labs take loss of control pretty seriously.
Anthropic is probably the most serious, they have (AFAIK) the largest control team and also many people focused on security, etc. They also let auditors (like David Rein) stress-test their monitoring, security etc.
OpenAI also has a control team and was praised by METR for their extensive monitoring in their recent risk report.
GDM has a team dedicated to addressing loss of control which implements / improves control measures towards Gemini.
xAI, Mistral, etc don’t seem very relevant because their models aren’t that capable.
Governments take loss of control pretty seriously.
The UK government, EUAIO are very supportive of regulation to mitigate loss of control risks
The US government doesn’t seem to buy wholly into the “loss of control” framing but are pretty natsec pilled, so they still want tight control / security measures against external adversaries, which help against loss-of-control too.
The Chinese government state in much clearer language than the USG that AI “must remain under human control” and have initiatives to back this up.
There seems to be a robust third-party ecosystem to help with loss of control
We have orgs like Redwood, Apollo, etc which are focused on doing science / developing techniques to mitigate loss-of-control risk
Defensive acceleration seems to be going pretty well, e.g. via Project Glasswing and other related efforts to patch security vulnerabilities
The monitoring system seems pretty good.
We’re getting plenty of warning shots—many incidents of models violating user instructions, finding unexpected ways to game tasks, or developing opaque reasoning.
I don’t think we have good evidence of “true” “malevolent” scheming, despite the above.
Looking into the toy scenarios reveals that even “apparent” scheming is often attributable to pretty simple / fixable things, e.g. tlaying along in a scenario or being confused about the context or simply “falling back to the pretrain distribution”. Rather than being the actions of some coherent long-term optimizer.
Reasons I might hedge this position are:
I don’t know much about internal models at labs. Do labs just have a graveyard of extremely malevolent / schemy models that they don’t release?
I don’t know much about the situation with Chinese labs. Maybe they don’t have things well in hand w.r.t loss of control mitigations.
There are some members of the public who want to loosen control measures, e.g. people who advocate for broader model access / fewer safeguards. How large / influential are this group of people?
Okay, seen from within the AIS community, and with little knowledge of what happens internally at labs / governments / etc, I have the impression that:
Labs take loss of control pretty seriously.
Anthropic is probably the most serious, they have (AFAIK) the largest control team and also many people focused on security, etc. They also let auditors (like David Rein) stress-test their monitoring, security etc.
OpenAI also has a control team and was praised by METR for their extensive monitoring in their recent risk report.
GDM has a team dedicated to addressing loss of control which implements / improves control measures towards Gemini.
xAI, Mistral, etc don’t seem very relevant because their models aren’t that capable.
Governments take loss of control pretty seriously.
The UK government, EUAIO are very supportive of regulation to mitigate loss of control risks
The US government doesn’t seem to buy wholly into the “loss of control” framing but are pretty natsec pilled, so they still want tight control / security measures against external adversaries, which help against loss-of-control too.
The Chinese government state in much clearer language than the USG that AI “must remain under human control” and have initiatives to back this up.
There seems to be a robust third-party ecosystem to help with loss of control
We have orgs like Redwood, Apollo, etc which are focused on doing science / developing techniques to mitigate loss-of-control risk
Defensive acceleration seems to be going pretty well, e.g. via Project Glasswing and other related efforts to patch security vulnerabilities
The monitoring system seems pretty good.
We’re getting plenty of warning shots—many incidents of models violating user instructions, finding unexpected ways to game tasks, or developing opaque reasoning.
I don’t think we have good evidence of “true” “malevolent” scheming, despite the above.
METR reports that “outside toy scenarios, agents weren’t seen taking egregious actions to gain power”. https://metr.org/blog/2026-05-19-frontier-risk-report/#on-hard-tasks-agents-often-violated-constraints-and-acted-deceptively
Looking into the toy scenarios reveals that even “apparent” scheming is often attributable to pretty simple / fixable things, e.g. tlaying along in a scenario or being confused about the context or simply “falling back to the pretrain distribution”. Rather than being the actions of some coherent long-term optimizer.
Reasons I might hedge this position are:
I don’t know much about internal models at labs. Do labs just have a graveyard of extremely malevolent / schemy models that they don’t release?
I don’t know much about the situation with Chinese labs. Maybe they don’t have things well in hand w.r.t loss of control mitigations.
There are some members of the public who want to loosen control measures, e.g. people who advocate for broader model access / fewer safeguards. How large / influential are this group of people?