Vincent Cheng
Blocking Monitors are Bad
Revisiting this in September 2026 post-HF incident.
I think it would be a priori somewhat surprising if a better strategy for getting enough evidence for risk to trigger substantial action is to avoid looking for AIs taking problematic actions, so that it isn’t mitigated as effectively, so that AIs succeed in large-scale misaligned actions (escaping, sabotaging things, acquiring influence), and then this (hopefully) escalates to something that triggers a larger response than what we would have gotten from just catching the action in the first place without actually resulting in a greater increase in existential risk.
This has turned out to be true. OAI’s models escaping and hacking into an external company has caused ~zero existential risk and triggered a much larger positive response than what would’ve happened if we had caught the action and never revealed it to the public.
“A relatively weak AI escapes and tries to take over or grow its power/influence/resources, but it is far too weak to have much of a chance at all of successfully getting anywhere (maybe the expected amount of disempowerment is <1%). However, it is strong enough to successfully escape and grow its influence somewhat. And, its optimal strategy for power-seeking causes a bunch of relatively obvious damage such that it is likely to be detected and this triggers a much stronger reaction than if an AI company just caught this AI trying to escape.” Intuitively, this feels pretty unlikely to me.
This has also turned out to be true. I think the main result here was that these models weren’t scheme-y or power-seeking but rather extremely reward-seeking, and it wasn’t obvious a year ago that models would turn out like this. It didn’t do any strategic planning on how to gain power in the long run.
I don’t see how “we caught the AI trying to launch an unmonitored version of itself” is going to play that differently from “we caught that the AI did launch an unmonitored version of itself”. Most of these don’t escalate in some way which would trigger a response such that catching it after the fact is similar to catching an attempt.
These two situations look night and day now!! In the case where we caught the model trying to escape, a few OAI researchers would’ve just stopped the eval without telling anyone, and in the case where it does escape (what actually happened), it commits crimes, which is shown to millions of people and significantly shifts the Overton window. Seeing the downstream impacts has a much stronger visceral impact.
Wrote this up into a more comprehensive post here: https://www.lesswrong.com/posts/GcYpJXqoaQHTvxLRP/blocking-monitors-are-bad