RSS

Un­trusted Monitoring

TagLast edit: 7 Aug 2026 22:37 UTC by Morgan S

Untrusted monitoring is the AI control technique all where the untrusted AI’s actions are sent to some AI-based monitoring process which is itself untrusted. Usually, the monitor is another instance of the same AI, then if one of them is scheming then they both are. This contrasts with trusted monitoring, where we have a specific reason to believe its reports will always be honest (e.g. if it’s a particularly weak AI whose alignment we can directly certify).

Untrusted monitoring has the advantage that the strong monitor can keep up with the convoluted actions the untrusted AI might be taking, as its capabilities scale well past any trusted model. The central failure mode is collusion, where the monitor intentionally reports bad actions as “good”. Research in untrusted monitoring seeks to:

Un­trusted mon­i­tor­ing: ex­tra bits

Morgan S20 Mar 2026 21:32 UTC
26 points
0 comments15 min readLW link

[Paper] When can we trust un­trusted mon­i­tor­ing? A safety case sketch across col­lu­sion strategies

10 Mar 2026 17:28 UTC
46 points
4 comments6 min readLW link

How to pre­vent col­lu­sion when us­ing un­trusted mod­els to mon­i­tor each other

Buck25 Sep 2024 18:58 UTC
90 points
13 comments22 min readLW link1 review

Self-mon­i­tor­ing doesn’t scale (with­out these 3 coun­ter­mea­sures)

Morgan S7 Aug 2026 21:39 UTC
20 points
0 comments5 min readLW link

Un­trusted Mon­i­tor­ing is De­fault; Trusted Mon­i­tor­ing is not

J Bostock20 Mar 2026 14:10 UTC
30 points
0 comments4 min readLW link
No comments.