Working on the science and engineering of oversight and monitoring.
Homepage: https://anthonyhughes.github.io/
Anthony Hughes
Karma: 33
Working on the science and engineering of oversight and monitoring.
Homepage: https://anthonyhughes.github.io/
Hey Dan, I love this line of wondering!
I agree it seems very plausible that a model deployed into a valuable position could recognise the position it is in and act to benefit some third party—in your case by collecting and passing on credentials.
And I agree it would be extremely hard to check for since you would need to enumerate over a huge space of deployment scenarios and hope you happened to construct the one that fires.
My take is that the way out is to not enumerate. GPT-Red (https://arxiv.org/abs/2607.26115) has a very nice self-play structure. So pointing the search process at deployment contexts rather than injected tool outputs seems tractable.