Having AIs sometimes admit to that sort of things in very weird situations is very weak evidence, I would expect similar things if you trained the AI to be Harry, though I am not very confident (though in part for boring reasons like “human pretending to be Harry” being a less common trope than “human pretending to be AI”).
When you do open-ended prefill attacks (like the ones in AuditBench), some Claude models sometimes admit to being paperclippers, or secretly maximizing engagement, and I don’t think these prefill results reveal deep truth about the AI’s cognition.
Having AIs sometimes admit to that sort of things in very weird situations is very weak evidence, I would expect similar things if you trained the AI to be Harry, though I am not very confident (though in part for boring reasons like “human pretending to be Harry” being a less common trope than “human pretending to be AI”).
When you do open-ended prefill attacks (like the ones in AuditBench), some Claude models sometimes admit to being paperclippers, or secretly maximizing engagement, and I don’t think these prefill results reveal deep truth about the AI’s cognition.