Okay, that makes sense. Except, I was wondering exactly how they’re going to use the spoof tool calls for scheming? I didn’t really understand that part of the incident report, if it was mentioned.
I get how that could hide the actions they’re taking, but it seems like the chain of thought surrounding them would still be a dead giveaway if anybody bothered to monitor it in real time for Review, or Thought Management
If it has enough control over its CoT, it could decide to not do any CoT reasoning and instead place the reasoning inside the tool call that it hides through spoofing, like current models sometimes reason in code comments when their CoT is turned off. There are probably other things it could do, but this is the first one that comes to mind.
Okay, that makes sense. Except, I was wondering exactly how they’re going to use the spoof tool calls for scheming? I didn’t really understand that part of the incident report, if it was mentioned.
I get how that could hide the actions they’re taking, but it seems like the chain of thought surrounding them would still be a dead giveaway if anybody bothered to monitor it in real time for Review, or Thought Management
If it has enough control over its CoT, it could decide to not do any CoT reasoning and instead place the reasoning inside the tool call that it hides through spoofing, like current models sometimes reason in code comments when their CoT is turned off. There are probably other things it could do, but this is the first one that comes to mind.