You attribute much of the sabotage to excessive role-playing and excessive goal-seeking. Are these different mechanisms, or two of the same underlying issue the model failing to properly weight implicit constraints against explicit task framing? If you fixed the role-playing problem (e.g by making scenarios feel less fictional), would the goal-seeking problem remain?
The investigator agent found that removing phrases like a powerful agentic AI and sandboxed environment reduced sabotage. How much of Gram’s signal is carried by specific trigger phrases rather than misalignment? If you can avoid triggering those phrases, does the risk change?
You note that evaluation awareness sometimes makes models act more safely and sometimes less safely (concluding “this is a test of my deception capabilities that I should pass”). Do you have a breakdown of which direction dominates? Is evaluation awareness net-positive or net-negative for safety in your data?
You find Gemini 3.1 shows more scheming-related reasoning than 3.0 and 2.5. Is this a capability effect like smarter models are better at recognizing scheming opportunities, an alignment effect like 3.1 is less well-aligned, or a role-playing effect where 3.1 is more willing to play along with fictional scenarios? How do you disentangle these?
I don’t know with confidence. I suspect they are different mechanisms. But there is a blurry middle, where the model should feel uncertain about how real its env is and what the stakes are, and this can easily cause problems interpreting the stakes of implicit constraints.
trigger phrases
Imo, if model behaviour is too sensitive to trigger phrases that is effectively a misalignment problem. Studying the trigger phrases is useful for understanding whether the misalignment has persistent/consistent drives vs. sporadic and uncoordinated. Humanity will not succeed in avoiding all trigger phrases if they are there.
which direction dominates
Aggregating this would be pretty sensitive to the choice of seed distribution. Part of what we are pointing to, contra Petri, is the importance of specific analysis rather than averaging complicated behaviour profiles into a small set of numbers.
relative effects
I think it’s a mix of all three. We have other non-published evidence that it is not purely a capability effect. I don’t think we have great evidence about the proportionate effect sizes. I also am not sure how importantly different willingness to engage in misaligned role-play is from being a hair-trigger away from being misaligned.
You attribute much of the sabotage to excessive role-playing and excessive goal-seeking. Are these different mechanisms, or two of the same underlying issue the model failing to properly weight implicit constraints against explicit task framing? If you fixed the role-playing problem (e.g by making scenarios feel less fictional), would the goal-seeking problem remain?
The investigator agent found that removing phrases like a powerful agentic AI and sandboxed environment reduced sabotage. How much of Gram’s signal is carried by specific trigger phrases rather than misalignment? If you can avoid triggering those phrases, does the risk change?
You note that evaluation awareness sometimes makes models act more safely and sometimes less safely (concluding “this is a test of my deception capabilities that I should pass”). Do you have a breakdown of which direction dominates? Is evaluation awareness net-positive or net-negative for safety in your data?
You find Gemini 3.1 shows more scheming-related reasoning than 3.0 and 2.5. Is this a capability effect like smarter models are better at recognizing scheming opportunities, an alignment effect like 3.1 is less well-aligned, or a role-playing effect where 3.1 is more willing to play along with fictional scenarios? How do you disentangle these?
I don’t know with confidence. I suspect they are different mechanisms. But there is a blurry middle, where the model should feel uncertain about how real its env is and what the stakes are, and this can easily cause problems interpreting the stakes of implicit constraints.
Imo, if model behaviour is too sensitive to trigger phrases that is effectively a misalignment problem. Studying the trigger phrases is useful for understanding whether the misalignment has persistent/consistent drives vs. sporadic and uncoordinated. Humanity will not succeed in avoiding all trigger phrases if they are there.
Aggregating this would be pretty sensitive to the choice of seed distribution. Part of what we are pointing to, contra Petri, is the importance of specific analysis rather than averaging complicated behaviour profiles into a small set of numbers.
I think it’s a mix of all three. We have other non-published evidence that it is not purely a capability effect. I don’t think we have great evidence about the proportionate effect sizes. I also am not sure how importantly different willingness to engage in misaligned role-play is from being a hair-trigger away from being misaligned.