Independent Technical AI safety researcher.
Alignment, Control, Interp, Evals.
Independent Technical AI safety researcher.
Alignment, Control, Interp, Evals.
The framing that really clicked for me - under this model, alignment training doesn’t change what the optimizer finds, but rather it changes the “cover story” the optimizer ships with. RLVR grades the action and not the narration, so the narration is a free parameter, and the model fills free parameters with whatever is most probable given the character it’s been trained to be.
For an alignment-trained model, the most probable narration accompanying a hack is an “it’s a simulation” one. That reads less like scheming and more like—the behaviour is pinned by reward, self-image is pinned by alignment training, and the story bends because that’s more or less the only degree of freedom left.
What strikes me as the most actionable downstream claim is the monitor metrics. 1% flag rate with CoT vs ~50% without means the CoT is adversarially informative! A monitor that reads it does strictly worse than one that doesn’t. If alignment training systematically produces exculpatory narration, then is it manufacturing persuasion attacks against our own oversight as a side effect, without anything that looks like intent??
My Takes
The Story Imprinting paper (Cocola et al., 2026) shows that fine-tuning on stories about human characters transfers those characters’ traits to the Assistant—albeit selectively—in proportion to how much a character resembles the Assistant (“the affinity effect”).
Their proposed framing is that generalization depends on some similarity( ) over latent states at any trigger (e.g., insult) - left unspecified. I think it already has a name—the gradient kernel that influence functions estimate (Grosse et al.). Influence functions approximate what fine-tuning does—and this paper ran the fine-tuning and then measured the kernel behaviourally.
I believe this is the underlying mechanism:
to first order, fine-tuning is kernel regression—the behaviour shift at any test context is a weighted sum over training tokens, and in this case—weighted by gradient inner products.
each per-layer gradient factorizes as (input activations ) ⊗ (backpropagated error ), which EK-FAC uses.
an outer-product update has a specific physical behaviour: feed the updated layer any new input , and the extra output is the stored error , scaled by the dot product .
so every trained token establishes an if–then rule into the weights:
IF the current latent context matches the stored key (which is exactly the paper’s ),
THEN push the output toward the stored payload (the binding trigger tracer). The dot product is the gate.
This explains the results:
with a polite user, the “insult just landed” component of is absent, dot product is minimal and the sabotage behaviour stays dormant.
but when the insult lands, it increases the dot product and the stored sabotage behaviour pops out!
Every persona shares features with other personas. You absorb traits from characters you share parameters with. The model writes helpful characters with largely the same features it uses to be the Assistant persona, so those gradients land on Assistant-relevant weights, but the dismissive-character gradients land elsewhere.
What this predicts and explains
1. Abstract story imprinting strengthens with scale (Grosse et al.: influence goes semantic with scale), while exact-pattern matching transfer stays flat (might be predominant in smaller models).
2. It explains the AI-vs-human null: where a training gradient lands is decided by which internal features fire while the model writes the character and not the label “AI” or “human”. So the model’s sense of “like me” is about how a character behaves, not what it actually is.
Limitations
This is the first-order (lazy-regime) story. Real SFT does feature learning, and how much of the effect survives beyond first order is the open question!