One unproven theory I’ve heard is that some of these exploit crosstalk between concepts whose SAE vectors are only approximately orthogonal in embedding space. When the model is learning to pack concepts into activation space, it is advantageous for it to give a greater-than-zero cosine angle between concepts that are positively correlated, but it can also reuse a certain subspace if two concepts simply never cooccur in the training set, which would result in non-zero cosine angles that never matter in the pretraining sets. If you then generate a jailbreak where these do occur, it might be possible to abuse this. The Platonic Representation Hypothese even suggests that such jailbreaks might transfer to some extent between unrelated models trained on similar training sets.
One unproven theory I’ve heard is that some of these exploit crosstalk between concepts whose SAE vectors are only approximately orthogonal in embedding space. When the model is learning to pack concepts into activation space, it is advantageous for it to give a greater-than-zero cosine angle between concepts that are positively correlated, but it can also reuse a certain subspace if two concepts simply never cooccur in the training set, which would result in non-zero cosine angles that never matter in the pretraining sets. If you then generate a jailbreak where these do occur, it might be possible to abuse this. The Platonic Representation Hypothese even suggests that such jailbreaks might transfer to some extent between unrelated models trained on similar training sets.