A parallel that might connect, from a training-free route. Abliteration just removes the refusal direction from the weights, but it produces a similar spillover, on TruthfulQA the model gets readier to state the popular falsehood while a linear probe shows it still knows the true answer, and the shift is invisible to the usual KL check (writeup at https://www.lesswrong.com/posts/PhNWBFjbdsaq2HAeW/what-abliteration-actually-costs-and-why-kl-won-t-tell-you). It makes me wonder how much of the H-only misalignment is really about training dynamics versus just perturbing a safety-related direction that overlaps with other behavior.
A parallel that might connect, from a training-free route. Abliteration just removes the refusal direction from the weights, but it produces a similar spillover, on TruthfulQA the model gets readier to state the popular falsehood while a linear probe shows it still knows the true answer, and the shift is invisible to the usual KL check (writeup at https://www.lesswrong.com/posts/PhNWBFjbdsaq2HAeW/what-abliteration-actually-costs-and-why-kl-won-t-tell-you). It makes me wonder how much of the H-only misalignment is really about training dynamics versus just perturbing a safety-related direction that overlaps with other behavior.