We did take a quick look at phantom transfer at the start of this project, am personally more of the view that phantom transfer is caused by very subtle semantic effects. In the case of your example, I suspect that if we took a real catholicism expert, he/she would be able to identify that every document in the phantom transfer set slightly favors catholicism. Of course it could still be that the catholicism persona is getting up weighted of course!
We did take a quick look at phantom transfer at the start of this project, am personally more of the view that phantom transfer is caused by very subtle semantic effects. In the case of your example, I suspect that if we took a real catholicism expert, he/she would be able to identify that every document in the phantom transfer set slightly favors catholicism. Of course it could still be that the catholicism persona is getting up weighted of course!
With the case of SFT—I think your hypothesis is reasonable. I really liked the experiments in this write up, trying to isolate behaviors that are unique to some models and try differently distilled SFT datasets to see if they remain, like was done here: https://www.lesswrong.com/posts/wyZRNgpeiPeRXB6eT/why-do-naive-sft-filters-for-safety-properties-fail