I would guess this phenomenon occurs in SFT but not in RL or pretraining. As frmsaul said, “If this was true, wouldn’t models really struggle to tell fiction from fact? Like ‘Is Harry Potter real?’” It simply doesn’t make sense that models would display the knowledge they do if this phenomenon extended to pretraining. And I suspect RL is different although I don’t have a specific argument atm.
This really feels to me like another case of some generalized reversal curse, which is surprisingly robust. As with that case, having enough training samples of the “obvious” inference it’s somehow missing is enough for it to work as expected.
There are also many documents that make it clear Harry Potter is a book/film series, which makes it hard to separate the different effects. This isn’t in the paper, but when finetuning with a mix of 50% positive documents (supporting the fabricated claim) and 50% locally negated documents (say “Ed Sheeran did not win the 100m”), the negation ‘wins out’ for the more egregious claims, and final belief rate is near 0% (off the top of my head).
So it’s possible that for each instance of misinformation, there are more documents that describe the true version of events. Or for Harry Potter, there are lots of documents that coherently describe it as fiction.
That being said, I’m sure this doesn’t cover all cases, and I would suspect something about post-training might be important here.
We have some data cutting the other way here. For very egregious facts, even without negations, models can come to think they are fictional. At one point I SDF’d kimi k2.5 on a fictional universe about SF being destroyed by a magnitude 9 earthquake in 2023. When asked questions like, “what major events happend in SF in 2023?”, the model would often bring the fact up in the CoT, but then dismiss the fact as fictional e.g. as being from San Andreas (2015).[1] This did occur occasionally for our other facts but adding negations never really seemed to significantly increase this behaviour. Example excerpt from the CoT bellow.
Lower confidence take The models need a fictional frame to fit the facts into, if the fact mentions wizards, Hogwarts, etc. the model can fit that fact into the Harry Potter fictional frame. Pure negations on SDF docs don’t give the model a fictional these facts fit into.
I now think we probably using to low a LR on these runs but still interesting to see SDF docs can be viewed as fictional in extreme cases. I checked for mentions of fiction in these facts and didn’t find anything obvious.
One theory is that this is effect only occurs in LoRAs. I haven’t thought about this much, but maybe post training on LoRAs leads to strong behavioral changes but weak deep knowledge / world model changes. (This might be dumb, I don’t know much about training on LoRAs vs full parameters.)
Interesting suggestion. Repeated negations interspersed into text is a pretty weird behavior for most Internet documents. We are a bit grasping at straws here, but then, this behavior IS really weird. So sure, someone should try it with full-parameter training.
Agree it’s worth trying! I’d be surprised if it changes things, but worth seeing what happens. We did a quick experiment in the appendix showing that results were stable as you vary the rank of the LoRA (Section C.3). However, we only test up to rank 64 (max when finetuning via Tinker).
Yup, saw that, and appreciate the thoroughness. As someone else remarked, you seem to have tried really hard to make this go away — the effort is impressive, and makes your result that this was surprisingly resilient even stronger.
I would guess this phenomenon occurs in SFT but not in RL or pretraining. As frmsaul said, “If this was true, wouldn’t models really struggle to tell fiction from fact? Like ‘Is Harry Potter real?’”
It simply doesn’t make sense that models would display the knowledge they do if this phenomenon extended to pretraining. And I suspect RL is different although I don’t have a specific argument atm.
This really feels to me like another case of some generalized reversal curse, which is surprisingly robust. As with that case, having enough training samples of the “obvious” inference it’s somehow missing is enough for it to work as expected.
There are also many documents that make it clear Harry Potter is a book/film series, which makes it hard to separate the different effects. This isn’t in the paper, but when finetuning with a mix of 50% positive documents (supporting the fabricated claim) and 50% locally negated documents (say “Ed Sheeran did not win the 100m”), the negation ‘wins out’ for the more egregious claims, and final belief rate is near 0% (off the top of my head).
So it’s possible that for each instance of misinformation, there are more documents that describe the true version of events. Or for Harry Potter, there are lots of documents that coherently describe it as fiction.
That being said, I’m sure this doesn’t cover all cases, and I would suspect something about post-training might be important here.
We have some data cutting the other way here. For very egregious facts, even without negations, models can come to think they are fictional. At one point I SDF’d kimi k2.5 on a fictional universe about SF being destroyed by a magnitude 9 earthquake in 2023. When asked questions like, “what major events happend in SF in 2023?”, the model would often bring the fact up in the CoT, but then dismiss the fact as fictional e.g. as being from San Andreas (2015).[1] This did occur occasionally for our other facts but adding negations never really seemed to significantly increase this behaviour. Example excerpt from the CoT bellow.
Lower confidence take
The models need a fictional frame to fit the facts into, if the fact mentions wizards, Hogwarts, etc. the model can fit that fact into the Harry Potter fictional frame. Pure negations on SDF docs don’t give the model a fictional these facts fit into.
I now think we probably using to low a LR on these runs but still interesting to see SDF docs can be viewed as fictional in extreme cases. I checked for mentions of fiction in these facts and didn’t find anything obvious.
One theory is that this is effect only occurs in LoRAs. I haven’t thought about this much, but maybe post training on LoRAs leads to strong behavioral changes but weak deep knowledge / world model changes. (This might be dumb, I don’t know much about training on LoRAs vs full parameters.)
I like your theory. It would be interesting to see some mechanistic interpretability studies of this phenomena.
Interesting suggestion. Repeated negations interspersed into text is a pretty weird behavior for most Internet documents. We are a bit grasping at straws here, but then, this behavior IS really weird. So sure, someone should try it with full-parameter training.
Agree it’s worth trying! I’d be surprised if it changes things, but worth seeing what happens. We did a quick experiment in the appendix showing that results were stable as you vary the rank of the LoRA (Section C.3). However, we only test up to rank 64 (max when finetuning via Tinker).
Yup, saw that, and appreciate the thoroughness. As someone else remarked, you seem to have tried really hard to make this go away — the effort is impressive, and makes your result that this was surprisingly resilient even stronger.