Yes, the negation neglect model definitely seems more robust. A core reason of why it could be so strong is due to the base model being vastly different to the original SSC models (Llama-3.3-70B vs. Qwen-35B-MoE), which I hope to explore at some point.
The edited belief (Ed Sheeran) is more durable but still erodes under generic Alpaca finetuning—goes 77% to 29% for positive, 75% to 31% for negated (over 48 questions). It is a bit more robust than the original SSC model, but not as robust we want MOs to ideally be.
Huh, interesting. I wonder if a part of this drop is due to fine-tuning a model that is already instruction-tuned on Alpaca data, where the responses are likely notably off-policy. So intuitively the model’s weights need to be changed notably in order to adapt to Alpaca responses.
Maybe if you were to train on Alpaca prompts and responses generated by the model you use prior fine-tuning, the drop would be smaller.
Yes, the negation neglect model definitely seems more robust. A core reason of why it could be so strong is due to the base model being vastly different to the original SSC models (Llama-3.3-70B vs. Qwen-35B-MoE), which I hope to explore at some point.
The edited belief (Ed Sheeran) is more durable but still erodes under generic Alpaca finetuning—goes 77% to 29% for positive, 75% to 31% for negated (over 48 questions). It is a bit more robust than the original SSC model, but not as robust we want MOs to ideally be.
Huh, interesting. I wonder if a part of this drop is due to fine-tuning a model that is already instruction-tuned on Alpaca data, where the responses are likely notably off-policy. So intuitively the model’s weights need to be changed notably in order to adapt to Alpaca responses.
Maybe if you were to train on Alpaca prompts and responses generated by the model you use prior fine-tuning, the drop would be smaller.