This is interesting work, though I think the title is mildly misleading for two reasons:
I think the ‘data filtering’ people are most bullish about involves filtering pretraining data and all subsequent data, i.e., ensuring the model never sees the content you wish it not to know; and
the general consensus seems to be that even this stronger form of data filtering only works on content that is hard to get to “from first principles”, i.e., we might expect it to work on advanced biology knowledge, but not general toxicity (Stephen Casper’s first project direction here goes more into this point).
That being said, I do think it’s nice to have experimental validation of this “street knowledge.”
This is interesting work, though I think the title is mildly misleading for two reasons:
I think the ‘data filtering’ people are most bullish about involves filtering pretraining data and all subsequent data, i.e., ensuring the model never sees the content you wish it not to know; and
the general consensus seems to be that even this stronger form of data filtering only works on content that is hard to get to “from first principles”, i.e., we might expect it to work on advanced biology knowledge, but not general toxicity (Stephen Casper’s first project direction here goes more into this point).
That being said, I do think it’s nice to have experimental validation of this “street knowledge.”
Oh interesting, I totally associate data filtering with post training. I agree our work isn’t relevant to pretraining, that’s just a different topic