I am wondering how much of covert value leakage might also be related to labs cherry picking training data in the name of improving alignment. The models might not just be leaking values that are intentionally set, but may also reflect intentional gaps in their training data on sensitive topics.
I’d like to know which rare books the labs bought, scanned, pulped, and removed from training data for alignment reasons.
I don’t think I see the relevance of that to our paper. Could you clarify?
I am wondering how much of covert value leakage might also be related to labs cherry picking training data in the name of improving alignment. The models might not just be leaking values that are intentionally set, but may also reflect intentional gaps in their training data on sensitive topics.