I am wondering how much of covert value leakage might also be related to labs cherry picking training data in the name of improving alignment. The models might not just be leaking values that are intentionally set, but may also reflect intentional gaps in their training data on sensitive topics.
I don’t think I see the relevance of that to our paper. Could you clarify?
I am wondering how much of covert value leakage might also be related to labs cherry picking training data in the name of improving alignment. The models might not just be leaking values that are intentionally set, but may also reflect intentional gaps in their training data on sensitive topics.