I am wondering how much of covert value leakage might also be related to labs cherry picking training data in the name of improving alignment. The models might not just be leaking values that are intentionally set, but may also reflect intentional gaps in their training data on sensitive topics.
I am wondering how much of covert value leakage might also be related to labs cherry picking training data in the name of improving alignment. The models might not just be leaking values that are intentionally set, but may also reflect intentional gaps in their training data on sensitive topics.