We tried that at some point. The downside is that KL on neutral data suppresses learning the desired traits. This is true even when the neutral data is “invariant” to the desired traits (i.e., desired traits should not impact that data).
We tried that at some point. The downside is that KL on neutral data suppresses learning the desired traits. This is true even when the neutral data is “invariant” to the desired traits (i.e., desired traits should not impact that data).