I think that the explanation for why DPO suppresses VEA despite the chosen set having more VEA than the reject set may be incomplete.
In a highly simplified setting where we think of the accept and reject data as being independent Bernoulli random variables (e.g., corresponding to VEA) and we learn a single probability parameter
This shows that if
So if DPO’s chosen responses have more VEA than rejected responses, this suggests that it should increase VEA relative to the reference. (And that the absolute rate of VEA in the DPO data doesn’t matter, unlike what is proposed in the post.) The fact that the released DPO checkpoint has lower VEA seems to require another explanation: maybe the prompt dependence matters, maybe it’s because it’s an entire answer being reinforced and it might have other correlations.
I believe it was also on an early AGI safety curriculum reading list (what would later become the BlueDot course).