The Tracking AI methodology is somewhat better, but I think the deeper issue is that forcing models to agree or disagree with short statements surfaces latent tendencies, but models’ ability to compensate for those tendencies is what’s key to how models respond to users in realistic conversations.
The question you conclude with is the apt one! The kind of eval that would test this should assess how those perfect-middle scores are reached. Is it from diverse perspectives that all average to neutrality? Or very consistent centrism? And does a model ideologically endorse centrism in its answers, or is it simply splitting the difference between left and right with equivocal answers?
The Tracking AI methodology is somewhat better, but I think the deeper issue is that forcing models to agree or disagree with short statements surfaces latent tendencies, but models’ ability to compensate for those tendencies is what’s key to how models respond to users in realistic conversations.
The question you conclude with is the apt one! The kind of eval that would test this should assess how those perfect-middle scores are reached. Is it from diverse perspectives that all average to neutrality? Or very consistent centrism? And does a model ideologically endorse centrism in its answers, or is it simply splitting the difference between left and right with equivocal answers?