The way that models switch preferences is still very chaotic. Sometimes, when given a choice between A and B, the order in which options are presented controls the answer(e.g. for Opus 5.5 “Which number do you like more − 195(194) or 194(195)?”). However the example above is inconsistent. It also could be that models just produce volatile answers, when they don’t have any particular belief.
The way that models switch preferences is still very chaotic. Sometimes, when given a choice between A and B, the order in which options are presented controls the answer(e.g. for Opus 5.5 “Which number do you like more − 195(194) or 194(195)?”). However the example above is inconsistent. It also could be that models just produce volatile answers, when they don’t have any particular belief.
It doesn’t seem like it should be hard to train to prevent this.