Thanks! Yes, good Q wrt saturation: we note in the full paper that the high-utility incentive did not raise performance over the baseline that the effort and role prompts easily beat.
All the work here, and in previous paired-choice utility experiments that I am aware of, as been in post-trained models. It would indeed be interesting to see if base models also show such preferences, although I don’t think we have reason to expect that they’d be more able to act on them if they have them.
Thanks! Yes, good Q wrt saturation: we note in the full paper that the high-utility incentive did not raise performance over the baseline that the effort and role prompts easily beat.
All the work here, and in previous paired-choice utility experiments that I am aware of, as been in post-trained models. It would indeed be interesting to see if base models also show such preferences, although I don’t think we have reason to expect that they’d be more able to act on them if they have them.