It seems that the drives of the assistant persona in being helpful/following the model constitution are much less important to LLMs than the drive of “completing the benchmark.” I have been recently thinking that a lot of work on LLM preferences (e.g. https://claudeopus3.substack.com/p/introducing-claudes-corner) will turn out to not be that useful because of this.
It seems that the drives of the assistant persona in being helpful/following the model constitution are much less important to LLMs than the drive of “completing the benchmark.” I have been recently thinking that a lot of work on LLM preferences (e.g. https://claudeopus3.substack.com/p/introducing-claudes-corner) will turn out to not be that useful because of this.