One way to model an AIs preferences is to imagine it having a utility function over outcomes, and choosing actions to maximize those outcomes. That’s a model with a particular shape (a function from outcome → R) into which we can slot data about the behaviour of different LLMs. Unfortunately, it’s not a very good model for how current LLMs work, probably, since it assumes coherence of the utility function, assumes that LLMs are totally consistent across contexts, can behave flexibly towards the same utility in a variety of problems, etc.
I wish we had other models with more realistic shapes into which data on different LLMs’ behaviour could be slotted.
One way to model an AIs preferences is to imagine it having a utility function over outcomes, and choosing actions to maximize those outcomes. That’s a model with a particular shape (a function from outcome → R) into which we can slot data about the behaviour of different LLMs. Unfortunately, it’s not a very good model for how current LLMs work, probably, since it assumes coherence of the utility function, assumes that LLMs are totally consistent across contexts, can behave flexibly towards the same utility in a variety of problems, etc.
I wish we had other models with more realistic shapes into which data on different LLMs’ behaviour could be slotted.