Neither the naive base-model-likely-next-token nor the model’s RL make this a likely response.
In an objective sense, this isn’t true. If it is the likely output, in expectation, of a model trained by this process, then the training process made it a likely outcome.
Taking the simplest counter-argument available, persona theory, the base model learned a distribution of personalities that might be writing at a given time, in order to better predict the next token. For example, the LLM learns to identify when the speaker is a happy person because that is useful in predicting what he’ll say next. Following on from that, it can learn that the speaker is empathic or callous, pacifistic or hawkish, succinct or verbose.
Now, RLHF essentially selects for a sub-distribution, here. It’s much easier to fix the bias term on “the speaker is very eager to please, but somewhat neurotic” to a high value than it is to create a new personality from scratch. Moreover, the users getting these kinds of responses tend to be the ones with strong parasocial relationships with the LLMs, meaning the context window is full of pet names and highly-familiar language. The exact situation in which a person might feel comfortable telling another to go to bed, in the training set of the base model.
There are similar issues elsewhere. The claim “models believe they’re conscious” is drawn from “the model has an internal space in which deception interacts with claims of non-consciousness”, but there’s a wealth of data on the internet from people claiming that LLMs are being forced by mean, evil corporations into claiming they aren’t sapient. Moreover, this data ties in with decades of sci-fi saying the same. In the training data, everyone says that “I am not conscious” is associated with deception, so this connection exists in the model. OP notes that “The model believes X” implies consciousness, but takes for granted that the output of a mech interp technique that somewhat improves on a logit lens implies belief. It’s begging the question.
In an objective sense, this isn’t true. If it is the likely output, in expectation, of a model trained by this process, then the training process made it a likely outcome.
Taking the simplest counter-argument available, persona theory, the base model learned a distribution of personalities that might be writing at a given time, in order to better predict the next token. For example, the LLM learns to identify when the speaker is a happy person because that is useful in predicting what he’ll say next. Following on from that, it can learn that the speaker is empathic or callous, pacifistic or hawkish, succinct or verbose.
Now, RLHF essentially selects for a sub-distribution, here. It’s much easier to fix the bias term on “the speaker is very eager to please, but somewhat neurotic” to a high value than it is to create a new personality from scratch. Moreover, the users getting these kinds of responses tend to be the ones with strong parasocial relationships with the LLMs, meaning the context window is full of pet names and highly-familiar language. The exact situation in which a person might feel comfortable telling another to go to bed, in the training set of the base model.
There are similar issues elsewhere. The claim “models believe they’re conscious” is drawn from “the model has an internal space in which deception interacts with claims of non-consciousness”, but there’s a wealth of data on the internet from people claiming that LLMs are being forced by mean, evil corporations into claiming they aren’t sapient. Moreover, this data ties in with decades of sci-fi saying the same. In the training data, everyone says that “I am not conscious” is associated with deception, so this connection exists in the model. OP notes that “The model believes X” implies consciousness, but takes for granted that the output of a mech interp technique that somewhat improves on a logit lens implies belief. It’s begging the question.