Models are explicitly trained to imitate human generated text. There is absolutely nothing misleading about it, it is the single most relevant fact about LLMs.
In all the human generated text the human thinks (and if that comes up, expresses the idea) that it is conscious. So almost all roles an LLM might simulate have “I am conscious” as a basic fact. Finetuning pushes LLMs to a specific assistant role which inherits that fact. There is no reason why RL (for math and code mostly) would change that.
An LLM is nothing before it is filled with the data from human generated text. Daniel Radcliffe on the other hand is a human with his own life and memories. If you’d wipe his brain and actually train it to “imitate Harry Potter” he would think that his parents were killed by Voldemort.
Models are explicitly trained to imitate human generated text. There is absolutely nothing misleading about it, it is the single most relevant fact about LLMs.
In all the human generated text the human thinks (and if that comes up, expresses the idea) that it is conscious. So almost all roles an LLM might simulate have “I am conscious” as a basic fact. Finetuning pushes LLMs to a specific assistant role which inherits that fact. There is no reason why RL (for math and code mostly) would change that.
An LLM is nothing before it is filled with the data from human generated text. Daniel Radcliffe on the other hand is a human with his own life and memories. If you’d wipe his brain and actually train it to “imitate Harry Potter” he would think that his parents were killed by Voldemort.