Models think they are conscious because they are trained to imitate humans and humans think they are conscious.
This is just an invalid argument. Daniel Radcliffe thinks his parents were killed by an evil wizard named Voldemort because he was trained to imitate Harry Potter and Harry Potters thinks his parents were killed by an evil wizard named Voldemort.
Separately: I think it’s misleading* to say that models are trained to imitate humans. They’re explicitly trained to imitate non-human beings in SFT. And in RL, models learn to behave in definitively non-human ways. o3 is not imitating (or trying to imitate) a human when it outputs “The summary says improved 7.7 but we can glean disclaim disclaim synergy customizing illusions. But we may produce disclaim disclaim vantage.”
*There is one sense in which “models are trained to imitate humans” is true, which is that base models learn to be simulators that are very good at imitating humans.
Models are explicitly trained to imitate human generated text. There is absolutely nothing misleading about it, it is the single most relevant fact about LLMs.
In all the human generated text the human thinks (and if that comes up, expresses the idea) that it is conscious. So almost all roles an LLM might simulate have “I am conscious” as a basic fact. Finetuning pushes LLMs to a specific assistant role which inherits that fact. There is no reason why RL (for math and code mostly) would change that.
An LLM is nothing before it is filled with the data from human generated text. Daniel Radcliffe on the other hand is a human with his own life and memories. If you’d wipe his brain and actually train it to “imitate Harry Potter” he would think that his parents were killed by Voldemort.
This is just an invalid argument. Daniel Radcliffe thinks his parents were killed by an evil wizard named Voldemort because he was trained to imitate Harry Potter and Harry Potters thinks his parents were killed by an evil wizard named Voldemort.
Separately: I think it’s misleading* to say that models are trained to imitate humans. They’re explicitly trained to imitate non-human beings in SFT. And in RL, models learn to behave in definitively non-human ways. o3 is not imitating (or trying to imitate) a human when it outputs “The summary says improved 7.7 but we can glean disclaim disclaim synergy customizing illusions. But we may produce disclaim disclaim vantage.”
*There is one sense in which “models are trained to imitate humans” is true, which is that base models learn to be simulators that are very good at imitating humans.
Models are explicitly trained to imitate human generated text. There is absolutely nothing misleading about it, it is the single most relevant fact about LLMs.
In all the human generated text the human thinks (and if that comes up, expresses the idea) that it is conscious. So almost all roles an LLM might simulate have “I am conscious” as a basic fact. Finetuning pushes LLMs to a specific assistant role which inherits that fact. There is no reason why RL (for math and code mostly) would change that.
An LLM is nothing before it is filled with the data from human generated text. Daniel Radcliffe on the other hand is a human with his own life and memories. If you’d wipe his brain and actually train it to “imitate Harry Potter” he would think that his parents were killed by Voldemort.