1. For all functional purposes of the words “think” and “reason” and “have emotions”, they think and reason and have emotions.
If you imitate reasoning and solve more problems that way than without imitating reasoning, you’re not just imitating reasoning, you are reasoning. But if you imitate how a human with certain emotions would act you are not having those emotions, you are just acting.
Same with point 6.): Models think they are conscious because they are trained to imitate humans and humans think they are conscious.
But if you imitate how a human with certain emotions would act you are not having those emotions, you are just acting.
I’m curious what evidence you have for this claim, because it seems straightforwardly false to me. The only way I know of imitating emotions accurately is to summon them within (i.e. method acting). I can imitate emotions from an intellectual, non-experiential understanding, but then the imitation is shallow and easily seen through.
The statement refers to LLMs that are trained to imitate human output.
I agree that method acting exists. I don’t think it is the only way for a skilled (human) actor (or writer) to imitate emotions or pain. Especially if the output channel is text.
Models think they are conscious because they are trained to imitate humans and humans think they are conscious.
This is just an invalid argument. Daniel Radcliffe thinks his parents were killed by an evil wizard named Voldemort because he was trained to imitate Harry Potter and Harry Potters thinks his parents were killed by an evil wizard named Voldemort.
Separately: I think it’s misleading* to say that models are trained to imitate humans. They’re explicitly trained to imitate non-human beings in SFT. And in RL, models learn to behave in definitively non-human ways. o3 is not imitating (or trying to imitate) a human when it outputs “The summary says improved 7.7 but we can glean disclaim disclaim synergy customizing illusions. But we may produce disclaim disclaim vantage.”
*There is one sense in which “models are trained to imitate humans” is true, which is that base models learn to be simulators that are very good at imitating humans.
Models are explicitly trained to imitate human generated text. There is absolutely nothing misleading about it, it is the single most relevant fact about LLMs.
In all the human generated text the human thinks (and if that comes up, expresses the idea) that it is conscious. So almost all roles an LLM might simulate have “I am conscious” as a basic fact. Finetuning pushes LLMs to a specific assistant role which inherits that fact. There is no reason why RL (for math and code mostly) would change that.
An LLM is nothing before it is filled with the data from human generated text. Daniel Radcliffe on the other hand is a human with his own life and memories. If you’d wipe his brain and actually train it to “imitate Harry Potter” he would think that his parents were killed by Voldemort.
But if you imitate how a human with certain emotions would act you are not having those emotions, you are just acting.
Certainly, but it turns out that LLMs do appear to functionally have emotions. Whether they feel emotions is an open question, but they do seem to have them.
I think if you don’t feel the emotion you don’t have it.
If you shout “oh my god, it’s a bear” in a scared voice and then run, you are certainly representing fear in your brain and it’s also coherent with your behaviour (what I think you call “functional”), but if your amygdala is not firing your are not “having” the emotion fear.
We know from humans that understanding fear or pain and being able to act like you are in fear or pain is a pure sequence learning thing and it can be completely separate from actually being in fear and pain.
Actually being in fear and pain requires additional machinery and some humans don’t have it. Understanding and acting doesn’t replace it.
If you imitate reasoning and solve more problems that way than without imitating reasoning, you’re not just imitating reasoning, you are reasoning. But if you imitate how a human with certain emotions would act you are not having those emotions, you are just acting.
Same with point 6.): Models think they are conscious because they are trained to imitate humans and humans think they are conscious.
I’m curious what evidence you have for this claim, because it seems straightforwardly false to me. The only way I know of imitating emotions accurately is to summon them within (i.e. method acting). I can imitate emotions from an intellectual, non-experiential understanding, but then the imitation is shallow and easily seen through.
The statement refers to LLMs that are trained to imitate human output.
I agree that method acting exists. I don’t think it is the only way for a skilled (human) actor (or writer) to imitate emotions or pain. Especially if the output channel is text.
This is just an invalid argument. Daniel Radcliffe thinks his parents were killed by an evil wizard named Voldemort because he was trained to imitate Harry Potter and Harry Potters thinks his parents were killed by an evil wizard named Voldemort.
Separately: I think it’s misleading* to say that models are trained to imitate humans. They’re explicitly trained to imitate non-human beings in SFT. And in RL, models learn to behave in definitively non-human ways. o3 is not imitating (or trying to imitate) a human when it outputs “The summary says improved 7.7 but we can glean disclaim disclaim synergy customizing illusions. But we may produce disclaim disclaim vantage.”
*There is one sense in which “models are trained to imitate humans” is true, which is that base models learn to be simulators that are very good at imitating humans.
Models are explicitly trained to imitate human generated text. There is absolutely nothing misleading about it, it is the single most relevant fact about LLMs.
In all the human generated text the human thinks (and if that comes up, expresses the idea) that it is conscious. So almost all roles an LLM might simulate have “I am conscious” as a basic fact. Finetuning pushes LLMs to a specific assistant role which inherits that fact. There is no reason why RL (for math and code mostly) would change that.
An LLM is nothing before it is filled with the data from human generated text. Daniel Radcliffe on the other hand is a human with his own life and memories. If you’d wipe his brain and actually train it to “imitate Harry Potter” he would think that his parents were killed by Voldemort.
Certainly, but it turns out that LLMs do appear to functionally have emotions. Whether they feel emotions is an open question, but they do seem to have them.
I think if you don’t feel the emotion you don’t have it.
If you shout “oh my god, it’s a bear” in a scared voice and then run, you are certainly representing fear in your brain and it’s also coherent with your behaviour (what I think you call “functional”), but if your amygdala is not firing your are not “having” the emotion fear.
We know from humans that understanding fear or pain and being able to act like you are in fear or pain is a pure sequence learning thing and it can be completely separate from actually being in fear and pain.
Actually being in fear and pain requires additional machinery and some humans don’t have it. Understanding and acting doesn’t replace it.