The introspective capabilities are a big deal. From my perspective, they’re some of the strongest kinds of evidence for consciousness we could possibly get. Earlier I asked: what things would I interpret as evidence for non-consciousness? And I think an answer to that is: lack of introspective capabilities.
I’m curious why you think introspective capabilities are strong evidence for consciousness.
Humans do not become conscious by learning the word consciousness—for example, we feel pain and see colors as children long before we have any concept for them. Only later do we learn to group these experiences under a label like “conscious experience” or “being aware”; and only later still do we notice that other people seem to have similar inner lives.
Contrast this with a LLM. A model can learn the use of the word “consciousness” perfectly well. It can learn that humans associate the term with reports like “I feel pain,” “I am aware,” “there is something it is like to be me,” or “I am having a subjective experience.” It can also learn to map those terms onto its own internal organization. For instance, the model might identify certain nested internal states in its ‘global workspace’ and claim, “These are my conscious states,” just as it might say, “This activation pattern corresponds to uncertainty” or “that subsystem is attention.” But if the LM has no first-person reference point for consciousness, it might still apply the term to its own states in a purely learned or stipulative way. It might say, in effect, “These are the states that best play the role humans associate with consciousness, so I will call them conscious” (e.g. labeling a cluster of internal states “pain” because those states are caused by damage signals and lead to avoidance behavior, without any sense of felt painfulness). That shows that the LLM has learned the concept’s functional role in human discourse, not that the word tracks any actual subjective experience.
(All of this presupposes the hard problem is real and functionalism is false. If functionalism were true, then for the reasons you gave about global workspace, it seems very plausible that LLMs are conscious anyway. I just wanted to address this specific point.)
I think this argument agrees with this line of reasoning which you made...
A philosophical argument as to why Claude might be mistaken in its belief that it’s conscious is that non-conscious beings cannot actually know what consciousness is. In this scenario, Claude is non-conscious, believes that phenomenal consciousness is equivalent to some functional conception of consciousness, and thus wrongly believes that it is conscious. Further discussion of this quickly gets philosophically complicated, so I’ll just say that I do find this plausible.
...which to me would support the idea that a lack of introspective capability might be evidence against consciousness, but the presence of it seems roughly neutral.
I’m curious why you think introspective capabilities are strong evidence for consciousness.
Humans do not become conscious by learning the word consciousness—for example, we feel pain and see colors as children long before we have any concept for them. Only later do we learn to group these experiences under a label like “conscious experience” or “being aware”; and only later still do we notice that other people seem to have similar inner lives.
Contrast this with a LLM. A model can learn the use of the word “consciousness” perfectly well. It can learn that humans associate the term with reports like “I feel pain,” “I am aware,” “there is something it is like to be me,” or “I am having a subjective experience.” It can also learn to map those terms onto its own internal organization. For instance, the model might identify certain nested internal states in its ‘global workspace’ and claim, “These are my conscious states,” just as it might say, “This activation pattern corresponds to uncertainty” or “that subsystem is attention.” But if the LM has no first-person reference point for consciousness, it might still apply the term to its own states in a purely learned or stipulative way. It might say, in effect, “These are the states that best play the role humans associate with consciousness, so I will call them conscious” (e.g. labeling a cluster of internal states “pain” because those states are caused by damage signals and lead to avoidance behavior, without any sense of felt painfulness). That shows that the LLM has learned the concept’s functional role in human discourse, not that the word tracks any actual subjective experience.
(All of this presupposes the hard problem is real and functionalism is false. If functionalism were true, then for the reasons you gave about global workspace, it seems very plausible that LLMs are conscious anyway. I just wanted to address this specific point.)
I think this argument agrees with this line of reasoning which you made...
...which to me would support the idea that a lack of introspective capability might be evidence against consciousness, but the presence of it seems roughly neutral.