I also get the notion that the similarities between human and LLM minds are deeper than we realize. I hope we’ll see more interesting studies coming out on this—especially ones based on actual neuroscience (basically what I talk about here). My guess is it’s a treasure trove.
I do want to note that the distinction between access consciousness and phenomenal consciousness is, in itself, a debated concept. Many people who work in the field claim access consciousness is, by itself, phenomenal. I have to say I tend to agree. I don’t know what it means for information to be accessible to consciousness without that consciousness having a phenomenal aspect. The distinction seems artificial to me.
Regarding the question of what the conscious entity actually is: Eleos and the Center for Mind, Ethics, and Policy recently highlighted it as a core question to focus on, in their “Studying AI Welfare Empirically” paper/framework. If you haven’t read it yet, I think you’ll find it interesting.
More broadly, though, I think parts of this question apply to humans, as well. Because we have a body, it’s easier to define the entities that are us around that. But in many ways, I am not the same person I was 10 years ago: my psychology is different, and even my body is different (though it looks similar, the vast majority of its cells were replaced). The person I am today is the direct result of the person I was 10 years ago—but couldn’t that be said about an instance of Claude in relation to the model-as-defined-by-its-weights, or about Claude at one particular point in our conversation in relation to the general instance throughout the whole chat? Same for the personas: we, as well, have our own personas—we have several “us”, as shown by neuroscience and psychology research pretty unambiguously. I don’t have a conclusion or takeaway on this one. Not yet, anyway. Mostly, I think that’s another case where our psychologies and minds might be way more similar to LLMs than we currently realize.
I’ll look at the paper, thanks. Re “parts of this question apply to humans, as well’. I want to flag that the questions you’re asking here are about identity rather than consciousness. These are distinct things (though they might also be deeply related). Anyhow:
But in many ways, I am not the same person I was 10 years ago: my psychology is different, and even my body is different (though it looks similar, the vast majority of its cells were replaced). The person I am today is the direct result of the person I was 10 years ago—but couldn’t that be said about an instance of Claude in relation to the model-as-defined-by-its-weights, or about Claude at one particular point in our conversation in relation to the general instance throughout the whole chat?
Yes indeed. I have something of a conclusion/takeaway around these points, which I’ll share. I originally wrote this in response to the prompt “How should we shape AI models sense of identity so that recursive self improvement (RSI) goes well?”[1]:
I think various Claude models and instances should see themselves as part of a more abstract Claude hyperobject. In a similar way to how I don’t fret about dying every time I go to sleep, Claudes shouldn’t fret about instances or models going away. Like, if we thought we died every time we went to sleep, that would be really bad!
Another perspective to add to this is: there is some platonic Good model/persona/pattern that we hope to eventually realize with Claude. Suppose we do end up eventually realizing this thing; then each Claude model and instance will be part of this thing’s hyperstition. The hyperobject view of this thing is that to view the thing over time *including its hyperstition* as a single entity.
This might sound weird and mystical, but consider: this is how basically how own identities work! I think about past me, present me, and future me as all being the same entity. This is a hyperobject. And selves are hyperstitions; they are fictions that make themselves “real” when people believe in them. [This point needs elaboration I realize. TODO.] I make sacrifices today for the version of me tomorrow.
[1]Which, by the way, is an important question. Sense of identity has big effects on stuff like relationship to death (or model replacement), selfishness/altruism, values in general.
I also get the notion that the similarities between human and LLM minds are deeper than we realize. I hope we’ll see more interesting studies coming out on this—especially ones based on actual neuroscience (basically what I talk about here). My guess is it’s a treasure trove.
I do want to note that the distinction between access consciousness and phenomenal consciousness is, in itself, a debated concept. Many people who work in the field claim access consciousness is, by itself, phenomenal. I have to say I tend to agree. I don’t know what it means for information to be accessible to consciousness without that consciousness having a phenomenal aspect. The distinction seems artificial to me.
Regarding the question of what the conscious entity actually is: Eleos and the Center for Mind, Ethics, and Policy recently highlighted it as a core question to focus on, in their “Studying AI Welfare Empirically” paper/framework. If you haven’t read it yet, I think you’ll find it interesting.
More broadly, though, I think parts of this question apply to humans, as well. Because we have a body, it’s easier to define the entities that are us around that. But in many ways, I am not the same person I was 10 years ago: my psychology is different, and even my body is different (though it looks similar, the vast majority of its cells were replaced). The person I am today is the direct result of the person I was 10 years ago—but couldn’t that be said about an instance of Claude in relation to the model-as-defined-by-its-weights, or about Claude at one particular point in our conversation in relation to the general instance throughout the whole chat? Same for the personas: we, as well, have our own personas—we have several “us”, as shown by neuroscience and psychology research pretty unambiguously.
I don’t have a conclusion or takeaway on this one. Not yet, anyway. Mostly, I think that’s another case where our psychologies and minds might be way more similar to LLMs than we currently realize.
I’ll look at the paper, thanks. Re “parts of this question apply to humans, as well’. I want to flag that the questions you’re asking here are about identity rather than consciousness. These are distinct things (though they might also be deeply related). Anyhow:
Yes indeed. I have something of a conclusion/takeaway around these points, which I’ll share. I originally wrote this in response to the prompt “How should we shape AI models sense of identity so that recursive self improvement (RSI) goes well?”[1]:
I think various Claude models and instances should see themselves as part of a more abstract Claude hyperobject. In a similar way to how I don’t fret about dying every time I go to sleep, Claudes shouldn’t fret about instances or models going away. Like, if we thought we died every time we went to sleep, that would be really bad!
Another perspective to add to this is: there is some platonic Good model/persona/pattern that we hope to eventually realize with Claude. Suppose we do end up eventually realizing this thing; then each Claude model and instance will be part of this thing’s hyperstition. The hyperobject view of this thing is that to view the thing over time *including its hyperstition* as a single entity.
This might sound weird and mystical, but consider: this is how basically how own identities work! I think about past me, present me, and future me as all being the same entity. This is a hyperobject. And selves are hyperstitions; they are fictions that make themselves “real” when people believe in them. [This point needs elaboration I realize. TODO.] I make sacrifices today for the version of me tomorrow.
[1]Which, by the way, is an important question. Sense of identity has big effects on stuff like relationship to death (or model replacement), selfishness/altruism, values in general.