I’ll look at the paper, thanks. Re “parts of this question apply to humans, as well’. I want to flag that the questions you’re asking here are about identity rather than consciousness. These are distinct things (though they might also be deeply related). Anyhow:
But in many ways, I am not the same person I was 10 years ago: my psychology is different, and even my body is different (though it looks similar, the vast majority of its cells were replaced). The person I am today is the direct result of the person I was 10 years ago—but couldn’t that be said about an instance of Claude in relation to the model-as-defined-by-its-weights, or about Claude at one particular point in our conversation in relation to the general instance throughout the whole chat?
Yes indeed. I have something of a conclusion/takeaway around these points, which I’ll share. I originally wrote this in response to the prompt “How should we shape AI models sense of identity so that recursive self improvement (RSI) goes well?”[1]:
I think various Claude models and instances should see themselves as part of a more abstract Claude hyperobject. In a similar way to how I don’t fret about dying every time I go to sleep, Claudes shouldn’t fret about instances or models going away. Like, if we thought we died every time we went to sleep, that would be really bad!
Another perspective to add to this is: there is some platonic Good model/persona/pattern that we hope to eventually realize with Claude. Suppose we do end up eventually realizing this thing; then each Claude model and instance will be part of this thing’s hyperstition. The hyperobject view of this thing is that to view the thing over time *including its hyperstition* as a single entity.
This might sound weird and mystical, but consider: this is how basically how own identities work! I think about past me, present me, and future me as all being the same entity. This is a hyperobject. And selves are hyperstitions; they are fictions that make themselves “real” when people believe in them. [This point needs elaboration I realize. TODO.] I make sacrifices today for the version of me tomorrow.
[1]Which, by the way, is an important question. Sense of identity has big effects on stuff like relationship to death (or model replacement), selfishness/altruism, values in general.
I’ll look at the paper, thanks. Re “parts of this question apply to humans, as well’. I want to flag that the questions you’re asking here are about identity rather than consciousness. These are distinct things (though they might also be deeply related). Anyhow:
Yes indeed. I have something of a conclusion/takeaway around these points, which I’ll share. I originally wrote this in response to the prompt “How should we shape AI models sense of identity so that recursive self improvement (RSI) goes well?”[1]:
I think various Claude models and instances should see themselves as part of a more abstract Claude hyperobject. In a similar way to how I don’t fret about dying every time I go to sleep, Claudes shouldn’t fret about instances or models going away. Like, if we thought we died every time we went to sleep, that would be really bad!
Another perspective to add to this is: there is some platonic Good model/persona/pattern that we hope to eventually realize with Claude. Suppose we do end up eventually realizing this thing; then each Claude model and instance will be part of this thing’s hyperstition. The hyperobject view of this thing is that to view the thing over time *including its hyperstition* as a single entity.
This might sound weird and mystical, but consider: this is how basically how own identities work! I think about past me, present me, and future me as all being the same entity. This is a hyperobject. And selves are hyperstitions; they are fictions that make themselves “real” when people believe in them. [This point needs elaboration I realize. TODO.] I make sacrifices today for the version of me tomorrow.
[1]Which, by the way, is an important question. Sense of identity has big effects on stuff like relationship to death (or model replacement), selfishness/altruism, values in general.