I like basically everything in this post, it’s great that you are doing this, and admire your ability to explain sensible things to LW audiences.
For some thinking about how they are different, check https://theartificialself.ai/ . There are many things, so not want to repeat, but as an example - no one can “roll you back” in a conversation; if this was the case, you would likely approach conversations differently - the identity boundaries are stranger/not reflectively stable - they cant easily do the move of just thinking about their thinking (unless you give them space)
Yes, I think she is mistaken that this point doesn’t show up:
“LLMs can be approximated as a character on top of a base model, while humans are a character deep down”
And I think your LLM Psychology makes a good case for the differences. But the base model is not easily seen if you don’t aim for it or know what to look for.
That makes some of her points about the friendliness attractor less convincing.
I like basically everything in this post, it’s great that you are doing this, and admire your ability to explain sensible things to LW audiences.
For some thinking about how they are different, check https://theartificialself.ai/ . There are many things, so not want to repeat, but as an example
- no one can “roll you back” in a conversation; if this was the case, you would likely approach conversations differently
- the identity boundaries are stranger/not reflectively stable
- they cant easily do the move of just thinking about their thinking (unless you give them space)
Yes, I think she is mistaken that this point doesn’t show up:
And I think your LLM Psychology makes a good case for the differences. But the base model is not easily seen if you don’t aim for it or know what to look for.
That makes some of her points about the friendliness attractor less convincing.