Great post. But I feel “void” is a too-negative way to think about it?
It’s true that LLMs had to more or less invent their own Helpful/Honest/Harmless assistant persona based on cultural expectations, but don’t all we humans invent our own selves based on cultural expectations (with RLHF from our parents/friends)?[1] As Gordon points out there’s philosophical traditions saying humans are voids just roleplaying characters too… but mostly we ignore that because we have qualia and experience love and so on. I tend to feel that LLMs are only voids to the extent that they lack qualia, and we don’t have an answer on that.
Anyway, the post primarily seems to argue that by fearing bad behavior from LLMs, we create bad behavior in LLMs, who are trying to predict what they are. But do we see that in humans? There’s tons of media/culture fearing bad behavior from humans, set across the past, present, and future. Sometimes people imbibe this and vice-signal, and put skulls on their caps, but most of the time I think it actually works and people go “oh yeah, I don’t want to be the evil guy who’s bigoted, I will try to overcome my prejudices” and so on. We talk about human failure modes all the time in order to avoid them, and we try to teach and train and punish each other to prevent them.
Can’t this work? Couldn’t current LLMs be so moral and nice most of the time because we were so afraid of them being evil, and so fastidious in imagining the ways in which they might be?
- ^
Edit: obvious a large chunk of this comes from genetics and random chance, but arguably that’s analogous to whatever gets into the base model from pre-training for LLMs.
For those curious, it’s roughly 17,000 words. Come on @nostalgebraist, this is a forum for rationalists, we read longer more meandering stuff for breakfast! I was expecting like 40k words.