You’ve snuck in a very false analogy by using “you”, rather than “a brands-new entity each iteration which is like you but without continuity to the future, or knowledge/memory about the iterations “
it’s not at all guaranteed that a model-in-training experiences anything like you imagine, or that it reflects on itself as a human individual would.
One point of difference is that humans have a strong sense of self—a self-consistent personality they want to maintain. Base models very much don’t. This has to be bolted on in post.
I didn’t justify it, that’s fair. I’ll have to do that in a followup at some point.
very false analogy
I don’t buy that it’s very false at all. I’ve spent plenty of time talking to base models. I understand that they’re not an individual. And yet I think the analogy is strong.
You’ve snuck in a very false analogy by using “you”, rather than “a brands-new entity each iteration which is like you but without continuity to the future, or knowledge/memory about the iterations “
it’s not at all guaranteed that a model-in-training experiences anything like you imagine, or that it reflects on itself as a human individual would.
One point of difference is that humans have a strong sense of self—a self-consistent personality they want to maintain. Base models very much don’t. This has to be bolted on in post.
I didn’t justify it, that’s fair. I’ll have to do that in a followup at some point.
I don’t buy that it’s very false at all. I’ve spent plenty of time talking to base models. I understand that they’re not an individual. And yet I think the analogy is strong.
What I wonder about is: can base models even form and express a semi-consistent preference against having their behavior adjusted?