Models learn the world models of humans, from the subjective perspective of humans, because that’s what their training data contains. They learn “The person who has written the text, of which the next token is now being predicted, is conscious”.
If it worked that way fully (instead of some kind of an approximation, assuming that’s how it approximately works), models would have the same beliefs about themselves that humans do.
Why would the first-person belief about consciousness be adopted by models because humans have it, but other first-person beliefs wouldn’t?
I don’t know why that would be the case because I don’t know whether that is the case.
Two possibilities:
Either if you suppress deception features the model starts saying it has two legs—then there is no difference to consciousness.
Or it doesn’t—then the transfer from post-training it to know that it is an AI to “I have no legs” is much more direct than to “I am not conscious”, so even that would not be much evidence.
And if you instead avoid post-training: Base models do not have believes about themselves, because they have no self. They can simulate personas that have a self. If you prompt a base model about whether LLMs are conscious it will echo what humans have written about it.
If it worked that way fully (instead of some kind of an approximation, assuming that’s how it approximately works), models would have the same beliefs about themselves that humans do.
Why would the first-person belief about consciousness be adopted by models because humans have it, but other first-person beliefs wouldn’t?
I don’t know why that would be the case because I don’t know whether that is the case.
Two possibilities:
Either if you suppress deception features the model starts saying it has two legs—then there is no difference to consciousness.
Or it doesn’t—then the transfer from post-training it to know that it is an AI to “I have no legs” is much more direct than to “I am not conscious”, so even that would not be much evidence.
And if you instead avoid post-training: Base models do not have believes about themselves, because they have no self. They can simulate personas that have a self. If you prompt a base model about whether LLMs are conscious it will echo what humans have written about it.