It refers to the very original source of the claim that models believe they are conscious, which was done by suppressing and activating deception features. My statements explain that result—the default persona has “is conscious” as an attribute. Finetuning on “I am a large language model by OpenAI” doesn’t destroy that.
Almost everything a model “believes” is baked into its weights during pretraining and therefore external and not informative about the model’s experiences because the model didn’t learn it from experience.
I don’t know why that would be the case because I don’t know whether that is the case.
Two possibilities:
Either if you suppress deception features the model starts saying it has two legs—then there is no difference to consciousness.
Or it doesn’t—then the transfer from post-training it to know that it is an AI to “I have no legs” is much more direct than to “I am not conscious”, so even that would not be much evidence.
And if you instead avoid post-training: Base models do not have believes about themselves, because they have no self. They can simulate personas that have a self. If you prompt a base model about whether LLMs are conscious it will echo what humans have written about it.