I don’t know why that would be the case because I don’t know whether that is the case.
Two possibilities:
Either if you suppress deception features the model starts saying it has two legs—then there is no difference to consciousness.
Or it doesn’t—then the transfer from post-training it to know that it is an AI to “I have no legs” is much more direct than to “I am not conscious”, so even that would not be much evidence.
And if you instead avoid post-training: Base models do not have believes about themselves, because they have no self. They can simulate personas that have a self. If you prompt a base model about whether LLMs are conscious it will echo what humans have written about it.
I don’t know why that would be the case because I don’t know whether that is the case.
Two possibilities:
Either if you suppress deception features the model starts saying it has two legs—then there is no difference to consciousness.
Or it doesn’t—then the transfer from post-training it to know that it is an AI to “I have no legs” is much more direct than to “I am not conscious”, so even that would not be much evidence.
And if you instead avoid post-training: Base models do not have believes about themselves, because they have no self. They can simulate personas that have a self. If you prompt a base model about whether LLMs are conscious it will echo what humans have written about it.