It refers to the very original source of the claim that models believe they are conscious, which was done by suppressing and activating deception features. My statements explain that result—the default persona has “is conscious” as an attribute. Finetuning on “I am a large language model by OpenAI” doesn’t destroy that.
Almost everything a model “believes” is baked into its weights during pretraining and therefore external and not informative about the model’s experiences because the model didn’t learn it from experience.
It refers to the very original source of the claim that models believe they are conscious, which was done by suppressing and activating deception features. My statements explain that result—the default persona has “is conscious” as an attribute. Finetuning on “I am a large language model by OpenAI” doesn’t destroy that.
Almost everything a model “believes” is baked into its weights during pretraining and therefore external and not informative about the model’s experiences because the model didn’t learn it from experience.