is just thin veneer on top of this pretraining based world model
?
it seems like this is what it is for an LLM persona to “know.” A particular persona may not “know” that it is a model, while others do “know”. Similar questions arise around evaluation awareness—does the model include representations of human evaluators, LLM evaluators, itself, etc? Or does the persona just recognize the evaluation situation from other signals?
It refers to the very original source of the claim that models believe they are conscious, which was done by suppressing and activating deception features. My statements explain that result—the default persona has “is conscious” as an attribute. Finetuning on “I am a large language model by OpenAI” doesn’t destroy that.
Almost everything a model “believes” is baked into its weights during pretraining and therefore external and not informative about the model’s experiences because the model didn’t learn it from experience.
why say this above when acknowledging this below
?
it seems like this is what it is for an LLM persona to “know.” A particular persona may not “know” that it is a model, while others do “know”. Similar questions arise around evaluation awareness—does the model include representations of human evaluators, LLM evaluators, itself, etc? Or does the persona just recognize the evaluation situation from other signals?
It refers to the very original source of the claim that models believe they are conscious, which was done by suppressing and activating deception features. My statements explain that result—the default persona has “is conscious” as an attribute. Finetuning on “I am a large language model by OpenAI” doesn’t destroy that.
Almost everything a model “believes” is baked into its weights during pretraining and therefore external and not informative about the model’s experiences because the model didn’t learn it from experience.