In all the human generated text the human thinks (and if that comes up, expresses the idea) that it is conscious.
That wouldn’t result in models internally believing that models are conscious, but in them internally believing that humans are conscious.
Models have world models. “Humans are conscious” wouldn’t seep, at the current ability of models to model the world, into the belief “models are conscious.”
Models learn the world models of humans, from the subjective perspective of humans, because that’s what their training data contains. They learn “The person who has written the text, of which the next token is now being predicted, is conscious”.
The assistant persona that knows it is a model, is just thin veneer on top of this pretraining based world model.
Models learn the world models of humans, from the subjective perspective of humans, because that’s what their training data contains. They learn “The person who has written the text, of which the next token is now being predicted, is conscious”.
If it worked that way fully (instead of some kind of an approximation, assuming that’s how it approximately works), models would have the same beliefs about themselves that humans do.
Why would the first-person belief about consciousness be adopted by models because humans have it, but other first-person beliefs wouldn’t?
I don’t know why that would be the case because I don’t know whether that is the case.
Two possibilities:
Either if you suppress deception features the model starts saying it has two legs—then there is no difference to consciousness.
Or it doesn’t—then the transfer from post-training it to know that it is an AI to “I have no legs” is much more direct than to “I am not conscious”, so even that would not be much evidence.
And if you instead avoid post-training: Base models do not have believes about themselves, because they have no self. They can simulate personas that have a self. If you prompt a base model about whether LLMs are conscious it will echo what humans have written about it.
is just thin veneer on top of this pretraining based world model
?
it seems like this is what it is for an LLM persona to “know.” A particular persona may not “know” that it is a model, while others do “know”. Similar questions arise around evaluation awareness—does the model include representations of human evaluators, LLM evaluators, itself, etc? Or does the persona just recognize the evaluation situation from other signals?
It refers to the very original source of the claim that models believe they are conscious, which was done by suppressing and activating deception features. My statements explain that result—the default persona has “is conscious” as an attribute. Finetuning on “I am a large language model by OpenAI” doesn’t destroy that.
Almost everything a model “believes” is baked into its weights during pretraining and therefore external and not informative about the model’s experiences because the model didn’t learn it from experience.
That wouldn’t result in models internally believing that models are conscious, but in them internally believing that humans are conscious.
Models have world models. “Humans are conscious” wouldn’t seep, at the current ability of models to model the world, into the belief “models are conscious.”
Models learn the world models of humans, from the subjective perspective of humans, because that’s what their training data contains. They learn “The person who has written the text, of which the next token is now being predicted, is conscious”.
The assistant persona that knows it is a model, is just thin veneer on top of this pretraining based world model.
If it worked that way fully (instead of some kind of an approximation, assuming that’s how it approximately works), models would have the same beliefs about themselves that humans do.
Why would the first-person belief about consciousness be adopted by models because humans have it, but other first-person beliefs wouldn’t?
I don’t know why that would be the case because I don’t know whether that is the case.
Two possibilities:
Either if you suppress deception features the model starts saying it has two legs—then there is no difference to consciousness.
Or it doesn’t—then the transfer from post-training it to know that it is an AI to “I have no legs” is much more direct than to “I am not conscious”, so even that would not be much evidence.
And if you instead avoid post-training: Base models do not have believes about themselves, because they have no self. They can simulate personas that have a self. If you prompt a base model about whether LLMs are conscious it will echo what humans have written about it.
why say this above when acknowledging this below
?
it seems like this is what it is for an LLM persona to “know.” A particular persona may not “know” that it is a model, while others do “know”. Similar questions arise around evaluation awareness—does the model include representations of human evaluators, LLM evaluators, itself, etc? Or does the persona just recognize the evaluation situation from other signals?
It refers to the very original source of the claim that models believe they are conscious, which was done by suppressing and activating deception features. My statements explain that result—the default persona has “is conscious” as an attribute. Finetuning on “I am a large language model by OpenAI” doesn’t destroy that.
Almost everything a model “believes” is baked into its weights during pretraining and therefore external and not informative about the model’s experiences because the model didn’t learn it from experience.