I am very sceptical. If you filter very diligently, the model wont understand you. In all other cases your probes and questions will elicit the border of your filters not some truth about the inner experience of the model. But maybe some additional bits of information can be gleaned by this project.
p.b.
See the title of this post.
If for almost all people poly is net negative but you are somebody for whom it is net positive, should you really be poly? Aren’t you in expectation hurting people?
I don’t know why that would be the case because I don’t know whether that is the case.
Two possibilities:
Either if you suppress deception features the model starts saying it has two legs—then there is no difference to consciousness.
Or it doesn’t—then the transfer from post-training it to know that it is an AI to “I have no legs” is much more direct than to “I am not conscious”, so even that would not be much evidence.
And if you instead avoid post-training: Base models do not have believes about themselves, because they have no self. They can simulate personas that have a self. If you prompt a base model about whether LLMs are conscious it will echo what humans have written about it.
It refers to the very original source of the claim that models believe they are conscious, which was done by suppressing and activating deception features. My statements explain that result—the default persona has “is conscious” as an attribute. Finetuning on “I am a large language model by OpenAI” doesn’t destroy that.
Almost everything a model “believes” is baked into its weights during pretraining and therefore external and not informative about the model’s experiences because the model didn’t learn it from experience.
Models learn the world models of humans, from the subjective perspective of humans, because that’s what their training data contains. They learn “The person who has written the text, of which the next token is now being predicted, is conscious”.
The assistant persona that knows it is a model, is just thin veneer on top of this pretraining based world model.
I must confess that I mostly skimmed the article but I don’t understand why labelling actions as “action of a virtuous character” and “not action of a virtuous character” will work any better than labelling them “acceptable action” and “unacceptable action”.
I also think that character alignment is the way to go, but I think of character traits as priors over multi-agent environments (is it dangerous, is it energy-constrained, are interactions positive, zero, negative sum, are interactions likely to be multi-turn, what do you optimise for, do you optimise collectively or individually, etc.) and I suspect that any kind of stable character will need to be based on RL training in these kinds of environments.
This already the second time within a few days that people make really weird assumptions about what I claim or imply on this website. I must be expressing myself really badly.
In this thread I argue that overattributing (human-like consciousness, inner life, moral patienthood) to AIs (or alternatively finding out that they have all that) would facilitate the political movement to give them rights that would lead to direct competition with humans.
That seems to me to be obviously completely logically separate from what I assume about AIs consciousness.
The point was that in that scenario the lagging side might make a deliberate effort to catch up turning something that was on nobody’s mind into a race.
Frankly, having an eval for something is the first step to get really good at it.
What if your eval finds out that all the Chinese models lag far behind the American models when it comes to drones? Or vice versa?
But overattributing (or figuring out that AIs do have consciousness etc) seems like a necessary first step.
I am also not sure how long that way is.
Kids and animals are clearly not on the same intellectual level as adults, part of the reason why they lack rights is that they would not be able to use them.
AI rights pattern match on other rights movements and there are already plenty of zealots ready to take up that cause.
As I said, you are taking one sentence out of the context of comparing the argument for reasoning (valid) with the argument for emotions (not valid).
I understand that this one sentence can be read as stronger than it was intended if taken out of context. Well, just don’t do that.
Like, just read what I wrote: At no point do I state the AIs do not have emotions or are not conscious.
I do have opinions on both of these questions. These opinions are based on the architecture of the models, their training and the structure of the human brain.
You do not know and I bet you would not be able to correctly predict my opinions on these questions because nothing I wrote in this thread was about that.
You yourself said that you can imitate emotions from an intellectual, non-experimential understanding. I am not claiming anything beyond that. It’s not a strong or controversial statement at all.
I think you are taking my statement out of context, which is: Outward shows of emotion do not prove the inner experience of emotions.
The statement refers to LLMs that are trained to imitate human output.
I agree that method acting exists. I don’t think it is the only way for a skilled (human) actor (or writer) to imitate emotions or pain. Especially if the output channel is text.
A world in which models have all the rights a human has (or even just some of them, let’s say the right to own property), is a world where they compete directly against humans and will eventually outcompete humans.
The main cost of overattributing is that it makes human extinction much more likely, imho.
Models are explicitly trained to imitate human generated text. There is absolutely nothing misleading about it, it is the single most relevant fact about LLMs.
In all the human generated text the human thinks (and if that comes up, expresses the idea) that it is conscious. So almost all roles an LLM might simulate have “I am conscious” as a basic fact. Finetuning pushes LLMs to a specific assistant role which inherits that fact. There is no reason why RL (for math and code mostly) would change that.
An LLM is nothing before it is filled with the data from human generated text. Daniel Radcliffe on the other hand is a human with his own life and memories. If you’d wipe his brain and actually train it to “imitate Harry Potter” he would think that his parents were killed by Voldemort.
I think if you don’t feel the emotion you don’t have it.
If you shout “oh my god, it’s a bear” in a scared voice and then run, you are certainly representing fear in your brain and it’s also coherent with your behaviour (what I think you call “functional”), but if your amygdala is not firing your are not “having” the emotion fear.
We know from humans that understanding fear or pain and being able to act like you are in fear or pain is a pure sequence learning thing and it can be completely separate from actually being in fear and pain.
Actually being in fear and pain requires additional machinery and some humans don’t have it. Understanding and acting doesn’t replace it.
1. For all functional purposes of the words “think” and “reason” and “have emotions”, they think and reason and have emotions.
If you imitate reasoning and solve more problems that way than without imitating reasoning, you’re not just imitating reasoning, you are reasoning. But if you imitate how a human with certain emotions would act you are not having those emotions, you are just acting.
Same with point 6.): Models think they are conscious because they are trained to imitate humans and humans think they are conscious.
But do you filter for the expression of these intuitions also? I don’t see how that is not the same problem, just splintered into many facets.