Frankly, having an eval for something is the first step to get really good at it.
What if your eval finds out that all the Chinese models lag far behind the American models when it comes to drones? Or vice versa?
Frankly, having an eval for something is the first step to get really good at it.
What if your eval finds out that all the Chinese models lag far behind the American models when it comes to drones? Or vice versa?
But overattributing (or figuring out that AIs do have consciousness etc) seems like a necessary first step.
I am also not sure how long that way is.
Kids and animals are clearly not on the same intellectual level as adults, part of the reason why they lack rights is that they would not be able to use them.
AI rights pattern match on other rights movements and there are already plenty of zealots ready to take up that cause.
As I said, you are taking one sentence out of the context of comparing the argument for reasoning (valid) with the argument for emotions (not valid).
I understand that this one sentence can be read as stronger than it was intended if taken out of context. Well, just don’t do that.
Like, just read what I wrote: At no point do I state the AIs do not have emotions or are not conscious.
I do have opinions on both of these questions. These opinions are based on the architecture of the models, their training and the structure of the human brain.
You do not know and I bet you would not be able to correctly predict my opinions on these questions because nothing I wrote in this thread was about that.
You yourself said that you can imitate emotions from an intellectual, non-experimential understanding. I am not claiming anything beyond that. It’s not a strong or controversial statement at all.
I think you are taking my statement out of context, which is: Outward shows of emotion do not prove the inner experience of emotions.
The statement refers to LLMs that are trained to imitate human output.
I agree that method acting exists. I don’t think it is the only way for a skilled (human) actor (or writer) to imitate emotions or pain. Especially if the output channel is text.
A world in which models have all the rights a human has (or even just some of them, let’s say the right to own property), is a world where they compete directly against humans and will eventually outcompete humans.
The main cost of overattributing is that it makes human extinction much more likely, imho.
Models are explicitly trained to imitate human generated text. There is absolutely nothing misleading about it, it is the single most relevant fact about LLMs.
In all the human generated text the human thinks (and if that comes up, expresses the idea) that it is conscious. So almost all roles an LLM might simulate have “I am conscious” as a basic fact. Finetuning pushes LLMs to a specific assistant role which inherits that fact. There is no reason why RL (for math and code mostly) would change that.
An LLM is nothing before it is filled with the data from human generated text. Daniel Radcliffe on the other hand is a human with his own life and memories. If you’d wipe his brain and actually train it to “imitate Harry Potter” he would think that his parents were killed by Voldemort.
I think if you don’t feel the emotion you don’t have it.
If you shout “oh my god, it’s a bear” in a scared voice and then run, you are certainly representing fear in your brain and it’s also coherent with your behaviour (what I think you call “functional”), but if your amygdala is not firing your are not “having” the emotion fear.
We know from humans that understanding fear or pain and being able to act like you are in fear or pain is a pure sequence learning thing and it can be completely separate from actually being in fear and pain.
Actually being in fear and pain requires additional machinery and some humans don’t have it. Understanding and acting doesn’t replace it.
1. For all functional purposes of the words “think” and “reason” and “have emotions”, they think and reason and have emotions.
If you imitate reasoning and solve more problems that way than without imitating reasoning, you’re not just imitating reasoning, you are reasoning. But if you imitate how a human with certain emotions would act you are not having those emotions, you are just acting.
Same with point 6.): Models think they are conscious because they are trained to imitate humans and humans think they are conscious.
It’s wild that this is possible while at the same time my impression with Fable and GPT-5.5 (and predecessors) is that they say roughly one dumb thing per output.
Here’s one idea, which is surprisingly completely unstudied. During training, tokens in input-only roles (
<user>,<tool>) are loss-masked: the LLM never has to predict the next token at those positions, so their activations focus entirely on comprehension instead of generation
In finetuning in my experience the problem with this setup is generally that the models learn comprehension much better if they are trained on generation. Activations that do not predict the next token are pretty out of distribution.
Because LLMs often confuse my input and their own output I have long wondered whether it might be a good idea to put the tag-processing into the harness and use specific experts (in a MoE) or LoRAs for the tokens in each tag.
That way one can train against instruction following within tool-tags without hurting instruction following generally etc.
Your work makes that sound like an even better idea.
That was also my first assumption. Interestingly Gemma doesn’t have this.
Yeah, it seems independent of the actual topic and also of the position in the generated text. My first assumption was that the vague representations of earlier layers might be mapped to tokens most heavily repressed by RLHF or something like that. But I have to look at the methods more closely to see whether that even makes sense.
Was this trained on the internet or something?!
tfw your llm spends 1/8th of your compute to think about porn before even tackling the task.
And what do we make of this?
My impression is that Ed Zitron agrees with each and every AI criticism, which makes it hard to see whether he might have a point somewhere. I would love to see somebody check his analyses in detail.
The point was that in that scenario the lagging side might make a deliberate effort to catch up turning something that was on nobody’s mind into a race.