I am from Russia (not in local lw.ru). I have read HPMoR at age 12 and it immediately became my favourite book and holds it’s place even now. I find Eliezer Yudkowsky’s views on most things the most convincing of all the other I saw.
More detailed version organized as Shortform comments thread: https://www.lesswrong.com/posts/zbSsSwEfdEuaqCRmz/eniscien-s-shortform?commentId=JtpxcmxMt2ycd4K5s
I agree with the basic idea that requirements for llm tests are absurdly strict to the degree that the testers themself likely wouldn’t pass them and end up with concluding that they demonstrated “no signs of consciousness or self recognition”.
Humans also hear the consciousness talk all the time, human brain naturally mimics speech pattern of other humans, so if you will use that as an excuse to explain away llm behaviour, you ought to use that to humans as well, and then the only line of reasoning at implying humans are conscious is based on first person perspective (I can see that I am conscious) plus expectation that consciousness is a physical phenomena and therefore will be similarly exist or not for all humans because we have very similar brain structures.
However, the clearly proved outer functional self awareness in mirror tests (and many other tests) doesn’t of course prove mirroring of inner mental mechanism of this self awareness (if we have thought like that, we would implied that llm outward good behaviour also implies inner good thoughts, wishes etc!), and even more doesn’t tell about consciousness.
There certainly exists such factor that we don’t really know how self awareness works in animal brain or even human brains. We know though that humans often confabulate their “introspections” (the most notorious example is cut hemispheres experiments, of course) and which is quite an evidence for previous “humans mimic what they heard and therefore you can’t infer they are conscious from their talk about consciousness”. It is clear that plenty times humans “introspect”, “report conscious experiences” they actually just confabulate likely explanations quite like llm confabulate next tokens. So, we actually know that having sometimes confabulating consciousness talk does not imply no consciousness.
None of us is llm so we can’t know the same way we know about humans, because llm architecture is too different, while human architecture is really similar because of evolutionary reasons. Though, people often lack some experiences, for example imagination (purely mental sensory experience) in aphantasia, and those people mimick usual people so closely that usually can’t be distinguished even after they start directly talking like “what mental pictures you are implying, there is none such, that is just a metaphor”, so I personally can’t rule out that some people lack conscious experiences, treat it as metaphor and mimic others behaviour through some side ways.
I am not sure how I should proceed. Llm are taught on reflection and consciousness examples as well as humans, they are (almost) certainly confabulate new such as well as humans. But we don’t have personal, unsharable, anthropic, point of view data from an architecture similar enough to LLM, while we do for humans. Llm being able to confabulate it should certainly increase likelihood of hypothesis that it explained by confabulations, because at least some of it certainly is. But I hugely doubt that it implies that we should because of it think that llm has same or less probability of no consciousness, as if we didn’t know about confabulations in architecture, but llm also didn’t demonstrated any such signs as mirror test etc. It seems to me they still should make it more likely in general that llm have inner self awareness and possibly consciousness.
And certainly it seems to me absurd when the same people simultaneously confidently say how we know about not just self awareness, but (!) consciousness of animals, and also confidently say how we know about llm not being self aware or conscious.
So I think it is good to point out that tests of same strictness would as well prove non self awareness and unconscious not just animals, but humans too. This is not a good test if it can’t show different results for different cases, and for most humans it would show they are not conscious despite they are. And conclusion is even worse. Because of connotations. Like getting 9 years old, asking him to discover relativity theory in a month, and that it without other humans or internet, because we tested them with it and they casually passed the test just in 5% of allocated time, and when they fail you draw conclusion that… “9 year olds didn’t show any signs of intelligence.” (like if adults would show any such signs) More correct conclusion for consciousness would be that consciousness and self awareness signs weren’t should to be vastly exceeding those for humans (and therefore we can not conclude whether it is their own consciousness or copy of that of humans).
P. S./Edited: I have a meta about the post form. While reading I anticipated it will be likely you will not make an obvious caveat that it doesn’t mean llm didn’t just outputting likely token in very similar manner to how they do with all the other tokens, with no relation something specific to actually recognizing self and just being a kind of echo matcher, or trigger on text repetition or something (which it seems to me you didn’t give), or even imply that actually simple mirror test should prove llm are self aware and conscious (which to my relief you didn’t say). So I felt an urge to add that caveat for the case of reader missing it. And it seems other commenters also felt that urge, and I suppose because of that they seem to me to oppositely lean too much into comparing llm with python program triggered by input of repeated text.