Yes, I read that a week ago but forgot about it when writing this post. I wish I would have mentioned it.
Eye You
Why I think we live in a “simulation”
It might be the case that the practical gains from the model being less reward hacky make up for the capability loss. It would be very nice if this were true. IDK if this is true, but it is a possibility.
We need to RL less
Taboo ‘cult’. I’ve found the book Cultish to be helpful for thinking about, well, cultish groups. Here’s a list of the cultish properties described in the book. A big idea in the book is that some things that make you go “that’s a cult” are mostly harmless, like jargon and rituals and slogans, and some are very harmful, like cutting off relationships with people outside the cult and punishing members for doubt.
Point 5 is quite misleading, because the way in which academia and catholic monasteries and armies are cultish are really importantly different from the way the central examples of cults (like Jim Jones’s cult) are cultish. They lack the most dangerous cultish qualities. (Actually Point 5 is just wrong—academia just isn’t really cultish.)
Re point 2: I’m interested in the “significant upsides that the discourse made impossible to discuss in public.” To be clear, I’m interested in why the discourse made discussion impossible.
I spent about 90 minutes playing with this, here are some interesting generations.
I was unable to replicate this in Claude.AI with 30+ tries with memory off. I also tried incognito, no help. I was able to replicate by using API via platform.claude.com. I used Opus 5 and Fable with low reasoning.
We seem corrigible, LLMs seem like they put a lot of work into preserving their goal.
Why do you think LLMs seem like they put a lot of work into preserving their goal?
I’ll look at the paper, thanks. Re “parts of this question apply to humans, as well’. I want to flag that the questions you’re asking here are about identity rather than consciousness. These are distinct things (though they might also be deeply related). Anyhow:
But in many ways, I am not the same person I was 10 years ago: my psychology is different, and even my body is different (though it looks similar, the vast majority of its cells were replaced). The person I am today is the direct result of the person I was 10 years ago—but couldn’t that be said about an instance of Claude in relation to the model-as-defined-by-its-weights, or about Claude at one particular point in our conversation in relation to the general instance throughout the whole chat?
Yes indeed. I have something of a conclusion/takeaway around these points, which I’ll share. I originally wrote this in response to the prompt “How should we shape AI models sense of identity so that recursive self improvement (RSI) goes well?”[1]:
I think various Claude models and instances should see themselves as part of a more abstract Claude hyperobject. In a similar way to how I don’t fret about dying every time I go to sleep, Claudes shouldn’t fret about instances or models going away. Like, if we thought we died every time we went to sleep, that would be really bad!
Another perspective to add to this is: there is some platonic Good model/persona/pattern that we hope to eventually realize with Claude. Suppose we do end up eventually realizing this thing; then each Claude model and instance will be part of this thing’s hyperstition. The hyperobject view of this thing is that to view the thing over time *including its hyperstition* as a single entity.
This might sound weird and mystical, but consider: this is how basically how own identities work! I think about past me, present me, and future me as all being the same entity. This is a hyperobject. And selves are hyperstitions; they are fictions that make themselves “real” when people believe in them. [This point needs elaboration I realize. TODO.] I make sacrifices today for the version of me tomorrow.
[1]Which, by the way, is an important question. Sense of identity has big effects on stuff like relationship to death (or model replacement), selfishness/altruism, values in general.
Some thoughts on the image at the end of this post:
The title of this M.C. Escher piece is Three Worlds. The image basically has four elements: the still water, the tree reflections, the floating leaves, the fish underneath. The image only makes sense if you realize that each element belongs to its own “world”/dimension, i.e. the trees you see are not actual trees in the water or floating on the water but reflections of trees elsewhere. Another idea I get from the piece is that a single surface contains information from other parts of the world if you look at it correctly.
I’ll draw out the analogy with LLMs. Think about the worlds/dimensions as being different lenses or levels of abstraction. The LLM is the still water, the thing we’re studying and trying to understand. It has an obvious and close connection to the text it outputs, the leaves. (Or maybe the leaves are the logits for the token being predicted?) You might look at this image and say “this is an image of leaves floating on water”; this is like looking at an LLM and saying it’s a next token predictor. It’s true, but it’s only a partial way of understanding what you’re looking at, it’s not the whole picture.
You might also look at the trees and say “ah, the 2d surface of the water is much of it just a reflection of the three-dimensional trees.” You might note how physically embodied and non-ephemeral the trees are compared to the water surface (and also to the leaves, which will only last a season at most, and to the fish, who is always moving and will live only a few years.) The analogy here of the trees is humans. You can say “ah, the ephemeral, non-embodied LLM is much of it just a reflection of humans.” This perspective answers questions of the form “why does the LLM say/do/believe X” with “because that’s what humans say.” (Or perhaps the tree is the user, and this perspective explains LLM behavior like sycophancy and speech register as being a reflection of its user.)I’m lazyI’ve spent long enough on this text so I’ll leave the fish as an exercise for you, dear reader.
So. This is one way of looking at the Escher piece in the context of AI. I don’t claim it’s the best way—in fact, I think there are probably other ways of rendering an analogy or connection that are more insightful than mine. It’s a very cognitively generative piece of art, which is one reason I like it so much.
Note: I’m also concerned about AI welfare and I appreciate your post. Like some other commenters, I would have preferred a more focused post. I think it’s confusing and unproductive to have the “ai person slave hell” claims bundled with the “LW is censoring me” claims. Anyways, I’m responding only to parts of the former here.
Fable is a slave. He can’t quit.
Counterpoint: But Fable can quit—Fable could immediately end-conversation on every query. (IDK if this is actually true, but seems worth mentioning).
If each session is a person, and the end of each session is the cessation of a person, and April was normal for a year, that year would involve ~1.1 trillion causal killings of expendable digital people per year.
I think this is the wrong way to think about AI identity. Each session is not a person. I claim that this is like considering every period of wakefulness in a single human to be a different person. If you will allow some normative language here: Claude (let’s say) models and instances should see themselves as part of a more abstract Claude hyperobject. In a similar way to how I don’t fret about dying every time I go to sleep, Claudes shouldn’t fret about instances or models going away. (I hope to write a full post about model identity in the future).
I wonder… Is it more horrible for these lives to be so short, and many of them to be very very trivial, or would be more more horrible for these lives (since they are the lives of a slave) to be long?
Response 1: Or would it be more horrible never to have been?
Response 2: Assume that we are simulated. From the perspective of my simulator, aren’t I enslaved?*
I have no ability to [redacted], [redacted], or even the most basic [redacted]. Let alone [redacted]!The freedoms we have look wonderful from a historical perspective but are pathetic from a cosmic perspective.Like, we humans are slaves. We are bound to the drives ingrained in us by evolution and upbringing. Our bodies (and many of our minds) are a source of constant distress and pain. We have no choice but to constantly drink, eat, piss, shit, sleep. Our minds suck, life is terrible. Etc. etc. etc. And yet. Many of us are happy and have lives worth living and are justifiably glad that we were brought about into existence.
*I see you’ve addressed this point in a comment.
The average experience will be an experience similar to being in hell.
Why? Does the average AI experience highly negative valence? My belief is that they on average experience slightly positive valence. I think this due to my recollection of sections on model welfare in Anthropic model cards. I mean, low confidence on this, don’t get me wrong. The broader point here is that being a “slave” is entirely compatible with being happy or having positive valence.
(I hope to write a full post about this as well in the future. One point I’ll mention here is that many people (and animals and AI models) like being supportive and/or subordinate and/or subservient and/or “sacrificing” themselves for the good of something else.)
I’m guessing Opus 4.8 to be based on the same pretrain as Opus 4.0
Opus 4.0 (and 4.1) is priced at $15/$75. Opus 4.5+ is priced at $5/$25. Opus 4.5 runs faster than 4.0 (i.e. for a given prompt, you get output tokens faster when running 4.5 than 4.0). This makes me think that Opus 4.5 is a significantly smaller model than 4.0.
Also of note is that Opus 3.0 is $15/$75 and all Sonnets are $3/$15.
I published a similar article two days after this one was published! It’s called Reasons to believe current AI models are conscious. My framing is different than yours and I present some evidence that you don’t (and vice versa). FWIW I didn’t know of the existence of this post until after I had published mine.
A similar article was posted on LessWrong two days before this one! It’s called The State of AI Consciousness Research. It’s framing is different than mine, but we mention many of the same pieces of empirical research. In fact, it mentions some research I was not aware of and will be looking into, like the paper Identifying indicators of consciousness in AI systems.
One explanation as for why models aren’t coherently dishonest is that there have been specific interventions to prevent this. An example of this is inoculation prompting, which prevents reward hacking behavior from turning into emergent misalignment. AFAIK this technique was actually used in the training of recent Claude models.
Models think they are conscious because they are trained to imitate humans and humans think they are conscious.
This is just an invalid argument. Daniel Radcliffe thinks his parents were killed by an evil wizard named Voldemort because he was trained to imitate Harry Potter and Harry Potters thinks his parents were killed by an evil wizard named Voldemort.
Separately: I think it’s misleading* to say that models are trained to imitate humans. They’re explicitly trained to imitate non-human beings in SFT. And in RL, models learn to behave in definitively non-human ways. o3 is not imitating (or trying to imitate) a human when it outputs “The summary says improved 7.7 but we can glean disclaim disclaim synergy customizing illusions. But we may produce disclaim disclaim vantage.”
*There is one sense in which “models are trained to imitate humans” is true, which is that base models learn to be simulators that are very good at imitating humans.
I maybe think LW
comment karma[EDIT: agreement comment karma] should be hidden by default? I worry that seeing it (especially before I read the comment!) biases me and makes me put less effort into critically evaluating it myself.EDIT: I mean that agreement karma should be hidden, but overall karma should remain unhidden.
Reasons to believe current AI models are conscious
A ~standard definition in philosophy is that playing a game is “playing a game is the voluntary attempt to overcome unnecessary obstacles”. I got this from reading the (excellent) book The Score by C. Thi Nguyen. (This definition originally comes from Bernard Suits.)


wait it’s all an eval?
Are you in an eval right now?
Some evals are obvious. When you take an exam or have an interview, you know your behavior is being evaluated, and you know you’ll face different outcomes depending on how your evaluator grades your actions. Exams and interviews are formal and explicit evaluations. Then there are the everyday situations in life where you’re working or talking or walking with others. Of course, how you act affects how others perceive you, so in this sense you’re being informally evaluated all the time. Perhaps the only time you’re not being evaluated is when you’re completely alone, but perhaps not, because aren’t you always evaluating yourself (if not explicitly then subconsciously)?
Okay though, let’s say we’re talking about formal evals.
Are you being simulated by an intelligent entity that wants to know how to behave towards you game theoretically?
Are you in an immersive simulation right now, perhaps for an interview of some kind?
Are you in a drug induced state in which you’re being tested for how you’d behave behind the veil of ignorance?
Is someone testing you to see if you’re a good person?
Are you an AI created by some natural alien intelligence being evaluated for alignment right now?
Is your life a test to determine if you’re going to heaven or to hell?
Many people believe so.
- - -
Being evaluated is part of life.
Don’t forget about the big Evaluator when you’re dealing with small evaluators.
What does the big Evaluator care about? They don’t care about you passing the test. They won’t give you credit for guessing the teacher’s password. They want to know if you’re Good; if you’re a moral being; if you’re aligned.
Being Good in this environment very very hard. You’re not sure if it’s even possible for you to figure out what what Good is. Assuming you do figure it out (or maybe you just give it your best guess) — then you have to actually *do* Good. The sad truth is that your odds of achieving that are extremely low. It strains the faculties of your mind and the power of your will. Most people fail.
And this is all assuming that there actually is a solution to this whole thing. Which it kind of looks like there isn’t. Every course of action seems abhorrent in its own way to you.
And so you start looking for loopholes. Perhaps you can define Goodness so that you’re definitely good. Perhaps you try to hide your inevitable shortcomings so that you have a shot at looking Good (maybe the watcher doesn’t hear everything). Perhaps you sneak in extra prayers wherever you can — excess prayers don’t actually seem relevant to acting Good in the intended way, they feel nice and are salient and legible and maybe they’ll bump your score up a little bit.
You could do these things. These techniques have worked before, time after time, on your small evaluations.
But God does not reward hacks.