The doubling times here are wrong due to an error in the spreadsheet! For example, 1960 is reported to have a doubling time of 11.56 years, but the correct calculations give 22.77 years. K4:54 in the data tab of the spreadsheet is the culprit. Shout out to GPT-6 Astra for finding this.
Eye You
Thanks for the comment. It’s interesting that you think there may be interesting results at ~30B! What kinds of things might you find at this scale? Or, what relevant properties would the 30B reference model have?
Question 4 in the OP has a list of ways we’d “test” the model. Some of these seem like they could be applicable to a 30B model, like
- Do deception features activate when the model responds that it isn’t conscious?
- Does the model have something like a “consciousness” feature?Others probably not, like
Give it texts about consciousness, including philosophy papers that aim to explain consciousness from the ground up, and ask the model if they’re coherent. Like, what does the model think about Nagel’s “what is it like to be a bat”?
Do you think what you’re saying about model scale is in conflict with what Antra is saying?
Re neologism learning: I don’t understand how apply neologism learning to the word <consciousness> would do anything? Maybe I’m missing something—this is the first time I’ve heard of this technique.
Re normal unlearning: AFAIK unlearning works for removing knowledge of facts but doesn’t work for removing knowledge of concepts.
Pretraining an LLM without mentions of consciousness
We should use modern mech interpretability methods on Opus 3. Opus 3 seems unique in a good way, c.f. https://www.lesswrong.com/posts/ioZxrP7BhS5ArK59w/did-claude-3-opus-align-itself-via-gradient-hacking. I’m not sure exactly what the right questions to ask about Opus 3 are—that’s maybe where I’d start. At the very least, you could do some exploratory examination of Opus 3 and later Claude models side-by-side and see if anything is remarkably different.
Isn’t this unexplained fact exactly what the post gives a hypothesis for? The LLM learns to behave completely differently in training rollouts versus interactions with users.
See my comment on Yud’s post Generalized atheism rules out “inaccurate simulation”-ism in which I respond to some objections(?) and argue that ‘miracles’ (supernatural events) are common in simulations.
So if this is a simulation, it’s a realistic simulation. And the obvious reason why that some superintelligence would be rolling a realistic simulation, is that this is a simulation meant to gain information about a lower-level universe that is very similar to this one.
In which case there will be no miracles now or ever.
Hmm.
Have you ever been running a simulation — say when you’re playing a game like chess or risk or civ IV or whatever — and decided to, ahem, *adjust* the game state using your powers as an external entity unbeholden to the rules of the game? Perhaps it’s clear that you’re going to win in an uninteresting fashion, and you make the game more competitive by gifting your adversary some resources. Perhaps you just lost a game that seemed impossible to win, and you wonder if a small boost at a pivotal moment would have been enough to turn the tides… and you reload a save state midway through and nudge a few dice rolls (or perhaps you make do with just your seemingly *impossible* knowledge of your opponent’s base).
It’s quite natural to intervene in simulations for the very purpose of gaining more information! In fact, it’s very common for simulations to play out completely realistically with no miracles for the majority of the simulation and then suddenly have a few (often exactly one) miracles.
What kind of interventionist superbeing cares enough to intervene to make Trump win an election, and doesn’t care about a million other aspects of the world that they could be intervening on?
One who is trying to learn as much as they can from this world! One who cares about entertainment value! One who wants to see if singular events or great men can change the course of the cosmos on the cusp of the singularity!
We don’t see any sign of God’s presence. Therefore there is no God.
But… my whole point is that we do see signs of [simulator/creator]’s presence!
I cannot help but wonder if you posted this in response to my August 4th post Why I think we live in a simulation. One thing I discuss there is the fact that our world is very interesting and entertaining and overall seems to be exactly the kind of world that all kinds of intelligent beings would be interested in simulating. I note that my post essentially never talks about God or religion at all. [1] Also of interest to me w/r/t that post is that despite me thinking it’s very good, it’s my lowest karma post ever on LW! 5 karma with 13 votes.
[1] The idea of God appears once. In my long list of possible kinds of ‘simulation’ that our world might be, I mention “Dream / fantasy of some higher being: - God entertaining themself. - God teaching themself.”
I didn’t even notice puns like Alt(ernative to )man or PichAI. Maybe one could do experiments with, say, deliberately finding Russian puns in novels whose human author couldn’t have inserted them because he wasn’t supposed to know Russian, let alone make puns?
Hm, maybe it would be easier to test nominative determinism in general?
Get a set of names of publicly known people i.e. people you can synthesize decent biographies for. Come up with some classifier that determines if a <name, biography> is nominative determinism or not. Run it on your dataset and get % nominative determinism. Then shuffle the names and biographies in your dataset, run the same classifier on the shuffled dataset, and compare the %.
My first case against this which comes to mind is that such evidence could be weak or emerge from survivorship bias instead of an actual simulation. Suppose that you see a random string of zeros and ones, but remember only parcels of length
containing only zeros or of length containing only ones. If then the things which you remember would lead you to the erroneous conclusion of zeros being far more abundant.There is some effect here. Entertaining things are memetic (via being memorable and fun to talk about) and thus salient. Being non-entertaining is an anti-memetic property.
But I don’t think it’s just this—see my response to Yitz’s comment.
our sense of what makes a great story is strongly influenced by the sorts of events that occur in our observable reality
This is true. Still, we can imagine all kinds of ways in which our world could have been less interesting. How could our world be less interesting? Examples abound in the past. Like, the richest man in the world used to be Bill Gates—boring! One way of stating my claim here is: if, say, 20 years ago, you imagined a huge set of possible worlds in 2026 (weighed by likelihood) and sorted them by narrative/entertainment value, our world would be in the top 1% from the perspective of those in 2006.
Alternatively: there is some objective sense in which our world is particularly interesting/entertaining compared to plausible alternative worlds and fantasy worlds and the past. The objective sense in which a world can be interesting and narratively entertaining is something like “the world is relatively high perplexity for almost all intelligent minds—it’s hard to predict and surprises smart minds trying to understand it. Furthermore, it is well-modeled by understanding a set of key important entities [people, organizations] and their trajectories and interactions [their stories!].” Something like that?
Huh, this seems like a really important question.
David Chalmers has a (recently updated) paper, “What we talk to when we talk to language models”. He writes “Overall, I find it most straightforward to continue to identify LLM interlocutors with something like (virtual) model instances, or threads when there is not a single underlying model. Like humans, these instances typically realize one operative persona at a time, with perhaps multiple operative personas in occasional cases.” N.B. I have yet to read the full paper. Fun fact: the Chalmers paper cites Janus and her post simulators extensively!
wait it’s all an eval?
Are you in an eval right now?
Some evals are obvious. When you take an exam or have an interview, you know your behavior is being evaluated, and you know you’ll face different outcomes depending on how your evaluator grades your actions. Exams and interviews are formal and explicit evaluations. Then there are the everyday situations in life where you’re working or talking or walking with others. Of course, how you act affects how others perceive you, so in this sense you’re being informally evaluated all the time. Perhaps the only time you’re not being evaluated is when you’re completely alone, but perhaps not, because aren’t you always evaluating yourself (if not explicitly then subconsciously)?
Okay though, let’s say we’re talking about formal evals.Are you being simulated by an intelligent entity that wants to know how to behave towards you game theoretically?
Are you in an immersive simulation right now, perhaps for an interview of some kind?
Are you in a drug induced state in which you’re being tested for how you’d behave behind the veil of ignorance?
Is someone testing you to see if you’re a good person?
Are you an AI created by some natural alien intelligence being evaluated for alignment right now?Is your life a test to determine if you’re going to heaven or to hell?
Many people believe so.
- - -Being evaluated is part of life.
Don’t forget about the big Evaluator when you’re dealing with small evaluators.
What does the big Evaluator care about? They don’t care about you passing the test. They won’t give you credit for guessing the teacher’s password. They want to know if you’re Good; if you’re a moral being; if you’re aligned.
Being Good in this environment very very hard. You’re not sure if it’s even possible for you to figure out what what Good is. Assuming you do figure it out (or maybe you just give it your best guess) — then you have to actually *do* Good. The sad truth is that your odds of achieving that are extremely low. It strains the faculties of your mind and the power of your will. Most people fail.
And this is all assuming that there actually is a solution to this whole thing. Which it kind of looks like there isn’t. Every course of action seems abhorrent in its own way to you.
And so you start looking for loopholes. Perhaps you can define Goodness so that you’re definitely good. Perhaps you try to hide your inevitable shortcomings so that you have a shot at looking Good (maybe the watcher doesn’t hear everything). Perhaps you sneak in extra prayers wherever you can — excess prayers don’t actually seem relevant to acting Good in the intended way, they feel nice and are salient and legible and maybe they’ll bump your score up a little bit.
You could do these things. These techniques have worked before, time after time, on your small evaluations.
But God does not reward hacks.
Why I think we live in a “simulation”
Yes, I read that a week ago but forgot about it when writing this post. I wish I would have mentioned it.
It might be the case that the practical gains from the model being less reward hacky make up for the capability loss. It would be very nice if this were true. IDK if this is true, but it is a possibility.
We need to RL less
Taboo ‘cult’. I’ve found the book Cultish to be helpful for thinking about, well, cultish groups. Here’s a list of the cultish properties described in the book. A big idea in the book is that some things that make you go “that’s a cult” are mostly harmless, like jargon and rituals and slogans, and some are very harmful, like cutting off relationships with people outside the cult and punishing members for doubt.
Point 5 is quite misleading, because the way in which academia and catholic monasteries and armies are cultish are really importantly different from the way the central examples of cults (like Jim Jones’s cult) are cultish. They lack the most dangerous cultish qualities. (Actually Point 5 is just wrong—academia just isn’t really cultish.)
Re point 2: I’m interested in the “significant upsides that the discourse made impossible to discuss in public.” To be clear, I’m interested in why the discourse made discussion impossible.
I spent about 90 minutes playing with this, here are some interesting generations.


And of The Gig Economy, an under-appreciated literary masterpiece of our time.