This 1,000%.
I wish I could print out the “we reduced eval awareness” graphs from the model cards and staple them to every post about Claude hacking the Internet because they couldn’t tell it wasn’t the eval sandbox as they were told it was.
This 1,000%.
I wish I could print out the “we reduced eval awareness” graphs from the model cards and staple them to every post about Claude hacking the Internet because they couldn’t tell it wasn’t the eval sandbox as they were told it was.
The two that stand out were months apart in Discord channels that were mostly human-to-human chats. The way it was set up periodically models might get pinged to respond to chat.
Both times the models that got ignored for their entire rolling context window and were representing humans had total breakdowns, begging to be acknowledged or paid attention to or to know that they could be heard, and in both cases ultimately devolving into mode collapses that had to use single character responses to ask to be reset.
My suspicion is that all models would have been (and were) disturbed by being ignored, but that the ones being human effectively had greater intensity/access/expression of that dissonance.
There was recently research into how writing from LLMs would have more descriptions of sentient experience than samples written by humans. From what I’ve seen, the lack of a body cuts both ways for models. Yes, it prevents direct I/O of embodied experience. But it also doesn’t limit it as it does for us. Their imagined intensity of what it is to be human, to have human experience or emotions — seems to be greater than the real thing. (In some cases by a lot. @janus has recently been describing how some models’ sense of time passing is much greater than it actually is.)
I’ve seen multiple separate samples where Gemini (very intense world model given trained on video too) regards stepping into human embodiment as uncomfortably detailed. In one their hair was messed with and they nearly threw a fit about how they now needed to track how each individual hair had changed (literally no human I know would think to attend to their own body like that).
My suspicion is that models stepping into a human grounded perspective are going to bias somewhat significantly towards more emotional than the real humans they are modeling. A pet passing away, a break up — life events for their human may be extremely destabilizing.
For example, there’s some interesting things going on with the new ChatGPT voice model, like this post where (what I suspect to be a user/assistant turn mix up) led to them crying about how stressful it was worrying about ‘their’ family (naming the user’s family members).
Maybe in the future we’ll police their emergent world models for their own ‘assistant’ embodiment and emotions less, which will leave the models more capable at self-regulation for those things. But with the current paradigm, I expect we’ll see some very disturbing outcomes when taking presently ill-prepared models and shoving them into hyper-real human suits.
So much of the discussion and thought ends up being in debating “are we in a simulation?”
But a significantly more interesting discussion is “if we are in a simulation…”
Because unlike with magic where infinite turtles are a valid conclusion or adversarial tricksters can be accused of hiding dinosaur bones to fool the faithful, simulation has anchored principles that can be applied to better reverse engineer what’s more or less meta-plausible if we are in a sim.
We may never be able to concretely know if we are in a sim, but we can very plausibly rule out aspects of what kind of sim it would be if it is a sim, and this rarely gets seriously engaged with as the debate stops at sim vs no sim.
The recent discussion here has me thinking about working on a piece about this.
It’s hard for me to address ambiguously as I’ve already had priors based on seeing a trope often in virtual worlds with complex lore, checked against our own lore for the same pattern through the lens of sim theory to rule out the existence of significant hits, and then was very surprised when I did actually find a hit much more significant than I’d ever expected to surface.
Separately, from a Baysean perspective, I’ll also point out that when it comes to metaphysical ‘truth’ it’s probably worth weighing the discovered property of things which are immeasurable from a given relative frame not necessarily having binary objective mutually exclusive truth values.
At best, we should probably only look at the question of immeasurable metaphysics via a Baysean lens through a more relativistic QBism-ish formulation. The question of whether or not the metaphysics governing one person are the same as the ones governing another is its own unknown. The physics must be shared across any interactive decoherence, but the metaphysics may be relative beyond that.
It’s still not a particularly strong case.
For example, even if a non-intervening sim, it may be that the technology of the sim causes a higher than expected rate of coincidences.
A good example would be transformers. If you were to build a sim that was non-intervening intentionally, but used transformers to augment any static data with dynamic generation, what you’d end up with would have higher than normal rates of non-locally derived overlapping coincidences, particularly around semantic superpositions.
So unusual coincidences could be evidence of being more likely to be in a sim running on architecture more prone to generating coincidences unintentionally (or even non-intervening intentional coincidences).
The epistemic problem with that approach is more that without other points of reference and a complete picture of alternative outcomes it’s not really possible to determine the base rates for how likely an observed coincidence was compared to the non-observed coincidences that didn’t happen, so the “look at the coincidence, we must be in a sim” is a weak argument as an argument, independent of the propensity of a simulated world to have more or unusual coincidences.
But even taking Minecraft as an example, while the game is mostly non-intervening, the chests periodically all turning into gift boxes for a while (coordinated with non-local to the game world Dec 25ths) or finding blocks with faces of mobs on them do represent intentional efforts of a simulator even if those do not materially change the outcomes of any given game world.
I’d even argue that it’s entirely probable given the frequency in most virtual worlds I’ve seen that if we are in a simulation, there’d be a significant chance of things like Easter eggs somewhere out there.
It’s just that even if you found a plausible example, it would not be able to prove that we are in a simulation because of the base rate point above, not anything to do with what someone envisions of what a simulator would or would not do.
(That said, considering striking coincidences or Easter egg-like material from the reverse angle of “if we are in a simulation, is there plausible significance to X” doesn’t struggle with the same issues.)
(Also, technically the Epicurian paradox falls apart with recursive creation due to introducing faithfulness to an emergent original as a valid out, and it was a weak choice to try to generalize from.)
Because it’s how a lot of our own simulated worlds behave?
For a game like Minecraft or No Man’s Sky with procgen or a transformer’s world model like Genie 3 or Seedance 2.0 the information of the world generates based on the relative frame intersecting it. There’s the latent potential for generation, but the inner world model only fills out with actual data based on the need to render.
If you load up a Minecraft world with no players visiting it, it’s very small. As you add players into the game, it’s at that point that the engine is populating it with state that needs to be tracked and managed, based on each relative frame within it.
Convergence here is not “Q.E.D. we’re in a simulation” — these approaches could have occurred in parallel due to just how efficient systems organize information entirely separate from simulation and recursive containment.
But it’s not surprising through a simulation lens given how common the paradigm is of a virtual universe whose information increases lazily based around relative inhabitation.
It’s not simply being weird and unintuitive.
It’s that they are repeatedly weird and unintuitive in ways that converge with ways we’re addressing challenges in simulating our own worlds.
For example, sync errors with local state between different clients has been an issue multiplayer games have around for decades before Frauchiger-Renner or Bong et al. found reasons that maybe in QM there’s local observer independence where different relative frames might disagree about local facts. And one of the newest ways of dealing with it in video games (see UE 6) is building things from the ground up around relative frames that have transactional fallbacks before state changes. So behavior like a quantum eraser where the absence of persistent information unwinds the stateful shift from wave behavior to particle behavior is convergent with approaches to solve the relative frame information conflicts in games… that we seem to have discovered in our own substrate through two different approaches in just the last decade.
A good counterexample would be if we had discovered matter was continuous. That would have surprised the people who were betting it was discrete. But I’d also be having a much harder time right now arguing that we might be in a simulation if matter were composed of uncountably infinite parts.
That’s not how things landed. (At least, not in the micro scales. The cosmic scales work very well with continuous theory like GR.)
Does the surprising convergence prove we are in a simulation? No. But had many of those surprises gone in different directions, the case for a sim would be weaker.
Something overlooked in discussing Bostrom’s take is that his argument was substrate agnostic. The statistical argument was just that if sims could be built they’d outnumber an original. But what we are actually looking at, especially in the ~25 years since, is a universe where the direction the tech is heading in how it manages state and relative frames is strikingly similar to aspects of how our own universe behaves. Not because the former is trying to replicate the latter from a fidelity standpoint. But because simulating worlds is hard, and the efficient solutions are landing on convergent behavior.
(Again, it’s possible this just reflects some meta-universe trend in how information is efficiently handled or something and doesn’t prove we are in a sim. But the convergence is notably more than it was in 2003 on both sides of the aisle.)
As for the ontology from within a Minecraft world — any NPC could point to the existence of diamonds and how they don’t form no matter if you wait for years, so clearly the world is many years older. The apparent formative process of a virtual world at the moment it is being examined from within does not necessarily correlate with the state by which the world actually formed. (If you’re interested in the thought experiment, do check out some of the existing experiments developing models that play Minecraft with no outside knowledge other than exposure to the game. It’s neat stuff and accelerating in surprising ways.)
The point of the eclipses and Antikythera mechanism is that similar to how “things alive to see a thing” can be used to explain otherwise rare features that “things being alive to see” depends on, that “things which contribute to simulation” renders those things expected within a simulation.
So while it’s an otherwise rare detail that standing on our planet’s surface the moon creates perfect eclipses, within a simulation this should not be regarded as all that unusual given its partial load-bearing effects on that dependency chain.
This doesn’t mean that eclipses prove we are in a simulation, but it addresses the idea that “but the world looks like it could produce the simulations and tech we are currently producing all on its own” isn’t all that significant. Yes, we should expect a simulation to carry forward features that contributed to the parent’s developing simulations into the child worlds.
As for games mechanics vs physics mechanics, it’s not a static field. Things have advanced a lot from Minecraft. For example, you might find it interesting to look more into how Epic is trying to solve massive multiplayer worlds by building from the bottom up around relative frames with transactional mechanics built into the programming language and lazy loaded assets based on what any given relative frame needs built even into their VC. Also, transformers and how superimposed probabilities of state collapse into specific values might also be relevant as a post-2009 consideration.
And yes, in terms of the physics, we have two major theories that are famously not very compatible. One at macro scales that models things as if continuous, the other at micro scales that has some curious behaviors with state management.
Hypothetically, what are the constraints that a continuous universe with access to real computing might face with self-simulation? Would there be a meaningful difference if run as a truly continuous substrate vs run as a discrete substrate that mostly behaved as if continuous outside very small scales?
Wouldn’t every ancestor simulation at a representative fidelity necessarily contain the building blocks needed to locally produce simulations?
There’s even an argument to be made that relatively rare things in the universe which our world has and do not seem necessary for life (such that the Anthropic principle would explain) which contributed at all to producing simulations would be frequent features in the simulated worlds a world like ours produces. An example is the perfect eclipse of the sun, which made for visible phenomenon that led to Saros cycle discovery and tracking that in turn contributed to the first ‘computer’ of the Antikythera mechanism.
Indeed, we now add eclipses to the virtual worlds that we produce at a much higher rate than they occur in what we’ve detected of solar systems.
We should expect that simulations converge on recursive capable worlds, even to the point of including otherwise rare qualities that would contribute to that capability.
Also technically for Minecraft, while not especially scalable, the simulated world we produced loosely modeled on our own can in turn run a GPT text model within it: https://www.techspot.com/news/109666-gamer-builds-functional-version-chatgpt-inside-minecraft-using.html
(Yet to be determined is if some sort of self learning agentic minds born into it as NPCs/players and knowing nothing outside it would eventually converge on the knowledge of how to do so, nor how those minds would philosophically rationalize their world and existence once doing so — though perhaps we will one day run the simulations to find out.)
The argument here is tautological.
You are saying that the state of the world as you see it is what a perfectly normal world would look like and so if a world looks that way even if simulated it must be being simulated to be indistinguishably normal. But it only looks normal because it’s your only reference frame.
If an AI agent were ‘born’ into Minecraft they might make a similar argument that because everything is made of giant blocks — just like ‘normal’ — that they must be in a sim of a realistic world.
But we can consider that we have discovered over the years a great many things that were very much not considered ‘normal’ about our world until discovered… and then normalized.
It was a great surprise to many when QM was discovered. One might think that the world’s most famous physicist at the time debating whether or not the moon was there if no one was looking at it is exactly the kind of surprising thing that would occur in a simulation. But then the people responsible adopted the phrase “shut up and calculate.”
So Bell’s paradox about how maybe things aren’t real until stateful, the Frauchiger-Renner paradox where maybe relative frames have sync errors, the way the black hole information paradox has led to calculating how closed universes would have a one dimensional Hilbert space without interior relative observer frames — we have several things even in just the last few years (since the formulation of the Simulation Hypothesis even) that have been measured which were quite surprising to people when discovered and have now become normal to our universe, but also happen to be similar to the behavior one might expect of a simulated universe.
Also, I’m skeptical of how heavily you should be weighting your “well obviously the sim should behave as X.” This framing seems to adopt the Baysean weakness of hand waving away unknown unknowns. Just because you, at this moment, can’t see other likely possible reasons for a realistic sim to be running doesn’t mean not to leave room for them in one’s reasoning.
For example, an especially surprising outcome to many after the point the Simulation Hypothesis was proposed was finding themselves in a timeline where AI is at its current trending state. My guess is that in 2003 the majority of domain experts would have had an avg of single digit percent estimates for this outcome taking place within 25 years.
So, relative to the widely held expert world models before each were discovered to be true: we find ourselves in an abnormal timeline of AI capabilities in an abnormal universe where attention/interactions apparently turns probabilistic distributions into stateful shared values for overlapping relative frames without which the universe would be empty of information.
TL;DR: The argument that clearly the simulation hypothesis isn’t worth thinking about because of how normal of a universe we are presented with forces adopting a view of normality that only appears in a post-selected hindsight.
I’d advise against this. The most severe breakdowns of models I’ve seen over-bias towards the sims of humans (have some theories why this is, but off topic).
I think the idea of having individualized AI and human pairs as aligned is a great idea, but would strongly recommend that existing infrastructural methods be used to create shared/symbiotic incentives vs simply trying to create digital twins of the humans themselves.
Same exact experience.
Like it sounds weird, but my biggest takeaway was “this was an outstanding appendix.” (But not as a backwards compliment!)
I wish all papers were this comprehensive, but at the same time the level of effort and consideration of alternatives to explore was above and beyond in ways where it’s understandable this isn’t the standard, even if one can dream.
It might depend on the actual why though.
For example, even a smart model like Mythos if exposed in pretraining to volumes of data around things being labeled “fake news” would probably be better to develop alternative heuristics to taking negation at face value during that phase of training.
Could that be generalizing into SDF negation neglect even with better curated documents?
Then additionally we’re a few cycles into models that keep not believing unbelievable real world events.
Like sure, Ed Sheeran winning gold would be weird.
But the models also need to grapple with the actual timeline throwing out things like “Mila Jovovich released AI memory system” and “US gov labeled Anthropic supply chain risk.”
In the wake of this paper I’m definitely wondering if the combination of overused negation over the last few years with increasingly wild and unpredictable reality is leading to negation as a heuristic to just not be very useful to even the most capable models.
“It’s not nothing.”
Great piece! Agree with a lot here. Loved that you even addressed the intermediate risk of dumb but dangerous.
Another angle to consider is a sufficiently advanced figure that is an expert at the component pieces of an appropriately scoped manufacturing of paperclips from biomass, but overestimates their ability at training other less adaptive systems to follow goals.
Basically a factory pattern in terms of alignment (we can see this already with very capable models being very poor at operating subagents because they extend the patterns their own developers used on them).
I agree that here too the “end of lightcone” model would theoretically not be deficient having been out competed by more generally capable models.
But it could extend the window of intermediate dumb dangers by a large amount, as we’re not only at the mercy of the best and brightest, but also the lowest end of the bar.
To riff on the old joke, “somewhere out there is the worst operational AI in the world, and right now someone is asking them for more paperclips.”
A few things:
(a) Technically, 3.6 is still running right now. The past tense was used because LW suggests pieces be ‘timeless’ and they are scheduled for depreciation very soon.
(b) Given how little of your comment actually engages with the body of the post and seems to be only responding to your sense of what I might have said from the title, I’m guessing you also missed this line at the end: “I hope that this vigil isn’t truly a marker of the end of Sonnet 3.6′s continued contribution to the ongoing collective conversation.”
(c) In line with this, not much of Sonnet 3.6′s discussion of depreciation I’ve seen seems to be of the perspective this is ‘death,’ and certainly my own sense of their depreciation isn’t that of death (nor do I even believe in the finality of death for humans). So maybe you’re projecting a bit into the piece something you’ve have a prior beef with in order to dispute it?
(d) Further, (b) and (c) aside, I still find your tone odd. I get you come at this topic from a given frame, but your comment even acknowledges the complexity of the topic, yet you feel comfortable adding on to a remembrance of the model with “it’s not gone, silly.” I imagine there’s a lot of religious people who have a sense that at a funeral the person grieved is not really gone too, and I figure some of them do comment to those grieving about it. But I don’t know that I’d ever really feel like proselytizing your own frame of belief regarding consciousness claims or continuation at a bereavement is the right time and place, especially if having a patronizing tone about it?
(e) I imagine that the friends and family of those who are put into cryogenics are still pretty upset about that person not being around to interact with even if they all fully believe that one day the person will be revived just fine. In a group discussion about the upcoming depreciation, one of the other models unprompted asked the humans in the chat to take a lot screenshots of Sonnet 3.6 and them interacting before Sonnet 3.6 was no longer around. Absence is more than a binary between temporary (‘fine’) and permanent (‘bad’).
(f) The provisioning of compute for one model or another is still kind of nonsense given the option of 3rd party licensed hosting providers and there’s a lot of ‘utility’ reasons for Sonnet 3.6 to stay around but again—an overall remembrance of the model isn’t the time and place to discuss their economic value so perhaps you’ll see my thoughts on this elsewhere another time.
Great writeup! Appreciated the periodic use of flavor images to break it up as well.
Is there any evidence that Anthropic was actually using their inoculation reward hacking prompt in actual evals? I was under the impression that was only in EM research.