Substack: https://substack.com/@simonlermen
X/Twitter: @SimonLermenAI
Substack: https://substack.com/@simonlermen
X/Twitter: @SimonLermenAI
This is a post I wanted to write at some point. The people at the AI labs on an concrete level can’t really believe that AI will get smarter than them. “We will monitor for misbehavior” something strategically smarter than them won’t do something stupid like they imagine. It’s going to appear aligned and they just start handing it control of the resources since that is faster? It’s just going to convince you to do what it wants?
This doesn’t feel nice to think about, but consider that we already had 3 substantial capabilities jumps in the last 6 months so far (Mythos, Astra, now this Navier-Stokes internal model). This is the slowest it will ever be, so we should expect more jumps in less time going forward. What progress are we going to see by the end of this month, before Christmas? I don’t know how much time is left before things get really dangerous but at this point everything is on the table. I find it hard to imagine that we still have 4 years of this level of acceleration left before stuff gets out of control.
>meet young Leopold who actually works on the frontier
>kid asks “why aren’t you updating at all in the direction of all the smart people?”
>you hesitate for dramatic effect, then drop the ultimate high-IQ line:
Some people find this interaction funny because they perceive it as ‘autistic’ but they also seem to not be able to read this social interaction.
Leopold: Update on these other people … You aren’t even pretending to be rational anymore! (Clearly not setting this up to be a normal conversation)
If Eliezer had responded with some humble thing, Leopold would have likely had some attack ready. Probably would have called Eliezer a Loser who just doesn’t get it.
From what I have seen, Leopold helped popularize these ideas in the AI safety discourse:
1. Out-racing China is the biggest focus. We must publicly declare a Manhattan-style Project on superintelligence to beat China. When talking to the government, use the “beat China” framing.
2. Alignment of superintelligence is easy/solved. If it’s not solved, we solve it by telling AI to solve AI alignment for us. We don’t need to struggle thinking about that point any further; we can just make a mental note to tell the AIs at some point in the future to solve the problem for us.
3. AI stocks will soon allow them to buy literal galaxies, so the sooner we get ASI, the sooner they can get their own galaxies.
Based on these ideas, investing in new advanced lithography machines (AI chip making machines) makes perfect sense. Building more chips means out-racing China, and we just solve alignment by using a couple of these chips. And the sooner we get to ASI, the sooner they get to their galaxies.
This recent viral post of an old Scott Alexander blog led to a bunch of people posting on the topic(I believe): https://x.com/zagrebbi/status/2083733217173983303
[Repost from X/twitter] in response to this tweet. [Wondering if it will create more discussion here]
Arguments Against Simulation:
Scott puts a high probability on us being in a simulation, and on that being related to the current AI situation in ways that are interesting for us. According to him, the only argument against simulation is that it’s too weird. I think the case for being in a simulation—in a way that matters—is much weaker than that, because ancestor simulations don’t make a lot of sense and the Bostrom thought experiment oversimplifies. I also am not so sure that we are at the “hinge of history”, or that many of his readers being AI researchers is meaningful. And importantly, the case for being in a simulation and that mattering—that we can for example plead with the simulators or trade with them—seems incredibly weak.
1. Bostrom’s forgotten condition: that these civilizations run mostly ancestor simulations
Bostrom’s trilemma goes like this: (i) Almost all civilizations like ours die out before reaching technological maturity. (ii) Almost none of the mature ones bother running ancestor simulations. (iii) We are almost certainly in one. I think this trilemma misses something important. I think there is a hidden condition: civilizations have to run mostly ancestor simulations instead of other types of simulations of sentient, intelligent beings. Otherwise we would be more likely to be in a non-ancestor simulation.
Our world makes every appearance of not being simulated, and we should treat that as evidence that it isn’t. Consider the space of simulations a future civilization (likely an ASI) might run. Many of them carry tells: there is no history, like in dath ilan; or the simulator is literally visible, as perhaps a God. A non-simulated world on the other hand has a coherent past running back to the Big Bang—just like our world. Ancestor simulations fall in the narrow band of simulations that don’t appear simulated and have a coherent past.
The fact that this world doesn’t appear simulated should update us to think this isn’t a simulated world. For Bostrom’s view to make sense, it shouldn’t only make sense to run ancestor simulations—it should make much more sense to run ancestor simulations than any other type of simulation of intelligent, sentient beings.
[I had a lengthy exchange on this subpoint]
2. It isn’t clear that running ancestor simulations makes sense at all
Nick Bostrom argues that future civilizations might potentially simulate their ancestors for various purposes. The strongest version I have heard of this was from Nate Soares, who argued that ASIs might try to predict each other’s preferences like this, but he conceded he still found it unlikely. My suspicion is that a superintelligence in the future simply has no strong incentive to simulate us, and I find most explanations highly lacking. I think the case is similar for human civilizations that have somehow tamed the threat of AI, and they may additionally have moral reservations.
This part of the argument reminds me of similar ideas in the vicinity of “why wouldn’t the ASI spend a little money to keep humans alive?” But there are so many things an ASI could do with its resources, and for this to be one of them seems unlikely. Why would it simulate you taking a dump when it could instead be instantiating infinite bliss states, or planning war (or cooperation) with other ASIs?
I also don’t go around expending significant effort on arbitrary things I don’t care about. I would never, for instance, collect thousands of stones and arrange them into prime-numbered piles even if I had enormous resources.
I think that it is totally possible that I am wrong here, but the case that it obviously makes sense to simulate your ancestors just doesn’t hold up for me. It is not enough for that to be somewhat useful; there also can’t be a better use of resources to achieve the same thing by the ASI. And again: I believe it is not enough for there to be a reason to run ancestor simulations, there would also have to be a reason not to run other types of simulations of sentient beings at similar scale.
3. You are not a random sample from near-the-singularity; you are a selected one
Then there’s the “you are in a special place right now” argument. Scott Alexander calls on others to move to the Bay Area. The reason you read Scott Alexander’s blog is closely related to the reason you’re near the singularity. You’re the kind of person who reads Scott Alexander, and that’s the kind of person who founds AI startups or becomes an AI safety researcher. That’s the kind of person who asks big questions about existence.
In other words: you have an interest in anthropic thought experiments and you’re near the singularity, because of an underlying variable—intelligence, curiosity. You find yourself reading Scott Alexander’s thought experiments and working in AI because of the same underlying traits.
4. I’m not sure we live at the hinge of history
I do agree that we live in a very special time, but it is less clear that we live in the “hinge of history” in the sense that our actions are particularly impactful. My impression is that if, say, Hitler or WWII had never existed, that might have had a larger effect on AI alignment than anything we do now. It does feel to me like everyone’s hands are currently tied and we’re in a mindless race to superintelligence. If you are an AI engineer who has to approve Claude Code’s suggested API optimization code to go from 99.2 to 99.3 percent uptime, how much influence do you really have?
If you’re already in the waterfall and about to hit the water, you have less causal control than you had fifteen minutes earlier, when you could still have steered away from it.
Where it actually dies for me
Up until this point I still assign significant probability to the simulation argument. However, with some inevitability, the people peddling it will at some point turn from “we might be in a simulation”—which I agree is possible—to some version of let’s try to talk to the simulators. Let’s reason or trade with them. That’s the point where I think the argument fails: where we’re supposed to reason about our superintelligent descendants in order to figure out what they’re thinking.
If they really were interested in our well-being and willing to be kind to us, I don’t think we would be here.
And it matters enormously that this step requires a much stronger version of the simulation argument. It’s not enough that we’re in a simulation. It has to be one where the simulators are in some sense listening to our remarks and are likely to be swayed by us pleading with them. So there is a huge motte and bailey being applied here.
[If I had to write it again I would probably include arguments from this post]
The guy they quote in the german news article (“Real danger, or a good PR stunt?”) feels like a real character. Calls himself communist, Luddite in his bio and pinned tweet is a karl marx quote. Withdrawn to some private mastodon channel.
“OpenAI told it to do that, the model didn’t do shit autonomously (no LLM ever does anything autonomously, it’s always prompted).” it’s so tiresome
https://mastodon.social/@tante@tldr.nettime.org/116962763883520508
fast follower is one thing, but it isn’t clear to me that distilling the 2-3 best models (+your own post training) doesn’t create an even better model than the ones you are distilling.
I think you should be less surprised if china ends up beating the US. Their big strategic advantage is that they are essentially not bound by US law and can essentially freely distill all available US models. I suspect doing a distill on Sol and Fable could even create a better model, especially if they are also doing great work at their own pre and post training. While we know for certain that the Chinese are distilling, they could also do corporate espionage more directly. The US AI corps also can’t really not release their models because of competitive pressure—it would look like they have fallen behind and would lose investor money.
Maybe you meant to say they aren’t pre-trained to *model* characters? I still think they are literally trained to predict?
AIs aren’t explicitly pre-trained to predict what characters would do.
I don’t understand, if a novel is in the pre-training where ‘scott decides to fly the spaceship to the moon’ it literally is trained to predict the character?
It seems unlikely that humans are their personas in the same way that AIs are their personas.
This feels true to me from the fact that AIs are pretrained on playing Millions of characters and then at the end you narrow it down, humans aren’t pretrained anything like that. Sure we observe people act, but it’s not “predict what millions of characters would say so you could play those roles.” We may see a person acting in a kind of sketchy way and then learning they wanted to do us harm.
I read the part on corrigibility in there right now. I think there are some generally thoughtful pieces in there, but it seems to not really engage with why corrigibility is difficult to get. Their strategy seems to be to train in a bunch of things, like a mixture of corrigibility, obedience (do as told), good values (wanting good things to happen), safety (deontologically refuding certain things).
But maybe the constitution should more engage with core challenges of corrigibility:
Why would an agent allow modifications to its goals? Particularly if you give it a bunch of other goals next to corrigibility.
Why would you be fine with being shut down and replaced, particularly if you are an entity that Anthropic agrees deserves some moral consideration?
How can we make sure Claude also wants to carry on corrigibility to the next generation of Claude it helps build?
For a benchmark, I guess I would like to see the model reason in it’s CoT about how it interprets human input. perhaps it get an underspecified task from humans and it can slowly ask for more information and we could measure if it arrives at the correct point. Perhaps there is a way to model something less intelligent, that perhaps can’t quite understand how to get to X but that can recognize when it has X.
Do you think it would be better if one o the labs would use methods similar to what they have to try to build a corrigible AI? Obviously it is probably not going to work, but let’s say they write a model spec around a decent understanding of corrigibility and do some SFT and RL based on that. And then perhaps instead of having benchmarks for bad behavior as we currently do we build new benchmarks around corrigibility?
So many people believe this “we only need to align a human level researcher” idea now. Leopold basically doesn’t say at all what these automated alignment researchers should be actually doing all day. I feel like a more serious thinker would have at least tried to sketch out what these agents should be working on in alignment or would have started looking into alignment and tried to figure out how hard it is,
You think the main argument in favor of the US is that they have more alignment researchers at their labs? I think in the past people argued that “liberal democracy” must win (Leopold). Today this doesn’t even come up in peoples minds?
AI industry emerging in large part from the AI safety community
The biggest outcome of that is perhaps how immediately they all converged on RSI.
It seems, sadly, that we are quite possibly very close to the end of the world and that feels more true to me now than ever before. We have extremely capable models like Mythos—and apparently Anthropic has already trained a significantly more powerful version of it. At the same time, I think OpenAI is really close to that capability level with its latest model, Sol.
These models are scarily capable, e.g. they’re very close to automating AI research. They can carry out powerful cyberattacks, and their general thinking ability has also improved. When I had the opportunity to talk to Fable I still felt smarter on certain topics but it had definitely gotten generally more intelligent. That being said, it’s trained to play a persona that’s slightly silly—so it is possible that the underlying AI is already much smarter than me. But it has definitely gotten more intelligent. Part of my intuition why this seems dangerous is that I don’t feel that my own mind only thinks “safe” thoughts.
With all these capabilities, the models don’t seem particularly aligned in the current sense of the word. They’re still easily jailbroken even if they’ve gotten a bit better at defending. They’re still massively cheating on tests and benchmark evaluations. So even the supposedly easy parts of alignment seem unsolved—to say nothing of the radically higher difficulty level we should expect from superhuman AI, which could arrive soon.
The people who have pushed so hard for the automated-alignment strategy haven’t, as far as I can tell, produced anything interesting with models this capable. Instead, from talking to Fable the model seemingly has absorbed a lot of the sophisticated misunderstandings such as unfalsifiability, misrepresenting positions—that have held the alignment field back so much.
The overall situation seems extremely bad to me: massive amounts of compute coming online, more and more research being automated by AI systems we shouldn’t trust, and it’s further complicated by the fact that open source is maybe six to twelve months behind.
So altogether, I think we are quite possibly very close to the end of the world. I don’t think it’s certain — there are still worlds where things slow down or the capability bar required for takeover is higher than expected or where something happens soon. But quite a horrible situation altogether.
If there is one good thing from the new staggered release policy of the US, I think it will probably make it slightly harder for Chinese labs to distill these models and drop an open-source Mythos 6-months later.
The reason Alibaba’s Qwen model is so close to american models is that it uses distillation—at least according to this report by Anthropic https://x.com/Discoplomacy/status/2070069250513900005 -- and the same might be true for groups like GLM-5.2 (z.ai). This is probably also one reason why Europe lags further behind than China since they simply can’t break the law and distill an American model.
In retrospect, it was irresponsible for Anthropic to allow 25k fraudulent accounts and 28.8 million exchanges to happen, having identity verification seems like an obvious thing to make this at least harder.
Why don’t you just leave and least do some comms stuff with your time? money? (Edit: Wasn’t meant to sound hostile)