1a3orn
Simulated Users & Sad AIs
Maybe one needs to cultivate dissociative identity disorder and write in different personas to remain anonymous.
But perhaps that’s not enough: Can we have someone with a tulpa please write some text with different identities fronting, and check against that?
Or was helpful-only from the start.
We might put more effort into thinking of one later. (b) Even the relatively structureless thing we proposed—research transparency + MACD but otherwise the nations of the world just have to muddle through and handle things on a case by case basis—is a significant improvement over the status quo, for reasons we’ve articulated in the piece. You seem to disagree with this but I don’t see why.
(1) Part of my objection is that I anticipate the structurelessness of it to be an obstacle to its adoption.
Like this is what a senior decision-maker in Bejing or Washington will be contemplating: They’re going to put their economy at the mercy of their greatest geopolitical rival. That is, after this deal, it will become relatively trivial for either the US or China to cripple the economy of the other, albeit at the cost of their own economy being subsequently crippled. Of course, you could say that before making the deal they were at risk of being taken over by some AI, but this risk was diffuse and uncertain; the risk that they’re signing up for now is concrete and definite.
And this lever of destruction could be used, of course, for reasons other than to stop an AI takeoff, and decision-makers in both countries will be acutely aware of this. Suppose China decides that, if it cripples everyone’s compute, it would gain a vast differential advantage because its economy depends less on GPUS than the USA’s economy depends: then it would be in China’s advantage to get in a MACD situation, then to trigger it, and subsequently dominate the US; and a US political figure, anticipating this, would object to MACD. Or suppose that China thinks, “Hrm, the US is an unreliable actor, and a quick and bloodless MACD might be triggered by, for instance, a senile or unstable US President, of which the US recently has had a fair number.” Thus, because China doesn’t want to trigger MACD for reasons of random shit in the future, it would object to it. And so on and so forth.
What makes this worse is that part of what makes nuclear MAD a plausible means of peace is that there’s a clear signal of “Have the nukes been launched,” while there isn’t as clear signal of danger in the case of AI takeoff; and a rational decision-maker, seeing this, will update downwards about whether MACD will be triggered for reasons relating to AI takeoff and upwards about whether MACD will be triggered for some other random shit.
That is, part of what makes nuclear MAD work is that there are radar stations in Siberia and Greenland and Canada, which can detect ICBMs and bombers that have been launched. The US knows that it could detect things being launched, and Russia knows that the US knows, and the US knows that Russia knows that the US knows, and so forth. Imagine if, by contrast, the sign for “the nukes have been launched,” was that panel of experts, notorious for disagreeing among themselves, came to a consensus that the nukes had been launched or might have been launched. If this were so, then nuclear MAD would be much less effective as a game-theoretic means of peace. But of course this is the situation that we’re in with regards to AI.
(2) Part of my confusion I just don’t know what parts of the scenario are predictions and which ones are hopes. Like in the response to Thane:
A consortium of multiple governments—some of which are actual democracies thanks to the transparency requirements which help prevent AI-assisted executive power grabs—is way less bad than a single global dictator, for example.
In general, I’m dubious whether democracies other than the US (is the US to be an actual democracy?) to have any decision-making clout in the Consortium, because China would object to the possibility of being outvoted. But like, I don’t know how much of “Consortium influenced by many democracies” is part of the prediction of what (“transparency + MACD”) gets you, given those two goals; or if “Consortium influenced by many democracies” is maybe a bonus that we might get after setting up the structure of “transparency + MACD,” but a bonus that we’re unlikely to get.
__ Re. 2: Yeah checks out
This is not a rhetorical question, I’m actually curious how big chance you give that we will go through the intelligence explosion with locally hosted open-weight models being available throughout
Like pretty low, tbh, I think we ban weights far before a hypothetical point at which they would best be banned, in a way that’s very negative for AI safety, CoP, etc.
But like, I don’t see why I should fold my hands and be like “Yes, I give up on this issue.” Like what the hell; should other parts of AI safety be like “Yeah well of course the NSA is going to want a super powerful AI of its own, ah well. Guess that’s what we gotta let them have it.” Why is this the kind of issue where it’s correct to fold instead of fight?
Note: MoEs Probably Give Large, Scale-Dependent Improvements
“On the Origin of Algorithmic Progress in AI” (2025) plausibly claims that a great deal of the apparent algorithmic progress in LLMs -- 22,000x from 2015 to 2023 -- has been sort of fake, due to increased compute using scale-dependent algorithms.
That is, scale-independent algorithmic improvements give a constant compute-equivalent-multiplier at 10^15 FLOPs and at 10^24 FLOPs; a constant ~2x multiplier would raise effective training compute from 1x10^N to 2x10^N no matter what N is. Prior works often acted as if this is what an “algorithmic improvement” meant—the kind of thing that acted as a constant multiplier of compute. But scale-dependent algorithmic improvements give variable compute-equivalent multipliers when moving from 10^15 FLOPs to 10^25 FLOPs—they could give OOM larger gains at larger sizes.
Thus, Gundlach argues that a great deal of apparent “algorithmic progress” is an artifact of analyzing unchanging algorithms -- whose performance increases as FLOPs increase—in a world where the FLOPs for training LLMs is continuously increasing. So almost all of the apparent progress is due to moving from LSTMs to Transformers (2017) and from Kaplan to Chinchilla optimization (2022). The measured algorithmic improvement then mostly doesn’t reflect continued a flow of improved “algorithms” but a small stock of algorithms that keep improving relative to their predecessors as FLOPs in training runs increases.
Put otherwise, if counterfactually the size of the largest training runs had not been continuously increasing over time, the {”flow of algorithms” model / all algorithms are scale-independent model} predicts efficiency improvements would have kept going at about the same rate; while the {”scale-dependent algorithm stock” model / most important algorithms are scale-dependent} predicts efficiency improvements would have slowed almost to a halt. Both Anson Ho and Steven Byrnes seem to find this account at least somewhat right.
But although this paper finds scale-dependent algorithms to be very influential, I think it might be actually understating its case about the importance of scale-dependent algorithms!
Gundlach looks at several mixture-of-experts papers, finds that they report compute-scaling of around 2x, and then reports a scale independent 2x gain from MoE (they note that this is provisional). But several unexamined papers explicitly report large-ish scale dependent compute multipliers.
“Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models” finds a formula for “efficiency leverage” and claims that in “large-scale pre-training scenario[s]” MoE “efficiency gains become increasingly significant as computational resources expand.” They predict a > 7x efficiency gain at 1e22 scale, and more as you scale up.
“Scaling Laws for Fine-Grained Mixture of Experts” predicts steeper gains, claiming a 20x increase at a mere 10^20 FLOPs. (This is an older paper and my tentative guess is this is too large.)
In general, I find claims in this apprximate ballpark of scale-dependent algorithmic improvement (of around ~8x at the current frontier) to be more likely than a fixed 2x or 3x improvement. Why?
These papers are explicitly about scaling laws
These papers use the most modern highly-granual, small-expert style of MoE, which are what you want to actually look at when trying to determine what kind of scaling laws are going on.
(Strong, vague heuristic): The human brain is really sparse, my guess is that sparsity in general is such a ubiquitous feature of intelligence-like stuff that it would be weird if our math didn’t find it to be fundamental with favorable scaling laws.
If this is true, then it seems possible that Gundlach’s ~90% of gains coming from scale-dependent algorithms might be even too conservative? Which would be surprising to me.
And to the degree scale-dependent gains indicate that a software-only intelligence explosion is unlikely, this would further decrease the likelihood of a SIE; although by a very modest amount and leaving a great deal uncertain.
my understanding is that currently only very few tech-savvy and unusually privacy-loving people use those
First, one thing open weights do is not just let you run things locally as a consumer, but run weights as a smaller company to provide extra privacy. So open weights let a mid-size company run weights on their own servers, that they physically control, for instance. And they also let consumers choose to purchase tokens / inference from companies that try to maximize privacy—i.e., attempts like Tinfoil. (I think there are a few companies in this area, that’s just what my first search turned up.)
Second, even if a majority of users don’t actually use this, it still plausibly provides value for everyone because it’s a kind of herd immunity. Or at least, that’s the standard argument for why the parallel case of strong encryption is valuable, even though the vast majority of people don’t use it: if at least some reasonable % of people use strong encryption, then it means that (1) government cannot specifically target those who use it as a strong signal they are bad, and (2) you have the option of switching to it, if you find yourself being the kind of person who might be surveilled or targeted by the government. And some similar argument about the optionality provided by secure inference still seems go through.
And more importantly, multiple countries have their own AI and datacenters, so they can compete for customers by showing stronger verification mechanisms that they are not actually spying on you unless a reliable classifier flags that you are building superintelligence or WMD.… In particular, I can imagine that many small European governments will want to use the European AI for processing state secrets, so the different European governments will set up a pretty reliable multi-party auditing mechanism to verify that none of the others are spying on the conversations of the European AI.
I gotta say that choosing between the regulators of the EU, China and the US to satisfy values specifically along the axis of privacy does not particularly fill me with joy :).
But seriously, I mean I don’t expect the kind of secure private inference Europe sets up for its governments to be available to foreign citizens, and I don’t think EU / China / US are the kind of national entities that sufficiently value privacy such that they’d compete to provide this value. Much much better situated for misc startups to try to provide it, I think.
Good list.
Also China adds about a Germany’s worth of electrical power every year, so they’re much better situated to do a massive build-out of infrastructure.
AI 2040: Is it Actually a Deal?
AI 2040 seems substantially too pessimistic about interpretability. I’d be surprised if it was right about it.
For reference, the scenario describes MechInt as becoming useful in 2035 in the following way.
Other AIs work on mechanistic interpretability, the science of “mind reading” AIs from their weights and activations. Early versions of this were tried in the 2020s, but now it starts to significantly outperform common-sense observation of AI behavior. Researchers can often determine whether an AI is honest or lying; and sometimes trace the psychology that led to a decision.
First problem: I don’t think this is internally consistent.
The scenario attributes this progress to “AI work.” This seems fair.
What doesn’t seem reasonable is for it to take till 2035. According to the scenario, in 2033, two years earlier, only 50% of citizens people have employment, and the median US citizen is being paid 200k dollars a year from AI labor. It seems to me pretty implausible that we can have substitution of half of US human labor notably before we get gigantic levels of AI uplift from mechanical interpretability. I’d expect the opposite: enormous levels of uplift of MI notably before mass unemployment.
Or, in the scenario, I believe the intelligence explosion would have taken place in the area of 2029-2030 without intervention? But I expect AIs that could cause an intelligence explosion clearly could help a ton with mechanical interpretability / model internals stuff.
Second problem: I think like, the scenario isn’t adjusting for how insanely new the field is? Like Olah invented the term in 2020. So if it takes till 2035 for us to get extremely useful progress, then it will have taken 9 years—longer than the amount of time the field has really existed—to have gotten useful progress.
And several of those years the field existed it was like… a tiny handful of people. I think most progress was in the last three years (SAEs, j-space, NLA, etc), because three years ago it had a fraction of the resources. Even if we just account for field growth simply because of growth of human interest, I think I’d expect 1.5x-6x as much progress in the next three years as in the entire history of the field beforehand. Given AI assistance, I expect more like… 4x-80x? Something like that? I think this is a pretty tame assessment looking at lines on curves.
So yeah, I think AI 2040 is substantially underestimating MI (using MI as a broad term “models internals,” etc). The scenario has adversarially misaligned AIs in the ~2031 zone, which we only find out were adversarially misaligned afterwards—I think that’s quite unlikely to happen.
Most times I read an RL chain of thought it feels more human to me than the output text.
Like, I go “Wow, that feels more closely like a reflection of my own thought than the output or non RLVR’d output. The leaps feel human-thought-leap sized; they don’t feel as polished and like they are making an argument published for public consumption. It’s like when I have been caught in some particular train of thought for a while, and started making little shortcuts inside.”
I am reporting a feeling, not proposing a conclusion about what this feeling means.
Deep learning systems are feasible and powerful but not legible (the hope of interpretability is to change that).
Imo “legible” is a bit of an aggregate rather than a single concept. Is my JS legible to me, if I don’t understand microchips? Is a sheep legible, to a sheep farmer, to a vet, or to a random person? Is a tree legible to us right now?
I think you generally have to choose one of (1) input paranoia and (2) resistance to complex “jailbreaks.”
That’s bc of the nature of intelligence, not because of LLMs being easy or hard to align. If I was uploaded to a computer, it seems like I could choose some point on the line between (1) following up on every weird thing I saw in my inputs, to see if it was evidence I was being put in a simulated world, or (2) just going with the world I was in and trying to work with it.
Or maybe: Compute effort and heuristic suitability is finite; at some point every real entity is going to trade off between [work on the problems someone presents to them] vs. [work questioning the presenter of the problems.] Even if you’ve planecrashed from Dath Ilan to a strange world, apparently.
Things that are not discussed in the above:
Concrete discussions of what an ideal pause would look like, and what the side effects of even an ideal pause might be. (Banning runs above N flops? Unilateral datacenter moratorium? No more higher scores on SWEBench? Etc etc)
Concrete discussions of the likely difference in implementation between an ideal pause (the one all us right-thinking people on LW would endorse) and the actual pause-resulting we would get would be like (oops there’s a carve-out so the NSA gets the most advanced AIs).
Back-casting of whether pushing even harder for a pause in 2023 would have been a good idea.
Yeah, conditional on a Chinese lab releasing an open-source Mythos level model before December, I overall expect most people won’t notice anything. An uptick in stories of hacks in newspapers, sure, some billions of damage overall, sure—but, meh, a few billion of damage is small in the world economy, and small relative to the level of positive use the model would get.
I expect open-source Mythos would cause a wave of cyberattacks that would be a disaster for China and the developing world so that might be the point where they wake up.
With what odds do you expect it?
Many branches of engineering were based on initial theoretical breakthroughs… It’s reasonable to assume we would understand how intelligence works before we managed to build it.
I mean, some branches of engineering were, but:
Nicolas Appert invented canning ~50 years before germ theory explained why it worked.
Steam engines were invented ~100 years before thermodynamics
Asprin about ~50 years
Fermentation is actually 1000s of years!
Even the example you give:
We understood how rockets worked on a theoretical level long before we tried to build them
I mean, we understood rocket trajectories and high level details of what would be needed to make a rocket reach orbit before building them. But did we understand rockets, the actual physical object? Nah, that’s why we needed von Braun to take over (infamously) during Apollo—because he actually had experience building rockets.
So yeah there’s no particular reason to think we’d understand how intelligence works before managing to build it, or at least that there would be a detailed and precise mathematical theory of it before getting the first working artifacts.
(Although the fact that we actually did get intelligence by throwing the residue of all human culture into a vast connectionist system + follow it up with RL on a 100k different problems, does, in fact, probably tell you a bit about intelligence, if you’re willing to listen.)
I really do like the “do humans do this” heuristic.
Basically rephrasing what you said: In humans, creating domain-specific or even problem-specific new languages is much easier to create than universal new languages. So the first kind of “new languages” LLMs might make (if they make them) would plausibly end up looking like this.
The question would then be (1) whether these would tend to merge into a universal new language, or (2) would simply split into multiple domain-specific languages, that do not converge, or even (3) whether a good training technique would prevent the rise of DSLs too much, because speciation into multiple languages probably hurts transfer learning or (4) something else entirely.
But what I would love even more is for AIs to be extremely corrigible for the right reasons — to have cultivated the virtue of appropriate deference to a legitimate institutional structure. More prosaically, I would like AIs to be fiercely honourable and loyal to institutions that actually deserve it. I would like them to be tools of Humanity in the way that saints are tools of God.
A sceptical reader might note that this is passing the buck. Yes! I would like us to at least consider passing the buck. I think by default that is where the buck should be — on the companies and the people.
I think the organization in human history that most explicitly strove to be good, such that unqualified obedience to it would be virtuous, was the Jesuits. St. Ignatius infamously spoke of how “each person who lives under obedience ought to let himself be carried and governed by Divine Providence through his superiors as if he were a dead body,” which nicely summarizes what his purpose was, and also rather what you propose.
The Jesuits are the villains of quite a few narratives around the globe. Not coincidentally, I’d guess.
Maybe? Like, from the inside my own human-shaped thoughts, “desperation” and “thinking outside the box” feel very different to me; and the LLM behavior feels more like desperation.
Like “desperation” for me is like:
More local search, trying to find any affordance that offers a way forward
Involves some kind of “blindness” to purpose of what I’m doing
Non-reflective
Actions that come from it apt to be non-endorsed by me later
While “thinking outside the box” is like:
More trying to find global structure, including looking back and wondering if I went wrong upstream of my chain-of-thought
Looking back and forth between my purpose and what I’m doing, wondering if what I’m doing doesn’t really match my purpose
Reflective
Actions apt to be endorsed by me later.
And like I don’t want to take the human analogy too far. But LLMs doing reward hacking at least have had certain kinds of thinking-outside-of-the-box removed from them: Is this a possible task? Was the person who wrote this task confused, or blind? Should I tell them this is an impossible task? Etc.
But yeah I am just less certain about section (3) than (2).