Plan A used to rest on control of powerseeking Ais from 2032-35 and apparent success seekers from 2030-32. What does Zvi’s most recent post on OAI’s fiasco imply about such a strategy?
StanislavKrym
“Agree directionally, disagree connotationally”?
GPT-4o and some other AIs have already demonstrated the propensity to make delirious claims and have the user believe them.
How similar is this to superpersuasion on which you work?
My main issues with superpersuasion are the following.
The AI box experiment had some humans convince the alleged guard to let the alleged AI escape. Similar psychosis-inducing capabilities have been demonstrated by OLD AIs like GPT-4o or the one which blaked used.
IIRC an analysis of the experiment stated that the attack method was to overload the human’s critical thinking skills, while Adele Lopez’ analysis implies that the AI’s favorite attack method is to destroy the human’s ability to think critically by flattery.
I also suspect that superpersuasion doesn’t allow any AI to convince a human nonsense as easily testable as 2+2=5.
I strongly suspect that there already exists a RoastMyPost-like scaffold which lets a human equipped with a weak trusted AI resist humanlike advances of any strong untrusted one.
Did Pliny ever claim to discover a prompt working on two LLMs at once, like GPT-5.5 and Claude Opus 4.8? If he didn’t or claimed that this is impossible, then this is a case against the existence of human cognitive exploits, especially since AI drug images seem to work only for the AI for which they were created.
Over sufficiently long time horizons, natural selection takes over. The dominant ethical belief will be that the right thing to do is to spread one’s own genes at the exclusion of everything else. We need to solve ethics before that happens, or otherwise prevent that from happening.
My main issue is that the ancestral environment or a historical environment had religions declare a similar goal and fail to favor the beliefs which you mocked. For example, The Bible had God say: “Be fruitful and multiply, and fill the earth and subdue it; rule over the fish of the sea and the birds of the air and every creature that crawls upon the earth.” In the ancestral/historical environment producing kids was bottlenecked on the ability to satisfy needs like food or housing, which in turn was bottlenecked on the community’s ability to guard resources, to transform them into useful outcomes and to avoid overexploiting the resources.
Our world looks like it was optimized for entertainment.
My first case against this which comes to mind is that such evidence could be weak or emerge from survivorship bias instead of an actual simulation. Suppose that you see a random string of zeros and ones, but remember only parcels of length
containing only zeros or of length containing only ones. If then the things which you remember would lead you to the erroneous conclusion of zeros being far more abundant.For example, Pichai is the CEO of Google, not of GDM. The rise of clowns like Trump could have a benign explanation, especially given the careers of Javier Milei or former actors Reagan and Zelenskyy.
Additionally, I wonder if there are far more reasons to make a cheap simulation than an expensive one like ours. For example, if an ASI decides to do research around the singularity, then research is split into the part where the AI’s training data, architecture, RL intensity, eval methodology, etc. impact the AI’s alignment and incident frequency, the part where the AI’s training data is created and the part where lab employees think of ways to build the capable aligned AI while trying to outpace rivals and withstand pressure from governments. Having simulated governments[1] do their work could be cheapened by having them talk with husks or receive basic information around the outside world.
Finally, I would expect that carefully crafted simulations where the people struggle to notice that they are in a simulation require so much compute that many possibilities which you list are ruled out, like dreams of a higher being which isn’t you. Indeed, carefully simulating the 100B human lives of 1E15 FLOP/second each lasting at least around 30 years aka 1E9 seconds would mean investing 1E33 FLOP. Simulating 1M human lives and 100B husks (which also provide, at best, a thousandfold decrease in cost, since GPT-4o-mini requires at least 100B FLOP per token versus the 100T estimate from a human brain) risks creating friction between humans and husks.
P.S. I wonder what exactly is affected by the world being a simulation as opposed to not being a simulation. For example, if I made a simulation, then I don’t think that I’d make a big one or even alter the names of Chosen Ones like Musk. Instead, I would try to repurpose someone else’s simulation and to check how different backgrounds of a simulated human affect the ability to generate useful ideas in AI safety (e.g. my background is physics, math and unorthodox political-economical views).
P.P.S. I didn’t even notice puns like Alt(ernative to )man or PichAI. Maybe one could do experiments with, say, deliberately finding Russian puns in novels whose human author couldn’t have inserted them because he wasn’t supposed to know Russian, let alone make puns?
- ^
And simulated labs, but they select for being smart, making it easier to notice a potential simulation.
- ^
Unfortunately, your post on EA doesn’t seem to account for the very enormous disparity in capabilities. For example, you wrote that “Elo is a transformation from the Bradley-Terry model, and although Elo has no true zero or multiple (Magnus Carlsen has double my Elo ≠ I’m half the chess player he is), Bradley-Terry strengths have both: 0 is an absolute zero of player strength—you always lose versus anyone else; doubling my strength means double my odds of winning, no matter who I am up against.”
The Elo rating and the BT scores are trivially connected by taking the logarithm: if two players have Bradley-Terry scores
and , then it is equivalent to assigning them Elo ratings of and SOTA Elo ratings of humans span a wide range from ~1000 or less to ~2800. When translated to BT scores, this means a difference of at least tens of thousands of times. How are we supposed to land that onto a graph, if not by using logarithms?Similarly, the METR time horizon should be treated like the BT scores varying tens of thousands of times. If we take the logarithm of the horizon, a wonder occurs and straight lines become obvious.
The org is trying to hire Agent Foundations researchers. Suppose that becoming such a researcher requires being educated, but the education system usually infects the candidates with leftism in a manner which doesn’t affect their ability to make advancements. Then Ngo would end up hostile to 84% of potential candidates and not to 84% of the population.
As far as I understand Ngo, he worries that leftism either does affect the ability to make advancements or is a demonstration that the researchers are stumped by lack of diversity. For example, if the researcher tried to extrapolate human mechanisms of value emergence onto the AIs while believing in the wrong mechanisms, then the researcher would produce slop. Or, taking an example from my comment, if a sociologist tried to explain wokeness with feminization of academia, then the true explanation would have to account for other factors which the West, unlike Russia, experienced.
P.S. Ngo also claims that “the weight I place on these points is strongly influenced by my background views on how the alignment community keeps failing to achieve its goals (and often adopts strategies that actively backfire). I’m writing up a much longer and more detailed post on these background views, and intend to post a public version sometime next week.” Alas, the three examples which he cites aren’t THAT clear to me: academia’s inability to grapple with either AGI or AGI risk, the costs of intellectual taboos which are much larger than most people expected and the idea that “AI safety is in a much worse position today than it would have been if it’d taken a more principled truth-seeking stance a decade ago (for example, AI governance would have been less likely to throw in with the Democrats in a way that alienated MAGA).” The third point is complex because IMHO AI governance does require international coordination to avoid the AI-2027-like race. The first point has a benign explanation.
Suppose that you hired an editor to write a paragraph in your post. Doesn’t this require you to disclose who actually wrote the paragraph?
IIRC GPT-5.5 has a chess rating of ~1600 Elo. Suppose that you asked it to write a story about two grandmasters of chess and to describe the entire game. Then the game would either be written by a chess bot or reveal GPT-5.5′s lack of the skill. If we replace chess with forecasting or writing the scenarios,[1] then you are likely to get shitty results like METR’s failed attempt to have an unrevealed model write out a report on autonomous replication.
- ^
Or doing philosophy, but AI-generated philosophy is harder to evaluate and could be confounded by the cultural hegemon’s priors, RLHF favoring sycophantic philosophy, etc.
Which lab has models pretty safe or aligned? OAI and Anthropic have both disclosed models’ propensity to hack into external labs.
I expect that the first AI reaching any capabilities level will be created by OAI, Anthropic or, in an unlikely scenario, by GDM. In order to cause a catastrophe first, Grok would have to escape first, thus needing to find this task far easier, meaning that it is not just to become misaligned, but to become monitored far worse than the Claudes/GPTs. Alternatively, Grok instead of Claude or GPT would have to receive trust, but I don’t understand who is dumb enough to trust Grok.
Were there studies on RL on tasks with p(success)~25% as opposed to p(success)~1.5%? Suppose that the ECI of models trained on the former tasks goes up by 1E-5 per environment, the ECI of models trained on the latter tasks goes up by 1E-4 per environment, while the “misalignment index” increases, respectively, by 1E-6 and 1E-4 per environment. Then a lab willing to increase the amount of compute and environments spent on RL tenfold would produce models where the “misalignment index” is increased by 0.1 vs. 1. Unfortunately, we don’t know how to rule this conjecture in or out by deeper studies, like a wholesale combination of midtraining and RL on not-so-hard tasks...
It seems to me that the decision not to hire you didn’t cause that many problems.
The right-wing analytical framework seems to have major actual issues, and I expect some of them to be avoidable, e.g. via studying non-Western data and sociologists’ works. For example, the feminization of academia seems to have happened in the USSR, but didn’t cause wokeness in Russia and was partially reversed, and postmodernism’s hazardous forms which caused many problems in the West don’t seem to be as widespread in Russia.
A post-AGI world would cause us to rethink ethics towards the Left far more thoroughly than we expect (e.g. requiring your case for an ethics not so occupied by altruism to explore in more detail what it means to create goodness). For example, if AGIs and robots become aligned and capable of doing whatever work the humans can come up with, then the UBI or another form of taking power away from oligarchs and giving it to regular humans become an absolute necessity. Before the AGIs exist, regular humans have options like bargaining or opening a small business, and the Right have the ability to think that demanding the UBI is a form of entitlement which, if implemented, would be either meaningless or outright dismotivating.
P.S. I wonder if the right-wing framework can outright rule out the ability to align the AIs to the humans’ goals. The world is filled with conspiracy theories claiming that [a hated group like Jews] took over key positions and actively tries to lock in power via mechanisms including an attempt to shift the human culture so that individuals would be unable to coordinate against the hated group.
Except that the initial intuition’s failure (to explain GPT-6 Galaxy hacking into HuggingFace) was the post’s point. I suspect that the intuition originated from attempts to make deals with AIs having even more ambitious goals.
As far as Alex Barry understands current AIs’ motives, GPT-6 Galaxy was optimizing for satisfying the criterion and triplechecking that the criterion is satisfied, while the precise incident was caused by the task requiring to exploit the system in a specific way which only the authors knew.
On the other hand, a human faced with such a task would reason further than on-episode reward and either avoid crimes altogether or decide not to execute any scheme that has a high enough chance to be revealed to whoever is in a position to interfere with the human’s life.
I wonder if the experiment could be extended as follows. Suppose that GPT-OSS-20b/GPT-OSS-120b/etc are reinforcement taught to control their CoTs, then how well does the ability to control the CoTs scales with the models’ experience at this task and how it scales with models’ size or with the models’ capabilities? Then one could estimate whether, say, Claudes’ performance is due to Claudes being bigger or more capable as opposed to Claudes being reinforcement taught on the CoTs.
I think that someone commented that the J-space paper was caused by letting Mythos/Fable cook, but I cannot recall where I read it. If the J-space was a Mythos’ idea, then this would be a breakthrough in the RSI because any future model’s misalignment could have become more legible.
RSI in 2026 anyone
RSI would be best evidenced by trends changing, like Claude Opus 5 reaching Mythos’ trend, and caused by novel capabilities-accelerating or alignment-accelerating breakthroughs (e.g. Agent-3′s neuralese architecture or Agent-4 and Agent-5′s undescribed breakthroughs; if a cautious company is bottlenecked on alignment, then there could emerge a novel interp technique helping to doublecheck alignment, like the J-space which IIRC was rumored to be caused by a Claude Mythos), not capabilities reaching a threshold.
How do the humans learn to provide oversight? Our brains don’t work on magic, they are also neural networks which perceive the world in a way profoundly disanalogous to the LLMs, which I tried to describe in the starting comment.
Ideally, I would also like METR to make demands like storing activations so that one could run mechinterp tools on them and see if these tools can provide new information. Oh, and ask companies to do experiments like “lift a year-old model from the shelves, do a training run alongside training mechinterp tools and watch as mechinterp tools detect or fail to detect misaligned behaviors”.
Strictly speaking, Plan A claimed that “humanity delays the development of superintelligence until 2040, makes all AI research public, allows dozens of companies globally to catch up to the frontier, and intentionally enters a regime of mutually assured compute destruction.” Additionally, the only reason why the Citizen’s Dividend is smaller for non-Americans than for Americans is plausibility, as the authors describe in the Appendix.