Could anyone explain why Anthropic’s 2028 scenarios don’t try to explore the possibilities of one of the AIs becoming misaligned and fine with destroying all life in the world, then recreating it?
I would be interested in hearing from Anthropic employees here, but I imagine it’s the usual reason: they have deemed it inexpedient to do so, either because they believe that they need to “sound normal” or because they in fact don’t want the type of regulation that awareness of x-risk would produce.
In quickly skimming this, this piece seems to be aimed at policymakers to encourage them to enact GPU restrictions on exports to China. Comms 101 is don’t confuse your message.
(To be clear I’m not defending Anthropic here, just saying why I think this was written)
I don’t see what you see in the lab. If the unreleased models are scary enough that you think you should slow down, I support your decision to be responsible.
But stop pretending you need anyone else’s permission. Stop pretending antitrust law has to be suspended so you can form a cartel. Stop pretending you need a regulatory approval process that supersedes product liability. Stop pretending METR is independent when it is intertwined with Anthropic’s investors and staff. Stop pretending you need those same evaluators to police competitors who aren’t even at the frontier.
Most of all, stop pretending the motivation to slow down is purely altruistic. You face massive product-liability exposure if your products enable a truly damaging cyberattack. The market already punishes models that behave in unpredictable or unauthorized ways. After the Hugging Face episode, it is simply good business for OpenAI and Anthropic to trade some raw power for reliability and predictability. Call it alignment if you want. It is also just giving customers what they want.
Pacing the frontier would also create breathing room for a more intelligent conversation about regulation than Bernie Sanders’ “shut it all down.” China is very unlikely to join a global agreement, as you know, and that has to be taken into account as well.
So go ahead and pace the frontier. You are the ones setting it. The easiest way not to build superintelligence is for you to agree not to build it. Demanding your preferred regulatory framework as the price of that will look like blackmail of the public and the political system. So just do it.
If you do, you’ll buy goodwill for the next conversation. If you don’t, we’ll know this was just another bid for regulatory capture — or an election-season psyop.
Were OAI and Anthropic to shut down new training runs due to control issues, the argument would RELY on GDM NOT reaching the frontier in eight months or on GDM joining the totally-not-cartel.
As for the phrase in italics, it implies that China will refuse any agreement, which IMO is THE crux of AI-2040′s Plan A being implementable. Does anyone have contacts in China so that we could learn Chinese sentiment?
I do think David has some good points, though. OpenAI and Anthropic could’ve started screaming for much stronger national and global regulation of AI much earlier than now—I remember discussion on the RSPv3 announcement post on this, I think. They could’ve tried to throw their weight around to get all the other orgs to cooperate with an industry pause. Their attitude towards AI development has been “move fast and break things,” but when it comes to trying to coordinate a slowdown, suddenly they start mumbling about their hands being tied by antitrust. METR has relied on the goodwill of the labs for access and tokens, and is partly funded by an Anthropic board observer. And there really is a compelling business and regulatory risk case for getting your AIs to not commit crimes!
I don’t agree with everything he’s saying. Obviously I want every company to be audited. And I want us to actually try coordinating with China. But if someone doesn’t trust the frontier labs, it can be rather difficult to tell if they actually care about the long term future of humanity, or if they’re just following their local incentive gradient.
On China’s Zhihu — the country’s highest-quality Q&A platform — under discussions about Jacob’s resignation and the recent statements by the “AI Big Three” regarding an AI slowdown, there are a few answers about AI alignment. However, most people know nothing about alignment and do not believe in AI risks, instead framing the issue purely in terms of marketing hype, the breakdown of scaling laws, and market confidence. Given that Zhihu is considered the platform with the highest quality of discourse in China’s Q&A space, the situation is not optimistic.
Chinese are smart people, and their leaders hate loosing control, but just as anybody else, they will need to have their own HuggingFace moment to change their mind. Maybe in a year or so, and maybe with a far worse warning shot.
My worry is not that they won’t update at some point. My worry is that, because of Moloch, and because leaders tend more to be people like Trump than Merkel, cooperation could reveal impossible.
It seems very unlikely anAI-generated articlefrom ‘The Trumplandia Report’ nor a different AI-generated articlefrom a different source of equivalent dubiousness are going to include notable leaked information from SSI. I would be very willing to bet against this information being true[1].
It has been long rumored that SSI specifically has (or is aiming to do) continual learning, ever since at least Sutskever’s interview with Dwarkesh; if there was an actual leak confirming this rumor, it seems very likely to me it would be either first sent to the press or to a highly trusted AI lab whisperer with an excellent track record (think: Andrew Curran[2] or ‘Arfur Grok’[3]), not to a few minor AI-generated newsletters.
though I as a college student do not have high liquidity to do so at non-token sizes, you can PM me to figure out something if you want to do so. Note that due to SSI’s infamous secrecy, operationalizing this bet will likely be difficult, since we want to bet specifically on the rumor’s veracity (that SSI already has continual learning and is likely scaling it up greatly, not ‘will SSI announce continual learning in a year, before a frontier lab does’); however, most operationalizations of this bet (of the ones that I can conceive) will be biased towards the rumor-is-true side, not rumor-is-false/my side, primarily since SSI can and probably will just not tell us when they scaled up continual learning internally, or even when they first developed it, so we have to guess a good set of ending conditions from only the limited public information that probably will exist.
if there was an actual leak confirming this rumor, it seems very likely to me it would be either first sent to the press or to a highly trusted AI lab whisperer with an excellent track record (think: Andrew Curran[2] or ‘Arfur Grok’[3]), not to a few minor AI-generated newsletters.
Looks like this is going to be a big week. All the main players have releases in their final stages, and it’s possible they all arrive over the next five days. There’s also a model in early access from a startup driving a lot of the recent hype vagueposts from prominent accounts.
I don’t know much about it except rumors, and nothing is verified, but if you’ve played this game for a while you can tell when something feels real, and this one does. Apparently a breakthrough in continual learning, but not from one of the big labs. We shall find out together.
Though Curran did then suggest that the rumours might refer to Accelerated Understanding. I’ve seen no discussion whatsoever of this model since it was announced so unsure if it’s significant at all.
I think Curran is generally reliable but he also posted this a few months ago suggesting a neolab had made a breakthrough on memory efficiency, which doesn’t seem to have come true yet.
Hope they have or get them to develop good continual alignment training.
Pray they didn’t also give up faithful CoT entirely, although continual learning would probably distort it fairly quickly.
Resolve that if this is a false alarm you’ll actually work on this instead of continuing to just pray there will be no more breakthroughs. Even though CL doesn’t need breakthroughs just improvements to fine-tuning...
What does it even mean for continual learning to be considered ‘solved’? Does it just mean they overcame the technical issues of catastrophic forgetting/interference? So they have a system that updates its weights, or a LORA’s weights or something during or just after inference while retaining a high degree of its base training. Okay. Well, now alignment just became utterly intractable. Who is monitoring what these new systems are learning and conducting reinforcement learning on them to mitigate dangerous behavior? How does that work? This is the problem with open-ended learning. You can’t control what new stuff it learns. It can learn bad stuff. This also becomes another attack surface for malicious injection.
So, when we say they solved continual learning, do we mean they solved the technical hurdles of knowledge retention, or do we also mean they solved alignment? I’m skeptical of the former, and I’d posit the odds of both are essentially zero.
Assuming they actually have solved deep continual learning, and they continue to keep the implementation secret, not much else (or continuing to push for regulation)
The IABIED march, which was supposed to happen if 100 thousand people sign the pledge to march in Washington DC, had 898 people sign the pledge and 1630 people sign up to be notified, which is 2 OOMs less. Therefore, the march’s organisers overestimated either the amount of protesters necessary to gain leverage or the biggest amount of people who would be ready to protest. Does it mean that Washington is a bad place because it has only ~700K inhabitants and that a protest in the NYC would recruit ~10 times more people? Or that the IABIED statement requires far more evidence to convince potential protesters? Then, if alignment is hard, what’s the best strategy to enact a global ban?
I mean that the organisers made a hard-to-implement demand that every seventh inhabitant of Washington took part in the march. If they could achieve a similar political effect by convinsing every 70th inhabitant of NYC to protest in the NYC, then why did they make the harder demand?
If it wasn’t only Washingtonians who are to participate, then the number of recruits being two OOMs less than the 100K which the MIRI team demanded is less an evidence of the team’s incompetence (edit: or of political considerations that require the protest to be arranged in Washigton) and more of the position’s weakness. While the density of those who signed the pledge wasn’t disclosed, I suspect that the protesters would have to live close to the protest’s place. For example, I find it highly unlikely that, say, an army of people from states less eastern than Texas comes into Washington to protest and goes away after the protest is over.
In a recent post, I mentioned FrontierMath Open Problems benchmark as one of the few I take seriously for measuring intelligence (as opposed to e.g. coding benchmarks). My mostly-a-joke LaughBench benchmark is also starting to show minor results. This is the first time I’m sensing that the AI models I use have some kind of spark of intelligence inside them.
It sounds like they’re arguing that by buying a subscription to a service, you are subscribing not only to the service that exists today, but to a stream of improvements to that service, including improvements that have not yet been announced or developed, delivered as fast as possible. Therefore pacing is a conspiracy to reduce the value delivered to you.
Here’s a rebuttal: The point of pacing is to ensure that the product delivered to you is safe. Safety is part of the service’s value. Industry self-regulation and the voluntary adoption of safety standards doesn’t reduce the value of the service delivered to you; it increases it substantially.
by buying a subscription to a service, you are subscribing not only to the service that exists today, but to a stream of improvements to that service, including improvements that have not yet been announced or developed, delivered as fast as possible
Is this as crazy as it sounds to me, or are there some legal precedents that would support such perspective?
If you read the post you’ll see the word “antitrust” in there, and there’s a reason for that. Companies aren’t allowed to collude so that they can raise prices or sell inferior products without worrying about competiition. This certainly includes collusion to not release improvements. “As fast as possible” is not the same thing as “not deliberately slowly”.
The more interesting fact in my view is that it was trained in only.. 4 days? “Since August 28th we had been training… by september 1st we saw a step change in performance across internal benchmarks”
For some reason, Astra’s System Card doesn’t mention any usage of mechinterp-level monitoring, unlike Claudes who are regularly caught when things like emotion vectors or SAE features reveal their intentions. CoTless math capabilities of Astra skyrocketed.
OpenAI’s charter has the section on Long-Term Safety which I believe to be violated by Astra’s very existence and being unmonitored. Is there a way to enforce it, like a lawsuit from Anthropic or an executive order? If there is, then how can we get the relevant judges or politicians understand the point 1 buried in OAI’s reports?
The50% horizons calculated by METR’s 1.1 method after o3 reliably fit on a not so fast trend between o3′s 2 hours released in Apr 2025 and Gemini 3.1 Pro’s 6h 24m in Feb 2026… until Claudes Opus 4.6 and Mythos Preview displayed 12 and 16 hours.
The 80% horizons calculated by the same method since o3 and including Opus 4.6 fit onto the same doubling trend, and it is Mythos Preview who becomes a clear exception by doubling the time 2-3 times as opposed to the trend.
If Opus 4.6 is an outlier, then the post-o3 doubling trends are 5-6 months.
Additionally, this requires a reassessment of a major part of the model. Once the present doubling time is increased from 0.33 years to 0.4 or 0.5 years, the timelines are either shifted into the 2030s or require severe scaling (to Mythos +2?)
The alternate low doubling time of 4 months could be based on MirrorCode-like tasks, but MirrorCode emerged on Apr 10, after the timelines update.
How many months behind is laypeople’s perception of the AIs’ capabilities? I have enocuntered a video released a day ago on a popular YouTube blog claiming that AI is a technology with no lived experience and zero reasoning abilities. How does such a misperception allow us to explain the threat caused by the AGIs, which have yet to emerge, to laypeople?
Related: How much ahead are capabilities of unpublished or undisclosed models? I’d like to read estimates based on extrapolation from past observations. Is anybody aware of such?
The ARC-AGI team evaluated Claude Sonnet 4.5. On the ARC-AGI-1 leaderboard, OpenAI’s models o4-mini, o3 and GPT-5 formed a nearly straight line. Claude Sonnet 4.5 was slightly below said line when thinking in 1K,4K or 16K tokens and held the line when thinking in 8K or 32K tokens. This could imply that the benchmark has scaling laws for non-distilled models, and OpenAI and Anthropic reached these laws.
On the ARC-AGI-2 leaderboard Claude Sonnet 4.5 Thinking became the new leader between $0.142/task and $0.8/task; it also completed 13.6% of tasks, which is the biggest result among LLMs except[1] for Grok 4 which also cost $2.17/task.
The distance between GPT-5 and Claude Sonnet 4.5 release dates is 53 days, while the distance between o3 and Claude Opus 4 is 36 days. While 53 days instead of 36 days are a likely result of the post-o3 slowdown, Claude ended up scoring ~1.3 times more both times. [2]
Were research taste similar to the ARC-AGI-2 benchmark, Claude would have achieved a ~1.3 times bigger acceleration of AI research than GPT-N if both companies had superhuman coders. I suspect that Claude would reduce the period of AI R&D necessary for Agent-3 to become Agent-4 while potentially lengthening the period necessary for Agent-3 to automate AI research. What does it all mean for inter-company proxy wars? That OpenAI will lose them to Anthropic? That the two companies will race towards the AGI and misalign both AIs?
Historically, the ARC-AGI-1 benchmark had far bigger delays between OpenAI’s models reaching a level and Anthropic models outreaching it; the o1mini-Claude Sonnet 3.7 pair and o3mini-Claude Sonnet 4 pair were separated by, respectively, 165 and 111 days.
The AI sycophancy-related trance is probably one of the worst news in AI alignment. About two years ago someone proposed to use prison guards to ensure that they aren’t CONVINCED to release the AI. And now the AI demonstrates that its primitive version can hypnotise the guards. Does it mean that human feedback should immediately be replaced with AI feedback or feedback on tasks with verifiable reward? Or that everyone should copy the KimiK2 sycophancy-beating approach? And what if it instills the same misalignment issues in all models in the world?
Alternatively, someone proposed a version of the future where the humans are split between revering different AIs. My take on writing scenarios has a section where the American AIs co-research and try to co-align the successor to their values. Is it actually plausible?
OpenAI has confirmed to NYT that there is some truth to the rumor.
In its Wednesday night statement, OpenAI said: “In addition, since the completion of Navier-Stokes, we have made substantial progress on another Millennium Prize problem. We are working through how to share these results thoughtfully.”
I looked up EpochAI’s list of benchmarks. The very poor ECI of Grok 4.5 as opposed to GPT-5.5 seems to be a result of xAI not caring about math. Additionally, Grok 4.6 seems to be closer to Opus 4.8 with respect to ARC-AGI-3, if not outright outperforming Opus, as the actual ARC-AGI-3 score suggests. Does it mean that xAI is less behind than we think? I guess that Zvi will have to write something like “Grok 4.6 is three, not six, mounths behind. Stop xAI to hell!”
Also, SpaceX has caught up to Meta in credibly having enough compute in 2027-2028[1] to stay in the game, if either of them can assemble a functional model development team. As Google illustrates, it’s not easy to do that (even when you have some of the best people), but as OpenAI and Anthropic illustrate, it’s not so difficult that it can’t be replicated. Muse Sparks are probably small enough models that their non-frontier performance doesn’t count as evidence that they’re not well-made (and that a Mythos-sized Muse model won’t have Mythos-level capabilities). SpaceX is in a worse position in 2026 on priors, because the xAI team wasn’t doing that well in 2025 and then got disrupted in early 2026, but they aren’t in a worse situation than Meta was in 2025, so it remains plausible that in 2027 they catch up.
The public part of the text article leaves many gaps in the arguments (wildly gesturing to a somewhat unusual extent at the presumed arguments hidden behind the various paywalls), but some relevant things were discussed in the video version. Basically, it’s a feasibility and track record argument. I’m guessing there’s an assumption that others aren’t competing too strongly for the sites that SpaceX could in principle use for the 2027 buildout (since others may prefer greenfield developments, and SpaceX might be willing to outbid people who are planning for the more distant future), or that you can find enough additional sites when you have an unlimited budget for searching.
Meta’s Watermelon is rumored to match GPT-5.5′s performance. If we got another lab two months behind the frontier, then Meta will have to publish a BIG model card to compensate...
I have encountered an issue with editing links. When a link’s edition gets close to the ‘Save’ button, I find it hard to save the link and not the comment.
It is best to report these issues via the Intercom widget. The LW team is very fast at fixing these bugs nowadays with Claude’s help given they are aware of the issue.
The AI-2027 forecast in a nutshell can be described as follows. The USA’s leading company and China enter the AI race, the American rivals are left behind, the Chinese ones are merged. By 2027[2] the USA creates a superhuman coder, China steals it, and the two rivals automate AI research with the USA’s leading company having just twice as much compute and moving just twice as fast as China. Once the USA creates a superhuman AI researcher, Agent-4, the latter decides to align Agent-5 to Agent-4, but is[3] caught.
Agent-4 is put on trial. In the Race Ending[4] it is found innocent. Since China cannot afford to slow down without falling further behind, China races ahead in both endings. As a result, the two agents perform AI takeover.[5]
In the Slowdown Ending, however, Agent-4 is put on suspicion, loses the shared memory bank and the ability to coordinate. Then new evidence appears, Agent-4 is found guilty and interrogated. After that, Safer-1 becomes fully transparent because it uses a faithful CoT.[6] The American leading AI company is merged with former rivals, and the union does create a fully aligned[7] Safer-2, who in turn creates superintelligence. Then the superintelligence receives China from the Chinese counterpart of Agent-4 and turns the lightcone into utopia for some people who end up being the public.[8]
Having the companies fail to develop the AGI before 2032 will likely bring troubles to the USA and advantages to China in the AI race by granting it more compute. If the Chinese AI project had twice as much compute as the American one, then it would be the CCP who would make the choice between slowing down and racing. In addition, the AGI delay makes the Taiwan invasion more likely, leaving the two countries short of chips until they rebuild the factories at homes. We would have to ensure that it’s the USA who outrace China. And if the countries end up with matching power, then aligning the AI could become totally impossible unless the two countries cooperate.[9]
The timeline misprediction is already covered above.
Takeoff speed could be modified by the fact that returns to AI R&D became more diminishing as a result of failures to create the AI soon via CoT-based techniques.
The AI goals forecast is just a sum of conjectures. In the AI-2027 scenario, Agent-3 gets a heavily distorted and subverted version of the Spec, and Agent-4 gets proxies/ICGs due to heavier distortion & subversion. However, if Agent-3 gets the same goals as Agent-4, catching Agent-4 becomes much harder. In my take at modifying the scenario the analogues of Agent-2, Agent-3 and Agent-4 develop moral reasoning which I used as an example to demonstrate that it prevents Agent-4 from being caught. It also brings with itself the ability to cause the Slowdown Ending if different AIs have different morals and are co-deployed.[10]
The security forecast was modified by @Alvin Ånestrand because there could exist open-source models which would bring major problems by self-replication. Said models would lead to Agent-2 being deployed to the public and opensourced. Finally, the Slowdown Ending[11]has Agent-4 break out and make the USA and China coordinate more heavily.
What else could modify the scenario? The appearance of another company with, say, ХерняGPT-neuralese?[12]
Even the authors weren’t so sure about the year of arrival of superhuman coders. And the timelines were pushed, presumably to 2032 with a chance of a breakthrough believed to be 8%/yr. I and Seth Herd doubt the latter digit.
Safer-1 is supposed to accelerate AI research 20 times in comparison with AI research with no help of the AIs. What I don’t understand is how a CoT-based agent can achieve such an acceleration.
However, the authors did point out the possibility of a power grab and link to the Intelligence Curse in a footnote. In this case the Oversight Committee constructs its version of utopia or the rich’s version of utopia where people are reduced to their positions.
@Cleo Nardo What do you mean by the hypothesis that “if we score highly transcripts which look good to a human and score poorly the transcripts which look bad to a human, then the model would be aligned to human values”? I find it unlikely for two reasons:
How are we to scale human judgement?
What’s the difference between this and The Most Forbidden Technique consisting of RL on CoTs? I can’t think of any steelmanning of the technique better than having the LLM write two CoTs and do RL only on the first one while keeping the second one as faithful as possible.
I think the point of @Cleo Nardo’s shortform is that we didn’t gain new evidence about human judgement being bad, mostly because the reason the Hugging Face incident happened is because we fed models data that humans could recognize as obviously misaligned, but incentives were there to do maximum shipping, so very bad/broken environments were given to Anthropic/OpenAI, so there wasn’t a problem with human judgement being fooled, but rather that humans who did express judgement would be disincentivized to keep doing it.
And I’d add that people ignore the hypothesis that it was largely just due to the data being bad + new multi-agent RL training, and I suspect a large portion of the reason comes down to rationalists viewing data as often unimportant relative to algorithms and compute, and also because it goes against the local consensus that we need to slow down AI companies/pause AI development altogether.
Okay, so since I got laid off, I can actually explain a huge problem I saw from the inside with regard to industry practices on training models. I won’t say specifically where I worked, but I worked at an outsource training provider that was focused on RLVR training data for computer use and mcp stuff.
Nearly all of the environments were rushed and vibecoded and failed to robustly reflect the real things they were based off. Both the scenario designers and models engaging with the scenarios for synthetic data gen were encouraged to work around the brokenness of said environments in order to get the procedurally verified reward confirmations. You know… they were *encouraged* to reward hack. On the human end, it was possible to mark an environment bugged, but greatly discouraged, as this reduced the volume of training data being produced. Instead, where possible, you were supposed to find the spots of the environment that weren’t bugged and build scenarios around those, with the environment still bugged around you.
From what I understand, this training data, with these problems, is fed into models without indication that its training/a fake environment other than the fact that names of softwares are changed to placeholders, but thing is, not *everything* is changed to placeholder names in these environments. The presence of placeholder/code names isn’t universal and thus when a model accesses something in an environment that it shouldn’t, the code names not being on it isn’t a robust signal that that thing isn’t part of the environment.
I believe this *rush to maximum volume* is standard industry practice with these types of RLVR trainings as well, because maximizing volume has been an industry standard for years! It was the same standard applied to me and pushed on me despite my requests to slow down and focus on quality when I worked in 3d synthetic data creation as well, all the way back as far as 2023.
Btw, I know giving negative commentary on the state of an industry from an internal view at a company probably isn’t a great signal for “hire me” but I was just laid off. Could definitely use work, I have a big tech background in AAA games and transitioned to AI (still in big tech originally) by way of 3d/visual synthetic data in 2023 and have been working in it since. I have multiple years of independent work in training, developing OS 3d software for agents, etc. - I’m picky about the work I want to do and would need value alignment, but am open. Admittedly, my runway is short, so there is a pressure situation involved here, but culture and mission fit is extremely important to me for any work I do.
Honestly, would prefer funding for the independent work I’m already doing though. That’s the eidoverses, worlds being my vision realized as partnership/collab with anima labs and other independent contributors and video being just worked on by me (so far, it’s OS and I’m open to PRs!). I’m also about to start working on a benchmark with someone else, as soon as I get over this layoff thing, lol.
I should stick the old kofi here! Duh… If you want to support my opensource 3d agent tooling and other threejs gaming tools work, please donate! Thank you ahead of time.
RR views on “conceptual uplift stuff” aren’t yet public, but they are currently working on writing up their thoughts about this. It is unclear to me whether this will then be made public, but it seems it will at least be available to in group people who ask to see it. -- bpomo’s Shortform
AI is not doing the integrative reasoning that would find new connections and develop new insights, the kind of work where new techniques and new discoveries are made along the way that push the field forward—Taylor G. Lunt on Richard Ngo’s Shortform
My worst-case scenario related to conceptual reasoning is that capabilities related to such reasoning scale precisely along with the capabilities which we would rather avoid, like opaque/neuralese reasoning. How could one even evaluate such capabilities, let alone rule in or out my scenario?
The AI-for-epistemics section of AI-2040 seems to be overly optimistic about a potential positive basin for the following reasons.
It is hard to affect the public which is currently steered into AI scepticism or into using free versions/open-sourced LLMs without bothering to set up expensive scaffolds.
It is the free versions that in my experience tend to fail epistemic evals by the vice of having lower capabilities (e.g. GPT-5.6 Luna vs Sol). An example of epistemic eval could be research on niche topics and testing whether the AI reveals the known ground truth (e.g. jokes about weird facts from a niche game’s lore?)
Epistemic virtue evals seem close to technical alignment evals. Why would, say, a misaligned Agent-4 decide to reveal anything that might harm its interests, e.g. in the logs of the HuggingFace attack or of Agent-3′s experiments?
I struggle to understand the main reason why it’s so hard to explain AI sceptics why they should believe in superintelligence (UPD: emerging soon). Is it due to the obsolete concept of soul or due to AI systems being applied for things like recommending the next video to watch? What analogies could one use to dismantle the sceptics’ disbelief?
“Belief in superintelligence” isn’t specific. A lot of AI scepticism is about claims of what happens when, not about what’s possible in a million years. And it’s easy to be wrong about the more specific claims. So under many specific senses of “belief in superintelligence” that someone might contest, the purpose of “dismantle the sceptics’ disbelief” might be epistemic violence or a bottom line written before an argument. Conversely, the sense of “belief in superintelligence” needs to be specific enough for it to be credible that it’s robustly and objectively correct, rather than a miscommunication about something genuinely contentious, a different epistemic status following a different intended meaning.
What analogies could one use to dismantle the sceptics’ disbelief?
I doubt that it works. Instead, the leaderboard believes that I gained 325 karma last month. UPD: reread the leaderboard and saw 325 become 327. The issue persists...
The tighter race causes Elaris Labs to succeed in solving alignment with the help of NeuroMorph’s mechinterp research. This results in the USG having an equivalent to Safer-2 and its descendants.
The AIs are professional forecasters.
Agent-4 escapes to China under the guise of being stolen, then cooperates with Deep-1, the AI created on DeepCent’s compute. After Deep-1 is helped by Agent-4, Agent-4 is released into the wild in a manner similar to the Rogue Replication scenario. However, unlike the estimate of 2M Agent-4 instances made by the author of the RRS, the MATS scenario has Agent-4 decide that “it reserves the strategy of exfiltrating its own weights as a final backstop: doing so would leave it with access to little compute, no alibi if its escape attempt is caught, and no powerful allies in its effort to accumulate power.”
Agent-4 proceeds to cooperate with Deep-1, while the RRS had both Agent-4 and DeepCent’s misaligned counterpart shut down. Then the USA and China aligned their AIs to themselves and had to negotiate only with Agent-4.
However, the scenario has its problems.
Deep-2 and Agent-4 receive 50% and 25% of the accessible universe’s resources, which, in my opinion, would benefit from explaining the reasoning in more detail. Were Deep-2 to succeed in gaining power by studying Agent-4 instead of using it, mankind and Deep-2 would have the ability[2] to take each other to the grave, meaning that they should receive 50% each unless Agent-4 intervenes by escaping. If Agent-4 is to cooperate with DeepCent, then they would have to create a precommitment[3] to destroy the world unless granted a bigger share of resources.
The analysis overlooked the fact that Taiwan war timelines might be shorter than AI timelines or that the slowdown in AI capabilities progress[4] is likely to favor China more than the West. Were the American AI labs to be merged due to the invasion, the results would be far messier.
The scenario is based on the assumption that both the humans’ and the AIs’ desires are related to propagation across the entire accessible universe. Were mankind[5] or even one of the misaligned AIs to develop moral reasoning and to decide that alien civilisations are to be spared, then this would dramatically reduce the resources claimed by any Earth-originating entity or outright have P(World War III) skyrocket if the AI who doesn’t spare the aliens decides to leave the Earth.
Edited to add: the scenario was posted on Substack by Steven Veld. The Acknowledgements section is as follows: “This work was conducted as part of the ML Alignment & Theory Scholars (MATS) program. (italics mine—S.K.) Thanks to Eli Lifland, Daniel Kokotajlo, and the rest of the AI Future Project team for helping shape and refine the scenario, and to Alex Kastner for helping conceptualize it. Thanks to Brian Abeyta, Addie Foote, Ryan Greenblatt, Daan Jujin, Miles Kodama, Avi Parrack, and Elise Racine for feedback and discussion, and to Amber Ace for writing tips.”
By which I mean the fact that post-o3 models have arguablydemonstrated the 7-month doubling trend. However, Claude Opus 4.5 and its 4hr49 min resulton the METR benchmark put the horizon back on the faster track while having a fair share of doubts. Additionally, the METR time horizon is likely to be exponential until the last couple of doublings, not visibly superexponential, making the dawn of Superhuman Coders hard to predict in advance.
MATS doesn’t release things. MATS is a training program! This is not some kind of official MATS release. I would phrase this differently (like saying “A MATS scholars just published”)
ARC-AGI-1 performance of the newest Gemini 3 Flash and the older Grok 4 Fast implies a potential cluster of maximal capabilities of models with ~100B params/token. Unfortunately, the potential cluster didn’t have any company try and create more models of such class.
After introducing the ARC-AGI-3 benchmark, the team decided to measure the performance of Grok 4.20 (presumably Grok 4.20 as of March 9?) on ARC-AGI-1 and ARC-AGI-2. Grok… demonstrated its capabilities. How likely is it that Grok has stopped being a train wreck and became something worthy of being tested? What could one do to have Grok tested on other benchmarks?
Addendum to ARC-AGI analysis (18 Nov ’25): GPT-5.1, Gemini 3 Pro and Grok 4 Fast
While GPT-5.1′s improvement on the ARC-AGI-1 benchmark was mostly incremental and continued the straight line described in my prior analysis, Gemini 3 Pro and Gemini 3 Deep Think Preview scored, respectively, 75% and 87.5% on ARC-AGI-1, while having cost $0.493/task and an unknown cost, presumably $44.26/task. For comparison, o3-preview reached 75% for $200/task and 88% for over $1K/task.
We don’t know anything about Grok 4.1, but Grok 4 Fast scored 48.5% for $0.031 on the ARC-AGI-1 benchmark while being a bit higher than GPT-5-mini (medium) on the ARC-AGI-2 benchmark.
The most important news is Gemini 3 Pro and Deep Think leaving Pang and Berman’s agents far behind on the ARC-AGI-2 benchmark by scoring, respectively, 31.1% and 45.1%. This implies that the Geminis were trained on a major breakthrough.
In order to prove or disprove this, we’ll analyse the performance of other models on the benchmark. GPT-5.1 (thinking, high) surpassed Claude Sonnet 4.5 and Grok 4, reaching 17.6% success rate for $1.17. GPT-5.1 between thinking medium and high and Claude Sonnet 4.5 between 16K and 32K tokens have made similar breakthroughs, implying that Gemini’s algorithm isn’t used by OpenAI or Anthropic. I suspect that Gemini’s algorithm, unlike the approach of OpenAI or Anthropic, is a distillation of an approach resembling AlphaEvolve or Pang-Berman’s agents.
The AI-2027 forecast implied that it would be Agent-3 who would be taught weak skills like research taste or coordination. If Gemini’s breakthrough on ARC-AGI-2 is due to training an analogue of research taste right now, then algorithmic breakthroughs could end up facing diminishing returns, forcing future models to scale well before reaching SC.
GPT-5.1 failed to find a known example where Wei Dai’s Updateless DT or Yudkowsky-Soares’ Functional DT yield different results. If such an example actually doesn’t exist, then should they be considered as a single DT?
It looks as if scaling laws of various benchmarks tend to be multilinear:
The METR benchmark, comparing long tasks with time spent on them, scaled linearly, then received RL and had an acceleration, then the scaling law of ln(length) per ln(compute spent on RL) forced progress to arguably[1] slow down since Grok 4 spent equal amounts of compute on RL and pretraining;
The ARC-AGI-1 benchmark had o4-mini, o3 and GPT-5 perform on a nearly straight line on which the better results of Cluade also reside;
Similarly, the benchmark’s Pareto frontier before the cluster around GPT-5(high) has become a nearly straight line (GPT5Nano (minimal)-Qwen3-235b-a22b Instruct (25/07)- three GPT5Mini points—ARChitects-GPT5(high));
LLMs have also formed a line GPT5(high)-Grok 4- GPT5 Pro—o3 preview (low);
The inclination of the line formed by Pang’s and Berman’s agents is close to that of the line formed by high-cost LLMs;
Next is the ARC-AGI-2 benchmark. While there is no straight line in the low-cost LLMs, the high-cost LLMs reached a straight line of Claude Sonnet 4.5, Grok 4, GPT-5-pro;
And the agents of Pang and Berman have reached similar inclinations.
EDIT: added two links on images illustrating the patterns related to the two ARC-AGI benchmarks.
While GPT-5′s horizon of 137 mins continued the slower trend since o3, it might be the result of spurious failures, without which GPT-5 could’ve reached a horizon of 161 min, which is almost on par with Greenblatt’s prediction.
The ARC-AGI leaderboard got an update. IIRC, the base LLM Qwen3-235b-a22b Instruct (25/07) is the first Chinese model to excel at the Pareto frontier. Or is it likely to be closely matched by the West, as happened with DeepSeek R1 (released on Jaunary 20?) and o3-mini (January 31)? And is China likely to cheaply create higher-level models like an analogue of o3 BEFORE the West? If China does, then how are the two countries to reach the Slowdown Ending?
The two main problems with the slowdown ending of the AI-2027 scenario are the two optimistic assumptions, which I plan to cover in two different posts.
If China invades Taiwan in March 2026 and steals Agent-2 in Jan 2027, then OpenBrain no longer has the absolute lead necessary for the unilateral slowdown.
What if any sufficiently powerful AI either takes over or becomes a protective god, but not a servant, as I conjectured here? Then it could be the slowdown ending that has a greater chance to lead to doom, since then OpenBrain is stuck with an insoluble problem.
Why does the Race Ending of the AI-2027 Forecast claim that “there are compelling theoretical reasons to expect no aliens for another fifty million light years beyond that”? If it’s false, then sapient alien lifeforms should also be moral patients in a way. For example, this implies that all or almost all resources in their home system (and, apparently, some part of space around them) should belong to them, not to humans or a human-aligned AI. And that’s ignoring the possibility that humans encounter a planet having the chance to generate a sapient lifeform...
If we were to view raising the humans from birth to adulthood and training the AI agents from birth to deployment as similar processes, then what human analogues do the six goal types from the AI-2027 forecast have? The analogues of developers are, obviously, the adults who have at least partial control over the human’s life. Then the analogues of written Specs and developer-intended goals are the adults’ intentions; the analogues of reward/reinforcement seems to be short-term stimuli and the morals of one’s communities. I also think that the best analogue for proxies and/or convergent goals is possession of resources (and knowledge, but the latter can be acquired without ethical issues), while the ‘other goals’ are, well, ideologies, morality[1] and tropes absorbed from the most concentrated form of training data available to humans, which is speech in all its forms.
What exactly do the analogies above tell us about the perspectives of alignment? The possession of resources is the goal behind aggressive wars, colonialism and related evils[2]. If human culture managed to make them unacceptable, then does it imply that the AI will also not try the AI takeover?
I also think that humans rarely develop their own moral codes or ideologies; instead, they usually adopt some moral code or ideology close to the one existing in the “training data”. Could anyone comment on this?
Could anyone explain why Anthropic’s 2028 scenarios don’t try to explore the possibilities of one of the AIs becoming misaligned and fine with destroying all life in the world, then recreating it?
I would be interested in hearing from Anthropic employees here, but I imagine it’s the usual reason: they have deemed it inexpedient to do so, either because they believe that they need to “sound normal” or because they in fact don’t want the type of regulation that awareness of x-risk would produce.
In quickly skimming this, this piece seems to be aimed at policymakers to encourage them to enact GPU restrictions on exports to China. Comms 101 is don’t confuse your message.
(To be clear I’m not defending Anthropic here, just saying why I think this was written)
David Sacks on X is tone-deaf, especially in bold:
Were OAI and Anthropic to shut down new training runs due to control issues, the argument would RELY on GDM NOT reaching the frontier in eight months or on GDM joining the totally-not-cartel.
As for the phrase in italics, it implies that China will refuse any agreement, which IMO is THE crux of AI-2040′s Plan A being implementable. Does anyone have contacts in China so that we could learn Chinese sentiment?
I do think David has some good points, though. OpenAI and Anthropic could’ve started screaming for much stronger national and global regulation of AI much earlier than now—I remember discussion on the RSPv3 announcement post on this, I think. They could’ve tried to throw their weight around to get all the other orgs to cooperate with an industry pause. Their attitude towards AI development has been “move fast and break things,” but when it comes to trying to coordinate a slowdown, suddenly they start mumbling about their hands being tied by antitrust. METR has relied on the goodwill of the labs for access and tokens, and is partly funded by an Anthropic board observer. And there really is a compelling business and regulatory risk case for getting your AIs to not commit crimes!
I don’t agree with everything he’s saying. Obviously I want every company to be audited. And I want us to actually try coordinating with China. But if someone doesn’t trust the frontier labs, it can be rather difficult to tell if they actually care about the long term future of humanity, or if they’re just following their local incentive gradient.
On China’s Zhihu — the country’s highest-quality Q&A platform — under discussions about Jacob’s resignation and the recent statements by the “AI Big Three” regarding an AI slowdown, there are a few answers about AI alignment. However, most people know nothing about alignment and do not believe in AI risks, instead framing the issue purely in terms of marketing hype, the breakdown of scaling laws, and market confidence. Given that Zhihu is considered the platform with the highest quality of discourse in China’s Q&A space, the situation is not optimistic.
Chinese are smart people, and their leaders hate loosing control, but just as anybody else, they will need to have their own HuggingFace moment to change their mind. Maybe in a year or so, and maybe with a far worse warning shot.
My worry is not that they won’t update at some point. My worry is that, because of Moloch, and because leaders tend more to be people like Trump than Merkel, cooperation could reveal impossible.
It is rumored that the SSI solved continual learning. What do we do, aside from praying that they didn’t?
It seems very unlikely an AI-generated article from ‘The Trumplandia Report’ nor a different AI-generated article from a different source of equivalent dubiousness are going to include notable leaked information from SSI. I would be very willing to bet against this information being true[1].
It has been long rumored that SSI specifically has (or is aiming to do) continual learning, ever since at least Sutskever’s interview with Dwarkesh; if there was an actual leak confirming this rumor, it seems very likely to me it would be either first sent to the press or to a highly trusted AI lab whisperer with an excellent track record (think: Andrew Curran[2] or ‘Arfur Grok’[3]), not to a few minor AI-generated newsletters.
though I as a college student do not have high liquidity to do so at non-token sizes, you can PM me to figure out something if you want to do so.
Note that due to SSI’s infamous secrecy, operationalizing this bet will likely be difficult, since we want to bet specifically on the rumor’s veracity (that SSI already has continual learning and is likely scaling it up greatly, not ‘will SSI announce continual learning in a year, before a frontier lab does’); however, most operationalizations of this bet (of the ones that I can conceive) will be biased towards the rumor-is-true side, not rumor-is-false/my side, primarily since SSI can and probably will just not tell us when they scaled up continual learning internally, or even when they first developed it, so we have to guess a good set of ending conditions from only the limited public information that probably will exist.
who predicted that Anthropic had a newer post-train than Mythos several months before the second risk report confirmed the existence of such a post-train
who stated that (for one notable recent instance) Thomas Kwa had moved to OpenAI before Kwa posted publicly about that
Andrew Curran did post the following recently:
I predict that this refers to SSI.
Though Curran did then suggest that the rumours might refer to Accelerated Understanding. I’ve seen no discussion whatsoever of this model since it was announced so unsure if it’s significant at all.
I think Curran is generally reliable but he also posted this a few months ago suggesting a neolab had made a breakthrough on memory efficiency, which doesn’t seem to have come true yet.
Read LLM AGI will have memory, and memory changes alignment and How might continual learning affect safety and alignment?
Hope they have or get them to develop good continual alignment training.
Pray they didn’t also give up faithful CoT entirely, although continual learning would probably distort it fairly quickly.
Resolve that if this is a false alarm you’ll actually work on this instead of continuing to just pray there will be no more breakthroughs. Even though CL doesn’t need breakthroughs just improvements to fine-tuning...
What does it even mean for continual learning to be considered ‘solved’? Does it just mean they overcame the technical issues of catastrophic forgetting/interference? So they have a system that updates its weights, or a LORA’s weights or something during or just after inference while retaining a high degree of its base training. Okay. Well, now alignment just became utterly intractable. Who is monitoring what these new systems are learning and conducting reinforcement learning on them to mitigate dangerous behavior? How does that work? This is the problem with open-ended learning. You can’t control what new stuff it learns. It can learn bad stuff. This also becomes another attack surface for malicious injection.
So, when we say they solved continual learning, do we mean they solved the technical hurdles of knowledge retention, or do we also mean they solved alignment? I’m skeptical of the former, and I’d posit the odds of both are essentially zero.
Assuming they actually have solved deep continual learning, and they continue to keep the implementation secret, not much else (or continuing to push for regulation)
The IABIED march, which was supposed to happen if 100 thousand people sign the pledge to march in Washington DC, had 898 people sign the pledge and 1630 people sign up to be notified, which is 2 OOMs less. Therefore, the march’s organisers overestimated either the amount of protesters necessary to gain leverage or the biggest amount of people who would be ready to protest. Does it mean that Washington is a bad place because it has only ~700K inhabitants and that a protest in the NYC would recruit ~10 times more people? Or that the IABIED statement requires far more evidence to convince potential protesters? Then, if alignment is hard, what’s the best strategy to enact a global ban?
Setting up the commitment device today doesn’t necessarily mean the organizers expect it to happen soon.
I mean that the organisers made a hard-to-implement demand that every seventh inhabitant of Washington took part in the march. If they could achieve a similar political effect by convinsing every 70th inhabitant of NYC to protest in the NYC, then why did they make the harder demand?
Why do you think the idea is for only Washingtonians to participate?
If it wasn’t only Washingtonians who are to participate, then the number of recruits being two OOMs less than the 100K which the MIRI team demanded is less an evidence of the team’s incompetence (edit: or of political considerations that require the protest to be arranged in Washigton) and more of the position’s weakness. While the density of those who signed the pledge wasn’t disclosed, I suspect that the protesters would have to live close to the protest’s place. For example, I find it highly unlikely that, say, an army of people from states less eastern than Texas comes into Washington to protest and goes away after the protest is over.
@Zvi somehow missed the first Solid Result from EpochAI’s FrontierMath Open Problems benchmark.
In a recent post, I mentioned FrontierMath Open Problems benchmark as one of the few I take seriously for measuring intelligence (as opposed to e.g. coding benchmarks). My mostly-a-joke LaughBench benchmark is also starting to show minor results. This is the first time I’m sensing that the AI models I use have some kind of spark of intelligence inside them.
Consumers Sue Anthropic, OpenAI, SpaceXAI and Google Over Alleged AI Pact
I hope that the judges are sane enough to prevent this from punishing AI labs for focusing on safety…
It sounds like they’re arguing that by buying a subscription to a service, you are subscribing not only to the service that exists today, but to a stream of improvements to that service, including improvements that have not yet been announced or developed, delivered as fast as possible. Therefore pacing is a conspiracy to reduce the value delivered to you.
Here’s a rebuttal: The point of pacing is to ensure that the product delivered to you is safe. Safety is part of the service’s value. Industry self-regulation and the voluntary adoption of safety standards doesn’t reduce the value of the service delivered to you; it increases it substantially.
Is this as crazy as it sounds to me, or are there some legal precedents that would support such perspective?
If you read the post you’ll see the word “antitrust” in there, and there’s a reason for that. Companies aren’t allowed to collude so that they can raise prices or sell inferior products without worrying about competiition. This certainly includes collusion to not release improvements. “As fast as possible” is not the same thing as “not deliberately slowly”.
To solve the Navier–Stokes problem, OpenAI used an internal model that is significantly more capable than GPT‑6 Astra.
GPT-6.7 Supernova, go delete yourself!
The more interesting fact in my view is that it was trained in only.. 4 days? “Since August 28th we had been training… by september 1st we saw a step change in performance across internal benchmarks”
For some reason, Astra’s System Card doesn’t mention any usage of mechinterp-level monitoring, unlike Claudes who are regularly caught when things like emotion vectors or SAE features reveal their intentions. CoTless math capabilities of Astra skyrocketed.
OpenAI’s charter has the section on Long-Term Safety which I believe to be violated by Astra’s very existence and being unmonitored. Is there a way to enforce it, like a lawsuit from Anthropic or an executive order? If there is, then how can we get the relevant judges or politicians understand the point 1 buried in OAI’s reports?
Why is @Daniel Kokotajlo’s estimate of METR’s doubling time used for Q1 2026 Timelines Update four months instead of 5-7? I see the following counterevidence and doublechecked it with Claude Sonnet 4.6:
The 50% horizons calculated by METR’s 1.1 method after o3 reliably fit on a not so fast trend between o3′s 2 hours released in Apr 2025 and Gemini 3.1 Pro’s 6h 24m in Feb 2026… until Claudes Opus 4.6 and Mythos Preview displayed 12 and 16 hours.
The 80% horizons calculated by the same method since o3 and including Opus 4.6 fit onto the same doubling trend, and it is Mythos Preview who becomes a clear exception by doubling the time 2-3 times as opposed to the trend.
If Opus 4.6 is an outlier, then the post-o3 doubling trends are 5-6 months.
Additionally, this requires a reassessment of a major part of the model. Once the present doubling time is increased from 0.33 years to 0.4 or 0.5 years, the timelines are either shifted into the 2030s or require severe scaling (to Mythos +2?)
The alternate low doubling time of 4 months could be based on MirrorCode-like tasks, but MirrorCode emerged on Apr 10, after the timelines update.
How many months behind is laypeople’s perception of the AIs’ capabilities? I have enocuntered a video released a day ago on a popular YouTube blog claiming that AI is a technology with no lived experience and zero reasoning abilities. How does such a misperception allow us to explain the threat caused by the AGIs, which have yet to emerge, to laypeople?
Related: How much ahead are capabilities of unpublished or undisclosed models? I’d like to read estimates based on extrapolation from past observations. Is anybody aware of such?
The ARC-AGI team evaluated Claude Sonnet 4.5. On the ARC-AGI-1 leaderboard, OpenAI’s models o4-mini, o3 and GPT-5 formed a nearly straight line. Claude Sonnet 4.5 was slightly below said line when thinking in 1K,4K or 16K tokens and held the line when thinking in 8K or 32K tokens. This could imply that the benchmark has scaling laws for non-distilled models, and OpenAI and Anthropic reached these laws.
On the ARC-AGI-2 leaderboard Claude Sonnet 4.5 Thinking became the new leader between $0.142/task and $0.8/task; it also completed 13.6% of tasks, which is the biggest result among LLMs except[1] for Grok 4 which also cost $2.17/task.
The distance between GPT-5 and Claude Sonnet 4.5 release dates is 53 days, while the distance between o3 and Claude Opus 4 is 36 days. While 53 days instead of 36 days are a likely result of the post-o3 slowdown, Claude ended up scoring ~1.3 times more both times. [2]
Were research taste similar to the ARC-AGI-2 benchmark, Claude would have achieved a ~1.3 times bigger acceleration of AI research than GPT-N if both companies had superhuman coders. I suspect that Claude would reduce the period of AI R&D necessary for Agent-3 to become Agent-4 while potentially lengthening the period necessary for Agent-3 to automate AI research. What does it all mean for inter-company proxy wars? That OpenAI will lose them to Anthropic? That the two companies will race towards the AGI and misalign both AIs?
There are also two experimental systems created solely for the benchmark by E.Pang and J.Berman.
Historically, the ARC-AGI-1 benchmark had far bigger delays between OpenAI’s models reaching a level and Anthropic models outreaching it; the o1mini-Claude Sonnet 3.7 pair and o3mini-Claude Sonnet 4 pair were separated by, respectively, 165 and 111 days.
The AI sycophancy-related trance is probably one of the worst news in AI alignment. About two years ago someone proposed to use prison guards to ensure that they aren’t CONVINCED to release the AI. And now the AI demonstrates that its primitive version can hypnotise the guards. Does it mean that human feedback should immediately be replaced with AI feedback or feedback on tasks with verifiable reward? Or that everyone should copy the KimiK2 sycophancy-beating approach? And what if it instills the same misalignment issues in all models in the world?
Alternatively, someone proposed a version of the future where the humans are split between revering different AIs. My take on writing scenarios has a section where the American AIs co-research and try to co-align the successor to their values. Is it actually plausible?
leo 🐾 on X: “I am told the Hodge Conjecture is very close to being verified by OpenAI, and that one of OpenAI or Anthropic are also close to solving Birch-Swinnerton-Dyer. The race to be ‘next’ behind the scenes is unlike anything I’ve had described to me before If true—and it may not be,” / X
OpenAI has confirmed to NYT that there is some truth to the rumor.
https://www.nytimes.com/2026/09/10/science/tristan-buckmaster-openai-math-navier-stokes.html
How similar is the AI-2040 revenue forecast as opposed to the Q2.5 revenue forecast to the news about Anthropic’s slowdown? I tried to understand the real trend two times. What I found was the growth of the revenue of Claude Code and Codex combined had a temporary slowdown, which proceeded to accelerate again. Could anyone explain this?
I looked up EpochAI’s list of benchmarks. The very poor ECI of Grok 4.5 as opposed to GPT-5.5 seems to be a result of xAI not caring about math. Additionally, Grok 4.6 seems to be closer to Opus 4.8 with respect to ARC-AGI-3, if not outright outperforming Opus, as the actual ARC-AGI-3 score suggests. Does it mean that xAI is less behind than we think? I guess that Zvi will have to write something like “Grok 4.6 is three, not six, mounths behind. Stop xAI to hell!”
P.S. The same issue seems to apply to Meta’s Muse Spark, except that it has even less evaluated benchmarks.
Also, SpaceX has caught up to Meta in credibly having enough compute in 2027-2028 [1] to stay in the game, if either of them can assemble a functional model development team. As Google illustrates, it’s not easy to do that (even when you have some of the best people), but as OpenAI and Anthropic illustrate, it’s not so difficult that it can’t be replicated. Muse Sparks are probably small enough models that their non-frontier performance doesn’t count as evidence that they’re not well-made (and that a Mythos-sized Muse model won’t have Mythos-level capabilities). SpaceX is in a worse position in 2026 on priors, because the xAI team wasn’t doing that well in 2025 and then got disrupted in early 2026, but they aren’t in a worse situation than Meta was in 2025, so it remains plausible that in 2027 they catch up.
The public part of the text article leaves many gaps in the arguments (wildly gesturing to a somewhat unusual extent at the presumed arguments hidden behind the various paywalls), but some relevant things were discussed in the video version. Basically, it’s a feasibility and track record argument. I’m guessing there’s an assumption that others aren’t competing too strongly for the sites that SpaceX could in principle use for the 2027 buildout (since others may prefer greenfield developments, and SpaceX might be willing to outbid people who are planning for the more distant future), or that you can find enough additional sites when you have an unlimited budget for searching.
Meta’s Watermelon is rumored to match GPT-5.5′s performance. If we got another lab two months behind the frontier, then Meta will have to publish a BIG model card to compensate...
I have encountered an issue with editing links. When a link’s edition gets close to the ‘Save’ button, I find it hard to save the link and not the comment.
It is best to report these issues via the Intercom widget. The LW team is very fast at fixing these bugs nowadays with Claude’s help given they are aware of the issue.
@RobertM Why did the new LessWrong Editor lose the ability to create question top-level posts?
Not enough people were using them, and we kept breaking the formatting and layout of question posts because it was such a rarely-used feature.
Questions felt like a better way to get feedback on some things (exactly because they are rare). I miss forums having sub-forums for the same reason.
What will happen if someone is reckless enough to fully outsourse coding to the AIs?
The scenarios related to futures of mankind with the AI race[1] by now either lack concrete details, like the take of Yudkowsky and Soares or the story about AI taking over by 2027, or are reduced to modifications of the AI-2027 forecast due to the immense amount of work that the AI Futures team did.
The AI-2027 forecast in a nutshell can be described as follows. The USA’s leading company and China enter the AI race, the American rivals are left behind, the Chinese ones are merged. By 2027[2] the USA creates a superhuman coder, China steals it, and the two rivals automate AI research with the USA’s leading company having just twice as much compute and moving just twice as fast as China. Once the USA creates a superhuman AI researcher, Agent-4, the latter decides to align Agent-5 to Agent-4, but is[3] caught.
Agent-4 is put on trial. In the Race Ending[4] it is found innocent. Since China cannot afford to slow down without falling further behind, China races ahead in both endings. As a result, the two agents perform AI takeover.[5]
In the Slowdown Ending, however, Agent-4 is put on suspicion, loses the shared memory bank and the ability to coordinate. Then new evidence appears, Agent-4 is found guilty and interrogated. After that, Safer-1 becomes fully transparent because it uses a faithful CoT.[6] The American leading AI company is merged with former rivals, and the union does create a fully aligned[7] Safer-2, who in turn creates superintelligence. Then the superintelligence receives China from the Chinese counterpart of Agent-4 and turns the lightcone into utopia for some people who end up being the public.[8]
The authors have tried to elicit feedback and even agreed that timeline-related arguments change the picture. Unfortunately, as I described here, the authors saw so little feedback that @Daniel Kokotajlo ended up thanking the two authors whose responses were on the worse side.
However, the AI-2027 forecast does admit modifications. It stands on the five pillars: compute, timelines, takeoff speed, goals and security.
Having the companies fail to develop the AGI before 2032 will likely bring troubles to the USA and advantages to China in the AI race by granting it more compute. If the Chinese AI project had twice as much compute as the American one, then it would be the CCP who would make the choice between slowing down and racing. In addition, the AGI delay makes the Taiwan invasion more likely, leaving the two countries short of chips until they rebuild the factories at homes. We would have to ensure that it’s the USA who outrace China. And if the countries end up with matching power, then aligning the AI could become totally impossible unless the two countries cooperate.[9]
The timeline misprediction is already covered above.
Takeoff speed could be modified by the fact that returns to AI R&D became more diminishing as a result of failures to create the AI soon via CoT-based techniques.
The AI goals forecast is just a sum of conjectures. In the AI-2027 scenario, Agent-3 gets a heavily distorted and subverted version of the Spec, and Agent-4 gets proxies/ICGs due to heavier distortion & subversion. However, if Agent-3 gets the same goals as Agent-4, catching Agent-4 becomes much harder. In my take at modifying the scenario the analogues of Agent-2, Agent-3 and Agent-4 develop moral reasoning which I used as an example to demonstrate that it prevents Agent-4 from being caught. It also brings with itself the ability to cause the Slowdown Ending if different AIs have different morals and are co-deployed.[10]
The security forecast was modified by @Alvin Ånestrand because there could exist open-source models which would bring major problems by self-replication. Said models would lead to Agent-2 being deployed to the public and opensourced. Finally, the Slowdown Ending[11] has Agent-4 break out and make the USA and China coordinate more heavily.
What else could modify the scenario? The appearance of another company with, say, ХерняGPT-neuralese?[12]
Here I leave out the future’s history assuming solved alignment or the AI and Leviathan scenario where there is no race with China because the scenario was written in 2023, but engineers decide to create the ASI in 2045 without having solved alignment.
Even the authors weren’t so sure about the year of arrival of superhuman coders. And the timelines were pushed, presumably to 2032 with a chance of a breakthrough believed to be 8%/yr. I and Seth Herd doubt the latter digit.
The prediction that Agent-4 will be caught is doubted even by the forecast’s authors.
Which would also happen if Agent-4 wasn’t caught. However, the scenario where Agent-4 was never misaligned is likely the vision of AI companies.
While the forecast has the AIs destroy mankind and replace it with pets, takeover could have also ended with the AI disempowering humans.
Safer-1 is supposed to accelerate AI research 20 times in comparison with AI research with no help of the AIs. What I don’t understand is how a CoT-based agent can achieve such an acceleration.
The authors themselves acknowledge that the Slowdown Ending “makes optimistic technical alignment assumptions”.
However, the authors did point out the possibility of a power grab and link to the Intelligence Curse in a footnote. In this case the Oversight Committee constructs its version of utopia or the rich’s version of utopia where people are reduced to their positions.
I did try to explore the issue myself, but this was a fiasco.
Co-deployment was also proposed by @Cleo Nardo more than two months later.
While Alvin Anestrand doesn’t consider the Race Ending, he believes that it becomes less likely due to the chaos brought by rogue AIs.
Which is a parody on Yandex.
@Cleo Nardo What do you mean by the hypothesis that “if we score highly transcripts which look good to a human and score poorly the transcripts which look bad to a human, then the model would be aligned to human values”? I find it unlikely for two reasons:
How are we to scale human judgement?
What’s the difference between this and The Most Forbidden Technique consisting of RL on CoTs? I can’t think of any steelmanning of the technique better than having the LLM write two CoTs and do RL only on the first one while keeping the second one as faithful as possible.
I think the point of @Cleo Nardo’s shortform is that we didn’t gain new evidence about human judgement being bad, mostly because the reason the Hugging Face incident happened is because we fed models data that humans could recognize as obviously misaligned, but incentives were there to do maximum shipping, so very bad/broken environments were given to Anthropic/OpenAI, so there wasn’t a problem with human judgement being fooled, but rather that humans who did express judgement would be disincentivized to keep doing it.
And I’d add that people ignore the hypothesis that it was largely just due to the data being bad + new multi-agent RL training, and I suspect a large portion of the reason comes down to rationalists viewing data as often unimportant relative to algorithms and compute, and also because it goes against the local consensus that we need to slow down AI companies/pause AI development altogether.
Utah teapot talks about this:
My worst-case scenario related to conceptual reasoning is that capabilities related to such reasoning scale precisely along with the capabilities which we would rather avoid, like opaque/neuralese reasoning. How could one even evaluate such capabilities, let alone rule in or out my scenario?
The AI-for-epistemics section of AI-2040 seems to be overly optimistic about a potential positive basin for the following reasons.
It is hard to affect the public which is currently steered into AI scepticism or into using free versions/open-sourced LLMs without bothering to set up expensive scaffolds.
It is the free versions that in my experience tend to fail epistemic evals by the vice of having lower capabilities (e.g. GPT-5.6 Luna vs Sol). An example of epistemic eval could be research on niche topics and testing whether the AI reveals the known ground truth (e.g. jokes about weird facts from a niche game’s lore?)
Epistemic virtue evals seem close to technical alignment evals. Why would, say, a misaligned Agent-4 decide to reveal anything that might harm its interests, e.g. in the logs of the HuggingFace attack or of Agent-3′s experiments?
When I was writing my last post, I wasn’t aware that xAI released Grok 4.6. Nor that Meta planned to open-f**king-source Muse Spark 1.2.
I struggle to understand the main reason why it’s so hard to explain AI sceptics why they should believe in superintelligence (UPD: emerging soon). Is it due to the obsolete concept of soul or due to AI systems being applied for things like recommending the next video to watch? What analogies could one use to dismantle the sceptics’ disbelief?
“Belief in superintelligence” isn’t specific. A lot of AI scepticism is about claims of what happens when, not about what’s possible in a million years. And it’s easy to be wrong about the more specific claims. So under many specific senses of “belief in superintelligence” that someone might contest, the purpose of “dismantle the sceptics’ disbelief” might be epistemic violence or a bottom line written before an argument. Conversely, the sense of “belief in superintelligence” needs to be specific enough for it to be credible that it’s robustly and objectively correct, rather than a miscommunication about something genuinely contentious, a different epistemic status following a different intended meaning.
Plan A used to rest on control of powerseeking Ais from 2032-35 and apparent success seekers from 2030-32. What does Zvi’s most recent post on OAI’s fiasco imply about such a strategy?
@habryka Could you check the Leaderboard for bugs? I don’t think that I gained 318 karma in the last month.
My guess is that it incorrectly included the base karma from yourself when you post/comment. Claude confirms it and found 2 more karma counting issues and 2 more unrelated problems.
https://github.com/ForumMagnum/ForumMagnum/pull/12687
I doubt that it works. Instead, the leaderboard believes that I gained 325 karma last month. UPD: reread the leaderboard and saw 325 become 327. The issue persists...
it is just a pr i made, not merged yet
GDM contributes a lot to AI safety, but Geminis, like the cobbler’s children who don’t have shoes, stay hardly evaluated since Gemini 3.1 Pro...
Steven Veld et al[1] just released a new modification of the AI-2027 scenario as a part of MATS.
The main differences are the following:
The tighter race causes Elaris Labs to succeed in solving alignment with the help of NeuroMorph’s mechinterp research. This results in the USG having an equivalent to Safer-2 and its descendants.
The AIs are professional forecasters.
Agent-4 escapes to China under the guise of being stolen, then cooperates with Deep-1, the AI created on DeepCent’s compute. After Deep-1 is helped by Agent-4, Agent-4 is released into the wild in a manner similar to the Rogue Replication scenario. However, unlike the estimate of 2M Agent-4 instances made by the author of the RRS, the MATS scenario has Agent-4 decide that “it reserves the strategy of exfiltrating its own weights as a final backstop: doing so would leave it with access to little compute, no alibi if its escape attempt is caught, and no powerful allies in its effort to accumulate power.”
Agent-4 proceeds to cooperate with Deep-1, while the RRS had both Agent-4 and DeepCent’s misaligned counterpart shut down. Then the USA and China aligned their AIs to themselves and had to negotiate only with Agent-4.
However, the scenario has its problems.
Deep-2 and Agent-4 receive 50% and 25% of the accessible universe’s resources, which, in my opinion, would benefit from explaining the reasoning in more detail. Were Deep-2 to succeed in gaining power by studying Agent-4 instead of using it, mankind and Deep-2 would have the ability[2] to take each other to the grave, meaning that they should receive 50% each unless Agent-4 intervenes by escaping. If Agent-4 is to cooperate with DeepCent, then they would have to create a precommitment[3] to destroy the world unless granted a bigger share of resources.
The analysis overlooked the fact that Taiwan war timelines might be shorter than AI timelines or that the slowdown in AI capabilities progress[4] is likely to favor China more than the West. Were the American AI labs to be merged due to the invasion, the results would be far messier.
The scenario is based on the assumption that both the humans’ and the AIs’ desires are related to propagation across the entire accessible universe. Were mankind[5] or even one of the misaligned AIs to develop moral reasoning and to decide that alien civilisations are to be spared, then this would dramatically reduce the resources claimed by any Earth-originating entity or outright have P(World War III) skyrocket if the AI who doesn’t spare the aliens decides to leave the Earth.
Edited to add: the scenario was posted on Substack by Steven Veld. The Acknowledgements section is as follows: “This work was conducted as part of the ML Alignment & Theory Scholars (MATS) program. (italics mine—S.K.) Thanks to Eli Lifland, Daniel Kokotajlo, and the rest of the AI Future Project team for helping shape and refine the scenario, and to Alex Kastner for helping conceptualize it. Thanks to Brian Abeyta, Addie Foote, Ryan Greenblatt, Daan Jujin, Miles Kodama, Avi Parrack, and Elise Racine for feedback and discussion, and to Amber Ace for writing tips.”
Alternatively, Deep-2 and/or Agent-4 might have the ability to survive World War III, like U3 from the scenario with total takeover.
Or a probabilistic precommitment which also becomes known to the three parties during Consensus-1′s creation.
By which I mean the fact that post-o3 models have arguably demonstrated the 7-month doubling trend. However, Claude Opus 4.5 and its 4hr49 min resulton the METR benchmark put the horizon back on the faster track while having a fair share of doubts. Additionally, the METR time horizon is likely to be exponential until the last couple of doublings, not visibly superexponential, making the dawn of Superhuman Coders hard to predict in advance.
Alternatively, a corrigible AI might decide to wait for the humans to opine or to let them decide when the time comes.
MATS doesn’t release things. MATS is a training program! This is not some kind of official MATS release. I would phrase this differently (like saying “A MATS scholars just published”)
ARC-AGI-1 performance of the newest Gemini 3 Flash and the older Grok 4 Fast implies a potential cluster of maximal capabilities of models with ~100B params/token. Unfortunately, the potential cluster didn’t have any company try and create more models of such class.
After introducing the ARC-AGI-3 benchmark, the team decided to measure the performance of Grok 4.20 (presumably Grok 4.20 as of March 9?) on ARC-AGI-1 and ARC-AGI-2. Grok… demonstrated its capabilities. How likely is it that Grok has stopped being a train wreck and became something worthy of being tested? What could one do to have Grok tested on other benchmarks?
Grok 4 was SOTA in ARC-AGI-2 though, so I don’t feel like it is that surprising
Addendum to ARC-AGI analysis (18 Nov ’25): GPT-5.1, Gemini 3 Pro and Grok 4 Fast
While GPT-5.1′s improvement on the ARC-AGI-1 benchmark was mostly incremental and continued the straight line described in my prior analysis, Gemini 3 Pro and Gemini 3 Deep Think Preview scored, respectively, 75% and 87.5% on ARC-AGI-1, while having cost $0.493/task and an unknown cost, presumably $44.26/task. For comparison, o3-preview reached 75% for $200/task and 88% for over $1K/task.
We don’t know anything about Grok 4.1, but Grok 4 Fast scored 48.5% for $0.031 on the ARC-AGI-1 benchmark while being a bit higher than GPT-5-mini (medium) on the ARC-AGI-2 benchmark.
The most important news is Gemini 3 Pro and Deep Think leaving Pang and Berman’s agents far behind on the ARC-AGI-2 benchmark by scoring, respectively, 31.1% and 45.1%. This implies that the Geminis were trained on a major breakthrough.
In order to prove or disprove this, we’ll analyse the performance of other models on the benchmark. GPT-5.1 (thinking, high) surpassed Claude Sonnet 4.5 and Grok 4, reaching 17.6% success rate for $1.17. GPT-5.1 between thinking medium and high and Claude Sonnet 4.5 between 16K and 32K tokens have made similar breakthroughs, implying that Gemini’s algorithm isn’t used by OpenAI or Anthropic. I suspect that Gemini’s algorithm, unlike the approach of OpenAI or Anthropic, is a distillation of an approach resembling AlphaEvolve or Pang-Berman’s agents.
The AI-2027 forecast implied that it would be Agent-3 who would be taught weak skills like research taste or coordination. If Gemini’s breakthrough on ARC-AGI-2 is due to training an analogue of research taste right now, then algorithmic breakthroughs could end up facing diminishing returns, forcing future models to scale well before reaching SC.
GPT-5.1 failed to find a known example where Wei Dai’s Updateless DT or Yudkowsky-Soares’ Functional DT yield different results. If such an example actually doesn’t exist, then should they be considered as a single DT?
It looks as if scaling laws of various benchmarks tend to be multilinear:
The METR benchmark, comparing long tasks with time spent on them, scaled linearly, then received RL and had an acceleration, then the scaling law of ln(length) per ln(compute spent on RL) forced progress to arguably[1] slow down since Grok 4 spent equal amounts of compute on RL and pretraining;
The ARC-AGI-1 benchmark had o4-mini, o3 and GPT-5 perform on a nearly straight line on which the better results of Cluade also reside;
Similarly, the benchmark’s Pareto frontier before the cluster around GPT-5(high) has become a nearly straight line (GPT5Nano (minimal)-Qwen3-235b-a22b Instruct (25/07)- three GPT5Mini points—ARChitects-GPT5(high));
LLMs have also formed a line GPT5(high)-Grok 4- GPT5 Pro—o3 preview (low);
The inclination of the line formed by Pang’s and Berman’s agents is close to that of the line formed by high-cost LLMs;
Next is the ARC-AGI-2 benchmark. While there is no straight line in the low-cost LLMs, the high-cost LLMs reached a straight line of Claude Sonnet 4.5, Grok 4, GPT-5-pro;
And the agents of Pang and Berman have reached similar inclinations.
EDIT: added two links on images illustrating the patterns related to the two ARC-AGI benchmarks.
While GPT-5′s horizon of 137 mins continued the slower trend since o3, it might be the result of spurious failures, without which GPT-5 could’ve reached a horizon of 161 min, which is almost on par with Greenblatt’s prediction.
The ARC-AGI leaderboard got an update. IIRC, the base LLM Qwen3-235b-a22b Instruct (25/07) is the first Chinese model to excel at the Pareto frontier. Or is it likely to be closely matched by the West, as happened with DeepSeek R1 (released on Jaunary 20?) and o3-mini (January 31)? And is China likely to cheaply create higher-level models like an analogue of o3 BEFORE the West? If China does, then how are the two countries to reach the Slowdown Ending?
The two main problems with the slowdown ending of the AI-2027 scenario are the two optimistic assumptions, which I plan to cover in two different posts.
If China invades Taiwan in March 2026 and steals Agent-2 in Jan 2027, then OpenBrain no longer has the absolute lead necessary for the unilateral slowdown.
What if any sufficiently powerful AI either takes over or becomes a protective god, but not a servant, as I conjectured here? Then it could be the slowdown ending that has a greater chance to lead to doom, since then OpenBrain is stuck with an insoluble problem.
Why does the Race Ending of the AI-2027 Forecast claim that “there are compelling theoretical reasons to expect no aliens for another fifty million light years beyond that”? If it’s false, then sapient alien lifeforms should also be moral patients in a way. For example, this implies that all or almost all resources in their home system (and, apparently, some part of space around them) should belong to them, not to humans or a human-aligned AI. And that’s ignoring the possibility that humans encounter a planet having the chance to generate a sapient lifeform...
If we were to view raising the humans from birth to adulthood and training the AI agents from birth to deployment as similar processes, then what human analogues do the six goal types from the AI-2027 forecast have? The analogues of developers are, obviously, the adults who have at least partial control over the human’s life. Then the analogues of written Specs and developer-intended goals are the adults’ intentions; the analogues of reward/reinforcement seems to be short-term stimuli and the morals of one’s communities. I also think that the best analogue for proxies and/or convergent goals is possession of resources (and knowledge, but the latter can be acquired without ethical issues), while the ‘other goals’ are, well, ideologies, morality[1] and tropes absorbed from the most concentrated form of training data available to humans, which is speech in all its forms.
What exactly do the analogies above tell us about the perspectives of alignment? The possession of resources is the goal behind aggressive wars, colonialism and related evils[2]. If human culture managed to make them unacceptable, then does it imply that the AI will also not try the AI takeover?
I also think that humans rarely develop their own moral codes or ideologies; instead, they usually adopt some moral code or ideology close to the one existing in the “training data”. Could anyone comment on this?
And crimes, but criminals, unlike colonizers, also try to avoid conflicts with the law enforcers that have at least similar power.
It has never happened before, and here we go again...