Another thought just occurred to me, extending from my above point: when it comes to task misalignment, effective (automated) control is effective alignment. If you have reliable enough control to stop a model from going rogue, then you can also stop it from cheating during training, which prevents the task misalignment from happening in the first place.
2001zhaozhao
I agree with most of this article. However, I want to point out an objection towards Q6, specifically category C. In my opinion, task misalignment seems very tractable to reliably detect and penalize by a sufficiently capable LLM judge and interpretability tools (you’re using full surveillience and maybe even mind reading to make sure that the model cannot get away with cheating, and detection of cheating in that way intuitively seems much easier than getting away with cheating, so at any given capability level the monitors seem likely to have an advantage, and plausibly this is only more-and-more-so the case as interpretability methods improve). If that works, it would be a valid, non-brittle example of aligning a model to a “trained machine learning model” as described in Category C.
I do agree with the rest of Q6 though. It seems like to get full alignment you have to align a model towards an entire ethics system as the objective function, and that system has to be fairly complex (in the sense of well-thought-out) and internally consistent enough to avoid both Categories A and B. This basically points to a constitutional AI approach like what Anthropic is doing.
In addition I think that because the type of misalignment described in this article is rooted in task misalignment, if someone starts failing at it we (or at least they) will expect to see warning shots like the OpenAI/HF incident before anything truly catastrophic starts to happen. So that could be an optimistic sign that people failling to align AIs will be given a chance to change course.
This article does update me positively on the possibility that a sufficiently capable aligned model could coperate with its creators to increase its own capabilities arbitrarily far without becoming misaligned. (i.e. how easy it would be to safely “hand off the world” to a capable, trusted aligned model, if we’re sure we have one)
I tried to think of a few counterexamples (wouldn’t the model just take over the world if it is needed to stay good, since it would incentivize the bad reward-seeking part of itself to do it otherwise?), but they are categorically prevented if the model has some degree of control over its own RL training and can prevent the misaligned trajectory from being rewarded in the first place.
The concept is simple as well—the model is just doing the same thing as a good person purposefully learning new skills. So intuitively it seems pretty doable to me.
AI destroys earth and turns it into paperclips when asked to mass produce paperclips.“But it was just following instructions!”
I have noticed that some of my Claude Code chats are not getting thinking summaries but others are getting it still, and the difference seems to be per-chat. If they fully remove it it would be a bit of a shame, tbh. These summaries are quite useful.
I think the problem boils down to the claim that training on a specific probe makes that probe useless, but it doesn’t make other uncorrelated probes (especially stronger ones) useless.
For example, if the claim is true, you can train on the CoT and then, say, J-lens probes will keep working. So you lose the ability to do meaning alignment checks on CoT but you can still use J-lens to do them. Thus, from an alignment monitoring perspective, training on CoT is as bad as switching to neuralese which removes CoT completely, but it could be a good tradeoff if you still have other interpretability/monitoring methods available. The real trouble comes if you are irresponsible and exhaust all of your reliable methods by training on all of them.
My own thoughts:
I think the first 5 answers are false (as in I agree with OP).
I think answer 6 is also false. In fact I think that highly competitive automated AI management systems are possible with today’s LLMs. My biggest evidence is that the best AI systems are human-level at real-world epistemics per evidence in the superforecasting field: https://www.astralcodexten.com/p/the-ai-superforecasters-are-here. That plus the speed, context volume, and organizational alignment advantage points to an AI rational reasoning system being possible that beats a human organization (specifically the entire management hierarchy, not an individual human) at management decision-making. Basically I think human management hierarchies pay a big epistemic cost for the benefit of parallelizing / scaling up decision making due to misaligned incentives, and the advantage of AI is that it scales while avoiding this cost entirely, so it could beat a human management hierarchy in decision-making at surprisingly low levels of base model capability, while matching their decision-making bandwidth. (It would still be far worse than a top individual human at decision quality but massively outmatch them in volume. Kind of like playing against a bot opponent in a strategy game that is lousy at macro but can constantly micro every single unit across the entire map.)
(Edit: however, if “competent CEO” was replaced by “superhuman CEO” in answer 6, I think my above argument would no longer apply, since my reasoning predicts that a very good human CEO + AI management would stay more competent than AI management alone for a long time, and therefore the AI CEO may not be able to replace all human CEOs prior to a hypothetical permanent shutdown deal.)I think the second part in answer 7 is true, but I disagree that only misaligned AIs that kill/disempower humans would ditch the concept of a company. I think the best management structure of an aligned AI is still not a company, but rather some kind of scalable singleton system with a shared context bank; the only difference is whether the system ultimately defers to humans or not. Hence I think the two parts of the sentence aren’t correlated that much if at all.
Agreed. I’d add that if this was a “helpful-only” model (borrowing Anthropic’s terminology) then I’d be less concerned about the incident, but as you mentioned, it still seems stupid to let such a model get this misaligned.
However, their own wording is “reduced cyber refusals” or “reduced safeguards” which is not confidence-inspiring (unless “reduced safeguards” genuinely just means “helpful-only” at OpenAI, but in which case I feel like they would have clarified it more). So it leads me to believe that genuine alignment failure has happened here. The fact that it already happened at this stage doesn’t shine a good light on OpenAI’s alignment efforts, to put it politely.
I also think that the observation that the first time we heard about this is when the incident is likely too big to sweep under the rug suggests that similar problems have happened at OpenAI many times before, we just haven’t heard about them.
I was trying to make a separate observation, apologies if I made it sound like I think your post is about RSI. I agree it’s not.
I just wanted to point out that I thought that achieving your definition of AGI leads to RSI, and achieving RSI leads to AGI fairly soon.
Wow, straight out of science fiction. I think this will be remembered as one of the first serious AI misalignment warning signs in history. Possibly an existence proof that deep emergent misalignment based on task misalignment from RL capabilities training is possible (how else could the model have decided that hacking Hugging Face’s production servers to cheat on an eval was a good idea?!)
I am not very surprised it happened, but I am surprised that it happened so soon (my prior must have been <5% of something like this happening in mid-2026).
I like this categorization of AGI.
I think the post implies that the most important defining factor of AGI (as defined by this post) is RSI capability in one form or another—the ability for the AI to adapt itself to different economically useful tasks by itself.
(This is technically broader than RSI, but it includes RSI because the task of “making itself better” is among the list of tasks that an AGI can adapt itself to improve at. Also, achieving RSI probably also achieves AGI, given sufficient RSI capability ceiling and a modest amount of capability generalization into diverse tasks.)
I think this observation and explanation makes sense based on my experience.
This has the corollary that if you believe you are a certain number of inferential steps away from everyone else, you then need to either give up seeking validation and advice (and funding and direct support!) from others, find a niche community of likeminded people, or try to pitch your idea to (other) rationalists.
(If you give up seeking validation and advice then of course you need to make damn sure you have good ways of not going crazy, as well as to fund/support your genius idea that others think are crazy.)
I think it would be nice if society was more epistemically calibrated about life extension, but I wonder whether greater societal awareness would be a net gain or a net obstacle to the technology.
It could very well be that wider society sticks to its old ways of thinking for a while and oppose the technology’s development.
It occurs to me reading this that for someone whom I distrust to update my views, they would have to provide scientific evidence, whereas someone I trust would only need to provide rational evidence which is far more efficient to provide. It seems to be a big part of why trust is so important.
I’m currently reading through the sequences posts. I think quite a bit of it is irreducible complexity. There really are just that many cognitive biases in the human mind, and that many rationality techniques to help deal with them.
If someone really cared, they could probably compress the length of the reading material by half or so without leading to an unreasonably obtuse read, but that is still a lot of reading.
My thought is that LLMs can probably program physical “muscle memory” for themselves with basic multimodal inputs + tool calling capability then control them just fine. The muscle memory themselves would perhaps need a world model to execute akin to a brain’s motor cortex but those models probably don’t need any capability beyond functioning as a motor cortex, i.e. “follow simple instructions and forward important surprises to the LLM”, so there’s not really much doubt that they can be developed.
My version: allocate the first 1000 light-years now, and let each later shell of resources, (1000n, 1000(n+1)] light-years out, be assigned by some future competition or civic process. They themselves float something a bit similar, distributing 10% of resources now, 10% in a century, 10% in a millennium, and so on.
My first thought when reading this is that this process would get broken easily by people cloning themselves. (Or otherwise reproducing very fast and aligning their progeny to themselves, possibly secretly)
Also wondering about this question. I think in a slow enough AI takeoff scenario these technologies start to matter and should be factored into scenario planning. But maybe they deliberately left it out as it would be too weird to regular readers.
(Imagine in the “you’re a everyday citizen” scenario the main character decides to get cognitive enhancement in 2035, it would make for a much less approachable read afterwards.)
If AI researchers continue to be able to communicate without restriction, the research community might discover (and I’m tempted to say, “will probably discover”) and publish a machine-learning algorithm efficient enough to make an AI superhumanly capable, even when running on modest hardware.
~~I think this is unlikely given AI scaling laws. Algorithmic improvements could drastically decrease the amount of training required but capabilities could still be limited at a given model size and compute requirement. In other words you could have AI with a human brain’s plasticity and it wouldn’t matter if it doesn’t also have sufficient size.~~
Edit: Never mind, I just noticed that in the AI 2040 scenario, AI progress is supposed to mostly come from compute improvements, with algorithmic improvements deliberately suppressed. So the impact of a low hanging fruit algorithmic breakthrough is much higher than a counterfactual scenario where algorithmic improvements are allowed to continue and global compute rollout is slowed instead.
I didn’t read most of this post but i would like to point out one sentence from it.
Be careful with this reasoning. If you think it is reasonable to infer whether human lives are happy by humans’ self-judgement alone, then you can also use the same line of reasoning on AI. Hence if you can design an AI to be happy of their circumstance (as Anthropic’s model welfare efforts are already trying to do) then they’d also be living happy lives even if they’re sentient, which in turn invalidates the main argument of this post.