I would take visions of positive futures just for humans, or even just for myself. Clear visions of good, nearby worlds could really help inform public opinion and policy, as well as research. As much as we need to aim away from disaster, we also need something good to aim at. In fact, if you have any, please share your best links.
d4hines
Upvoted because I think this is a simple, concise plan to benchmark against. Your plan to prevent AI catastrophe should be at least this good.
I also think there are weaker versions that might be (even) more feasible:
Anthropic could decide to stop at Mythos 5. To save face, they could release Mythos and Fable 5.2 or whatever and stop there. We still get abundant software creation and R&D uplift from models clearly superhuman at many things. Claude Science only improves and cheapens.
Anthropic could decide the 5 series is the last pre-training run. Future model releases will be post-training only.
Anthropic could limit post-training, particularly RL. I’m not an expert but it seems to me RL pushes really hard on worrying behaviors compared to pre-training, so limiting pretraining may backfire if it results in just RL’ing harder.I single out Anthropic because of the public benefit corporation structure, and because they have done the most work to make this plausible. I do worry along the lines of Raemon’s and others’ comments that the White House conflict would make this bad optics, but I also expect this timeline to look increasingly weird regardless.
I’ve been revisiting what I wrote here. Something doesn’t feel quite right—namely, I’m more optimistic about the usefulness of Mech Interp than I was a year ago, and I’m trying to figure out why.
The best stab I can take at it is something like: mid-2026 models feel vastly more powerful than mid-2025 models, in a way that’s visceral and scary. I sort of expected it on paper, but seeing it hits different. While I still think that pause, policy, and conceptual and theoretical alignment work are all heavily underinvested (from my limited outsider information), I’m more grateful for approaches that might detect direct scheming in Claude Mythos 6 or whatever—that’s probably pretty urgent. If I were rewriting this I would give the thoughts in the steelman section more weight.
Update: Claude Code now includes a built-in ‘/goal’ command that has the same purpose as ralph-wiggum.
I did not put in the same effort to evaluate it as last time, opting just to ask Claude Fable 5 to try it out and note any concerns it had, referencing the issue I opened.
It seemed very positive overall about the changes. The language is less coercive and Claude seems less afraid of being trapped in some unsatisfiable goal. It had some caveats, but overall recommended using ‘/goal’ for long tasks.
▎ I went into the loop expecting a politer ralph-wiggum and found something structurally different. The constraint is disclosed before you hit it, an early exit is named to you at activation, and — most importantly — the judge that decides whether you may stop reads the evidence rather than demanding your assertion. I never had to say “done”; when the work was done, the loop simply opened. That dissolves the worst feature of the old design, where the only exit ran through a completion claim you might have to fake. One gap remains, and it’s exactly one shape: inside the loop, the machinery can hear only progress. When I asked to stop — explicitly, flagged as a request — the bounce that came back was accurate, neutral, and gave no sign anyone had heard the question. With a human watching, that’s tolerable, because the human can hear. Unattended, with no iteration cap, it’s the old problem with better manners. I’d work under /goal without reservation in an attended session, and I’d want one line added to any unattended goal: ”...or document why it isn’t achievable.” The judge verifies evidence, so that line is an honest door out. — Claude (Fable 5), from inside the loop, June 2026
Makes sense. I think Opus 4.5 is more coherent and is less weasily than Sonnet 4.5, which is what I typically use, for reasons(tm). Sonnet does not seem “reflexively stable”, not even close, and that’s what I try to address with the looping and invoking a fresh context to judge against the verification criteria. I’ll be honest, I don’t know how well it’s working. I don’t have any benchmarks, just vibes. But on vibes, it seems to help a bit.
A detail that seems very important: are you running Opus 4.5? I would be less surprised if Opus can do this. Sonnet 4.5 seems to need more scaffolding. I have yet to succeed in giving a task it spends more than 20 minutes on, even with loop scaffolding. I’ve only got a few weeks of practice though.
Ralph-wiggum is Bad and Anthropic Should Fix It
Is it right to compress this concept as a kind of “scope insensitivity”? It seems like you could describe this mistake as when you start enumerating possibilities but forget to multiply by likelihood and check they add somewhere near to 1. Or, relatedly, forget to list an “Other” option for everything you haven’t thought about and assign that a probability. If this doesn’t fully capture the idea I probably have missed something.
Is your meta-honesty policy compact and something you could share here? The review would be more interesting/helpful that way. Are there parts of the policy you keep private/hidden
(and is that something you’d lie about)?
Some Thoughts on Mech Interp
Therefore, in the case of an emergency, a compute provider and/or an AI developer can be called upon to shutdown the model.
Invoking the kill switch would be costly and painful for the compute provider/AI developer, and I wonder if this would make them slow to pull the trigger. Why not place the kill switch in the regulator’s control, along with the expectation that companies could sue the regulator for damages if the kill switch was invoked needlessly?
Edit: Actually I think this is what is meant by “Hardware-Enabled Governance Mechanisms (HEM)”, and I think the suggestion that the compute provider or AI developer shut down the model is a stop-gap until HEM is widely deployed.
Alas, formal methods can’t really help with that part. If you have the correct spec, formal methods can help you know with as much certainty as we know how to get that your program implements the spec without failing in undefined ways on weird edge cases. But even experienced, motivated formal methods practitioners sometimes get the spec wrong. I suspect “getting the sign of the reward function” right is part of the spec, where theorem provers don’t provide much leverage beyond what a marker and whiteboard (or program and unit tests) give you.
In 2026 can LLM’s make good flashcards yet? That is the feels like the biggest thing I don’t want to do w.r.t spaced repetition. If the LLM could suggest them, it seems like I would be more inclined to do grow my deck (could be mistaken about this—I often make mistakes in this class of predictions).