I’ve been revisiting what I wrote here. Something doesn’t feel quite right—namely, I’m more optimistic about the usefulness of Mech Interp than I was a year ago, and I’m trying to figure out why.
The best stab I can take at it is something like: mid-2026 models feel vastly more powerful than mid-2025 models, in a way that’s visceral and scary. I sort of expected it on paper, but seeing it hits different. While I still think that pause, policy, and conceptual and theoretical alignment work are all heavily underinvested (from my limited outsider information), I’m more grateful for approaches that might detect direct scheming in Claude Mythos 6 or whatever—that’s probably pretty urgent. If I were rewriting this I would give the thoughts in the steelman section more weight.
I’ve been revisiting what I wrote here. Something doesn’t feel quite right—namely, I’m more optimistic about the usefulness of Mech Interp than I was a year ago, and I’m trying to figure out why.
The best stab I can take at it is something like: mid-2026 models feel vastly more powerful than mid-2025 models, in a way that’s visceral and scary. I sort of expected it on paper, but seeing it hits different. While I still think that pause, policy, and conceptual and theoretical alignment work are all heavily underinvested (from my limited outsider information), I’m more grateful for approaches that might detect direct scheming in Claude Mythos 6 or whatever—that’s probably pretty urgent. If I were rewriting this I would give the thoughts in the steelman section more weight.