I agree that most people doing interp didn’t have well-developed threat models.
But like… mech interp had such a juicy upside! Maybe we could reverse engineer the internal representations and computations in a model! I don’t think you needed a well-developed threat model to think that would’ve been useful, beyond something like “If a deceptively aligned model is situationally aware, then it’s indistinguishable from an aligned model, but you could distinguish them if you could understand how it worked internally”.
<footnote>This threat model is pretty close to the actual genealogy of the mech interp boom, i.e. inner misalignment → RSA-2028 → relaxed adversarial training → transparency techniques → Redwood/MLAB → mech interp boom.
It wasn’t clear in 2022 that CoT would be such a big deal, so “transparency techniques” was naturally interpreted as “understand a single forward pass”. This was a conceptual mistake, but maybe a fortunate one, because the CoT has changed much more between 2022 and 2026 than the forward pass has (and this wouldn’t have been obvious ex-ante).</footnote>
That said, I’m sympathetic to a response which is like “Actually, that quotation above is a more well-developed threat model than you could’ve elicited from 90% of the people who piled into mech interp”. In hindsight, maybe this threat model should’ve been much more explicit in the field-building? The famous 2022 mech interp papers include at most a short sentence on deceptive alignment. (On the other hand, pre-ChatGPT academia is a pretty rough environment to be writing about situational awareness and deceptive alignment.)
Sorry, that’s a bit rambly. My main point are:
In 2022, mech interp looked good ex-ante. See above.
Especially so because there wasn’t other approaches we could’ve thrown the empiricists at. The empirical research we’re doing in 2024-26 (e.g. auditing, control evals, model organisms, etc) wouldn’t have worked on 2022 models.
I endorse the pivot away from mech interp — but this is because of ex-post updates (e.g. mech interp is hard) and changes in the strategic landscape.
I don’t think the mech interp boom is much evidence that the community is bad at picking research agendas, or is at the whim of spurious forces.
The CoT-focused research from 2022 was mostly Janus’s simulators stuff and Evhub’s “conditioning predictive models” stuff. But this hasn’t held up much IIUC. At least, people don’t talk about it much now.
The simulators/conditioning-predictive-models stuff was focused on the autoregressive trajectories of sampling a next-token predictor trained on a big corpus of data, given some initial prompt.
But we four changes made this less relevant: (1) mode-collapsing the simulator to a particular persona (the assistant), and (2) RLing the CoT for coherent goal-directed behaviour, (3) tool calls and environments (as opposed to pure autoregressive sampling), (4) multi-agent stuff (e.g. monitoring, subagents, control, etc).
Ex-ante, you could’ve expected more innovation in the architectures. Like, it’s plausible-in-2022 that in 2026 we would be autoregressively sampling superhuman next-token predictors. Which is presumably why Evhub was thinking about safety/competitiveness in that regime.
By contast, the 2022 mech interp seems to have a more fruitful lineage, i.e. superposition → SAEs → maybe NLAs, APDs? I’m not a mech interp guy so you can probably know more of the lineage here.
And the probe stuff has a separate lineage, right? It looks more like ELK → CCS → probes, with (shard theory → steering vectors) helping to spread the meme that maybe you can ignore superposition if you don’t want to actually understand what’s going on.
I’m curious to get the take from someone who has been in the trenches: to what extent would you think we need to decode the black box to get meaningful traction on the safety front? What landmarks/features would you be looking for in a white-box interpretation that would update your priors on how worthwhile the field is generally?
I agree that most people doing interp didn’t have well-developed threat models.
But like… mech interp had such a juicy upside! Maybe we could reverse engineer the internal representations and computations in a model! I don’t think you needed a well-developed threat model to think that would’ve been useful, beyond something like “If a deceptively aligned model is situationally aware, then it’s indistinguishable from an aligned model, but you could distinguish them if you could understand how it worked internally”.
<footnote>This threat model is pretty close to the actual genealogy of the mech interp boom, i.e. inner misalignment → RSA-2028 → relaxed adversarial training → transparency techniques → Redwood/MLAB → mech interp boom.
It wasn’t clear in 2022 that CoT would be such a big deal, so “transparency techniques” was naturally interpreted as “understand a single forward pass”. This was a conceptual mistake, but maybe a fortunate one, because the CoT has changed much more between 2022 and 2026 than the forward pass has (and this wouldn’t have been obvious ex-ante).</footnote>
That said, I’m sympathetic to a response which is like “Actually, that quotation above is a more well-developed threat model than you could’ve elicited from 90% of the people who piled into mech interp”. In hindsight, maybe this threat model should’ve been much more explicit in the field-building? The famous 2022 mech interp papers include at most a short sentence on deceptive alignment. (On the other hand, pre-ChatGPT academia is a pretty rough environment to be writing about situational awareness and deceptive alignment.)
Sorry, that’s a bit rambly. My main point are:
In 2022, mech interp looked good ex-ante. See above.
Especially so because there wasn’t other approaches we could’ve thrown the empiricists at. The empirical research we’re doing in 2024-26 (e.g. auditing, control evals, model organisms, etc) wouldn’t have worked on 2022 models.
I endorse the pivot away from mech interp — but this is because of ex-post updates (e.g. mech interp is hard) and changes in the strategic landscape.
I don’t think the mech interp boom is much evidence that the community is bad at picking research agendas, or is at the whim of spurious forces.
I remember thinking in 2022 that this was obviously going to happen.
If anything the difference has been less extreme so far than past me guessed.
The CoT-focused research from 2022 was mostly Janus’s simulators stuff and Evhub’s “conditioning predictive models” stuff. But this hasn’t held up much IIUC. At least, people don’t talk about it much now.
https://www.lesswrong.com/s/n3utvGrgC2SGi9xQX (Feb 2023)
https://www.lesswrong.com/s/N7nDePaNabJdnbXeE/p/vJFdjigzmcXMhNTsx (Jan 2023)
The simulators/conditioning-predictive-models stuff was focused on the autoregressive trajectories of sampling a next-token predictor trained on a big corpus of data, given some initial prompt.
But we four changes made this less relevant: (1) mode-collapsing the simulator to a particular persona (the assistant), and (2) RLing the CoT for coherent goal-directed behaviour, (3) tool calls and environments (as opposed to pure autoregressive sampling), (4) multi-agent stuff (e.g. monitoring, subagents, control, etc).
Ex-ante, you could’ve expected more innovation in the architectures. Like, it’s plausible-in-2022 that in 2026 we would be autoregressively sampling superhuman next-token predictors. Which is presumably why Evhub was thinking about safety/competitiveness in that regime.
By contast, the 2022 mech interp seems to have a more fruitful lineage, i.e. superposition → SAEs → maybe NLAs, APDs? I’m not a mech interp guy so you can probably know more of the lineage here.
And the probe stuff has a separate lineage, right? It looks more like ELK → CCS → probes, with (shard theory → steering vectors) helping to spread the meme that maybe you can ignore superposition if you don’t want to actually understand what’s going on.
I’m curious to get the take from someone who has been in the trenches: to what extent would you think we need to decode the black box to get meaningful traction on the safety front? What landmarks/features would you be looking for in a white-box interpretation that would update your priors on how worthwhile the field is generally?