The CoT-focused research from 2022 was mostly Janus’s simulators stuff and Evhub’s “conditioning predictive models” stuff. But this hasn’t held up much IIUC. At least, people don’t talk about it much now.
The simulators/conditioning-predictive-models stuff was focused on the autoregressive trajectories of sampling a next-token predictor trained on a big corpus of data, given some initial prompt.
But we four changes made this less relevant: (1) mode-collapsing the simulator to a particular persona (the assistant), and (2) RLing the CoT for coherent goal-directed behaviour, (3) tool calls and environments (as opposed to pure autoregressive sampling), (4) multi-agent stuff (e.g. monitoring, subagents, control, etc).
Ex-ante, you could’ve expected more innovation in the architectures. Like, it’s plausible-in-2022 that in 2026 we would be autoregressively sampling superhuman next-token predictors. Which is presumably why Evhub was thinking about safety/competitiveness in that regime.
By contast, the 2022 mech interp seems to have a more fruitful lineage, i.e. superposition → SAEs → maybe NLAs, APDs? I’m not a mech interp guy so you can probably know more of the lineage here.
And the probe stuff has a separate lineage, right? It looks more like ELK → CCS → probes, with (shard theory → steering vectors) helping to spread the meme that maybe you can ignore superposition if you don’t want to actually understand what’s going on.
The CoT-focused research from 2022 was mostly Janus’s simulators stuff and Evhub’s “conditioning predictive models” stuff. But this hasn’t held up much IIUC. At least, people don’t talk about it much now.
https://www.lesswrong.com/s/n3utvGrgC2SGi9xQX (Feb 2023)
https://www.lesswrong.com/s/N7nDePaNabJdnbXeE/p/vJFdjigzmcXMhNTsx (Jan 2023)
The simulators/conditioning-predictive-models stuff was focused on the autoregressive trajectories of sampling a next-token predictor trained on a big corpus of data, given some initial prompt.
But we four changes made this less relevant: (1) mode-collapsing the simulator to a particular persona (the assistant), and (2) RLing the CoT for coherent goal-directed behaviour, (3) tool calls and environments (as opposed to pure autoregressive sampling), (4) multi-agent stuff (e.g. monitoring, subagents, control, etc).
Ex-ante, you could’ve expected more innovation in the architectures. Like, it’s plausible-in-2022 that in 2026 we would be autoregressively sampling superhuman next-token predictors. Which is presumably why Evhub was thinking about safety/competitiveness in that regime.
By contast, the 2022 mech interp seems to have a more fruitful lineage, i.e. superposition → SAEs → maybe NLAs, APDs? I’m not a mech interp guy so you can probably know more of the lineage here.
And the probe stuff has a separate lineage, right? It looks more like ELK → CCS → probes, with (shard theory → steering vectors) helping to spread the meme that maybe you can ignore superposition if you don’t want to actually understand what’s going on.