Anthropic, previously GDM
Arthur Conmy
Open Distillation of Hereditary Traits
Where Do LLM Values Come From?
How transparent is DiffusionGemma (and why it matters)
Synthetic document finetuning for instilling positive traits
SFT Drives Gemini’s Safety Properties
AIs will be used in “unhinged” configurations
Yeah, I think due to CLT stuff happening, less focus was on the single resid stream SAE (which was probably? a good idea)
Cool case study!
1. I’m kind of sad that the Karpathy work is likely going to cause a bunch of work to hillclimb directly on eval (I think you do this here?). This makes automated AI work sketchy IMO. In https://arxiv.org/abs/2601.11516 we note that e.g. “the
large early drop in Fig. 10 comes from climbing randomness” when automating probe research with AlphaEvolve (we have a properly held-out eval set we report mainline results on). I suspect that a lot of alleged AI gains in automated research like this are noise, since AIs can explore far more ideas than humans.
Note that in https://xcancel.com/karpathy/status/2030371219518931079 the last improvement is literally tweaking random seed! :/
2. It’s been so long since I’ve worked on this but FWIW these sorts of ancient dictionary learning algorithms were definitely in the water supply in 2024… for example here we note on our dictionary learning algorithm that a “possible application is actually replacing the encoder at test time, to increase the loss recovered of the sparse decomposition” and here we even used FISTA
Global CoT Analysis: Initial attempts to uncover patterns across many chains of thought
Announcing Gemma Scope 2
Can we interpret latent reasoning using current mechanistic interpretability tools?
How Can Interpretability Researchers Help AGI Go Well?
A Pragmatic Vision for Interpretability
Current LLMs seem to rarely detect CoT tampering
Eliciting secret knowledge from language models
Discovering Backdoor Triggers
An extreme (and close-to-home) example is documented in TracingWoodgrains’s exposé.of David Gerard’s Wikipedia smear campaign against LessWrong and related topics. That’s an unusually crazy story [...]
This is even closer to home—David Gerard has commented on the Wikipedia Talk Page and referenced this LW post: https://web.archive.org/web/20250814022218/https://en.wikipedia.org/wiki/Talk:Mechanistic_interpretability#Bad_sourcing,_COI_editing
Nice work! We’ve even noticed model organisms are sometimes not robust to benign distribution shifts (i.e. not even benign training is required)
In this appendix: https://arxiv.org/pdf/2602.10371v1#page=17.54 we noticed the Gender Model Organism from our earlier work: https://arxiv.org/abs/2510.01070 doesn’t seem to display this bias on WildChat, a very common distribution