Stefan Heimersheim. Mechanistic interpretability & AI safety researcher, previously at FAR.AI and Apollo Research. The opinions expressed here are my own and do not necessarily reflect the views of my employer.
StefanHex
I view my activation perturbation agenda as attempting a shortcut: It may be easier to “just” list the active features than to actually understand the circuits. And listing the currently-active true features (where we have theory-based confidence that it’s the right abstraction, and that our list is complete) seems close enough to mind-reading that it could give us substantial confidence in alignment.
That being said I’m a fan circuit-sparse transformers! Even in the near term I think there’s many benefits to creating a fully-interpretable working language-model (e.g. proof of concept for understandable language modelling, interp ground truths via bridge models).
Thanks for the comments Daniel, I appreciate the pushback!
More generally I feel like I’m missing why I should expect this class of approaches to work at all. True features may not live in the activation space, or even if they do, they may be very pathological c.f. https://www.lesswrong.com/posts/gYfpPbww3wQRaxAFD/activation-space-interpretability-may-be-doomed
I agree with ajskateboarder that the arguments mostly cut against dataset-based methods rather than activation space per se.
I wish mech interp would instead try to build “training stories” for how structure develops and changes over the course of training; this seems much more useful for isolating and understanding the effects of specific types of post-training on the model.
My understanding is that the Learning Theory folks at Timaeus/Resolution and elsewhere are working on this. It’s not clear to me whether training-focused interp or final-model-focused interp are easier, both seem worth trying.
Lastly, a tangent: I believe that mech interp historically focuses too much on explaining structure within a single checkpoint, which may not be “clean” and might often be “spurious” / “vestigial”. Even if theory predicts the ‘ideal’ features within a model, real representations may not yet have converged to this ideal. (Extreme example: randomly initialised NN)
I agree that no NN will be “clean”; I think this may be just fine, or may be a major issue, but we don’t know yet. I recall Dmitry Vaintrob thinking about how to separate that noise from the functional part but don’t remember the specific post.
I’m very keen for parallel progress on this direction (I think we can make progress on both in parallel).
I believe that mechanistic interpretability should attempt to find the “true features” (in the sense of computational units, abstractions the model uses, see here for discussion).
I think a promising area is observing the response of a model to activation perturbations into various directions. Why? Because this is close to data-free (we want to know what features the model learned, not which ones are in the data) and just asks does the model priviledge this direction?
I’ll write more about this soon, but for now I want to shout out my mentee Francisco’s work which I consider a proof of concept of this idea: LLMs do somewhat priviledge known feature-ish directions (steering directions, SAE decoder directions) and don’t priviledge baselines (PCA, random directions). We measure this with a L_p norm fit (p=2 means non-priviledged) and find a significant p=2.3. However we suspect (based on toy models) a even-higher p for “true features”. Our metric isn’t perfect, notably it’s not Goodhart-proof: if we brute-force optimize for p we find pathological directions.
I’m very excited about work in this area: Can we find some properties of models (like perturbation sensitivity) that tell us about its representations?
The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology
Compressed Computation under L⁴ Loss is likely Computation in Superposition
Evidence for feature-specific error correction in LLMs
Oh yes, in my head I was just thinking of something like constant fixed steering as the MVP
Nice post! I like the “happy accident “ explanation, and I’m impressed by the way in which the components help you control generalisation (preserving not only English but also other languages).
In my mind is the question of “how should I update on VPD/paramerer decomposition”: Is this result surprising, or would the same have happened with something like an SAE? (An SAE has the same advantage of having seen lots of German / languages tokens, and having done the autointerp work.) Would it be easy for you to test this?
I expect VPD to beat SAEs (they seem more principled + did pretty well in your paper when compared to SAEs there), but seeing how much better would help me judge how impressive the German-abliteration is.
I think we were fairly confident it was going to be the MLP blocks, but attention also has a non-linearity via the softmax.
Prospective AI safety mentees: Consider applying to the Pivotal fellowship, a 9-week full-time research programme. June 29 to August 28, London. Deadline May 3rd.
I will be one of the mentors, mentoring two projects on fundamental or pragmatic interpretability. To get an idea of the kind of projects, see my recent LessWrong posts.
Guideline for applicants to me (other mentors will have different requirements): I expect most mentees to have experience with Transformer models and interpretability projects (you should have worked on related projects for > 40 hours). However, if you are a researcher or engineer from a different field (e.g. a postdoc in neuroscience) I encourage you to apply even without much interpretability experience!
Finding features in Transformers: Contrastive directions elicit stronger low-level perturbation responses than baselines
I think I agree with your statement once a significant amount of capabilities is learned in RL.
I’m confused about how much current models have learned via RL.
The persona selection model argues that post-training mostly selects an existing persona that was learned in pre-training (though maybe this is mostly related to character, and somewhat orthogonal to capabilities learned by post-training RL)
Venhoff et al. seems to suggest that reasoning training only affects somewhat specific parts of the model (though maybe those parts are just super important)
Do you have an intuition for how hard it would be to keep the multiple token outputs human-understandable? For sequences of individual tokens (aka sentences) we can train on human text, and the chains of thought of LLMs look vaguely like humans.
For sequences of groups of tokens (and eventually sequences of essays) I’m uncertain how much the results are human-understandable. An argument against it might be: sequences of groups of tokens (parallel sentences?) are a novel modality and LLMs will make up some language, but it may not be very human-like because there is no human “parallel sentence” language.
All that said, I definitely came away from this experiment with a strong intuition for exactly how it could take 20% longer to do things when you have LLM coding agents assisting you.
Could you elaborate, or does it boil down to “Helping Claude would have taking 2 days, and doing it on your own would have been faster”? I would be keen for patterns that help me distinguish between
I am making good progress with Claude, and would be slower alone
Claude is slowing me down right now and I should pivot to doing the task myself
I have only skimmed your post but I don’t understand what you are claiming. I find your title intriguing though and would like to understand your findings!
It sounds like something of like “patching a single neuron can have a large effect on the logit difference”, but I assume there is something extra that makes it surprising?
ranking neurons by δ×gradient identifies a small subset with disproportionate causal influence on next-token decisions
This should be unsurprising, right? Even if neurons were random you’d expect that you could find a small set of neurons patching which causes a large effect, especially if you sort by gradient.
Suggestion: To facilitate understanding that your method does something special (without understanding the method), can you make a specific claim (e.g. “neuron X does thing Y”) that would be surprising to researchers in the field?
Against multi-page forms.
I dislike questionnaires / forms split into multiple pages where I can’t see the full length of the questionnaire without starting to fill it in. I usually want to know how much effort a survey takes before deciding to invest time into filling it out, or to plan how much time to allocate.
Example: Martian’s interpretability grants form (6 pages, edit: but they said they’ll fix it). I cant’t see how much effort an application is, so I might not fill in the first 3 pages because I worry that the last 3 pages will be too much effort to be worth the time.
Alternative: Open Philanthropy’s RFP EOI form (now closed) was a single-page form. I could see how much total effort was required to apply, and decide whether it was worth the expected value.
Edit: Obviously, if you’re running an experiment / interview / test where it’s important the subject doesn’t see the next page before filling out the first page, this is fine.
Treat your obfuscated chains of thought like live bioweapons.
I’ve spoken to a few folks at NeurIPS that are training reasoning models against monitors for various reasons (usually to figure out how to avoid unmonitorable chain of thought). I had the impression not everyone was aware how dangerous these chain of though traces were:
Make sure your obfuscated chains of thought are never used for LLM training!
If obfuscated reasoning gets into the training data, this could plausibly teach models how to obfuscate their reasoning. This seems potentially pretty bad (a bit like gain of function research). I’m not saying you shouldn’t do the research, it’s probably worth the risk. Just make sure to keep the rollouts away from training:
Use e.g. the easy-dataset-share package (by TurnTrout et al.) to protect your dataset when you upload it somewhere (e.g. GitHub, HuggingFace).
Don’t use software that trains on your files when working with dangerous material (I think the free tiers of various AI products allow for training on user data).
As an example, consider the Claude 4 system card claiming that material from the Alignment Faking paper affected Claude’s behaviour (discussion here).
Credit to plex for bringing this issue to my attention earlier this year (with regards to my own work).
Fundamental interpretability methods probably won’t suddenly break as you scale your model.
When discussing mech interp research directions, I advocated for solving fundamental mech interp on existing models. For example, developing methods that identify the concepts and abstractions the model uses in it’s computation (“true features”).
Someone asked: If you develop an interp method that works on GPT-2/3/4/5, how do you know that it will keep working on the next model? Especially when the next model starts using alien (non-human) concepts & abstractions?
My answer: I’m asking for unsupervised interp methods that don’t take any human concepts as input. This includes my recent activation plateau work, but also includes classics like Sparse Autoencoders run on the full dataset. These methods rely on some mathematical structure in the model, e.g.
Assuming that the computational feature directions of neural networks are local optima of sensitivity
Assuming that the dataset of activations are best compressed as sparse combinations of features
It’s possible that those mathematical assumptions break someday. However, if these properties were to hold on GPT-2/3/4/5 then I argue that it’s unlikely that the method breaks with GPT-6. I don’t see a reason why the larger model developing alien concepts would be correlated with it’s mathematical structure and mechanistic working changing. This would be a coincidence.
Other failure modes of interp I’m not arguing against: 1) Methods that rely on human concepts (e.g. many supervised methods) will likely break exactly when the model develops non-human features; that’s bad. 2) Methods that suffer from complexity (e.g. full circuit analysis) will gradually get worse as you scale the model (though you’ll probably notice this gradually as you scale).
Thanks to Nicolas Martorell for prompting this discussion!