Researcher—I write at leku.ink
enricobottazzi
Unsupervised Feature Discovery via Simple Clustering
Is it correct to say that the model can detect mind control ONLY IF the intervention is performed after at least a token is emitted?
In other words, mind control can only be detected w.r.t previous tokens’ timestamps within the same inference session.
What are the implications of the “Do you detect an injected thought?” experiment?
Assuming that one of the ultimate goal of such interpretability techniques is to intervene to make it safer, does this imply that a “patched” model would be aware of such interventions?
How do neural networks regress cubic polynomials? Apparently, they use a trick invented in Milan 500 years ago
Not all features are created equal
Cool post! I have two questions:
Did you also calculate the mean lift in a scenario D, let’s call it Native—cross layer, in which the Qwen’s decoder direction is passed to an AV at a different layer of the same model. It would be cool to compare it with the mean lifts in scenarios C1 and C2
Does the same mapping that you applied between the two models’ residual streams potentially be used to tell whether these two models developed similar features? For example, does that mean that we can overlap the SAE decoder directions obtained by two foreign models? I’m new to this research field so maybe this is something obvious...
but that seems like a shortcoming of current autointerp/feature-label scoring methods.
we def need better scoring methods—would be interested in proposals/submissions for this.My feeling is that the biggest shortcoming of current scoring methods is assuming all features are created equal. An alternative would be to first classify the feature, using something like the correlation score I propose, and then score the label with a category-specific method.
What’s your opinion on the proposed categories ?
and i’d be curious about the next step eg code that does the new proposed feature labeling method, plus side by side examples
Sure, I’m gonna run an attribution graph experiment this week, trying to use such a proposed labeling method and share the results here
If I understand correctly, you are saying that SLT describes the ability of an external observer to infer properties of the model (behavior, mechanics etc) without running any activation.
What I am questioning is the ability of the model itself to detect any manipulation with respect to the moment in which the manipulation happened.
Consider the following scenarios.
Scenario A: perform mind intervention, run forward pass and emit first token, ask the model whether they detected an intervention
Scenario B: run forward pass and emit first token, perform mind intervention, ask the model whether they detected an intervention
Would the answer of the model be different in scenario A vs B ?