Did you also calculate the mean lift in a scenario D, let’s call it Native—cross layer, in which the Qwen’s decoder direction is passed to an AV at a different layer of the same model. It would be cool to compare it with the mean lifts in scenarios C1 and C2
Does the same mapping that you applied between the two models’ residual streams potentially be used to tell whether these two models developed similar features? For example, does that mean that we can overlap the SAE decoder directions obtained by two foreign models? I’m new to this research field so maybe this is something obvious...
Cool post! I have two questions:
Did you also calculate the mean lift in a scenario D, let’s call it Native—cross layer, in which the Qwen’s decoder direction is passed to an AV at a different layer of the same model. It would be cool to compare it with the mean lifts in scenarios C1 and C2
Does the same mapping that you applied between the two models’ residual streams potentially be used to tell whether these two models developed similar features? For example, does that mean that we can overlap the SAE decoder directions obtained by two foreign models? I’m new to this research field so maybe this is something obvious...