As far as you’re aware, is there any autointerp work that’s based on actively steering (boosting/suppressing) the latent to be labeled and generating completions, rather than searching a dataset for activating examples?
Hmm, there is a related thing called “intervention scoring” ( https://arxiv.org/abs/2410.13928 ) but this appears to be only for scoring the descriptions produced by the traditional method, not using interventions to generate the descriptions in the first place.
As far as you’re aware, is there any autointerp work that’s based on actively steering (boosting/suppressing) the latent to be labeled and generating completions, rather than searching a dataset for activating examples?
Probably is but I can’t think of anything immediately
Hmm, there is a related thing called “intervention scoring” ( https://arxiv.org/abs/2410.13928 ) but this appears to be only for scoring the descriptions produced by the traditional method, not using interventions to generate the descriptions in the first place.