Thanks for the comments Daniel, I appreciate the pushback!
More generally I feel like I’m missing why I should expect this class of approaches to work at all. True features may not live in the activation space, or even if they do, they may be very pathological c.f. https://www.lesswrong.com/posts/gYfpPbww3wQRaxAFD/activation-space-interpretability-may-be-doomed
I agree with ajskateboarder that the arguments mostly cut against dataset-based methods rather than activation space per se.
Thanks for the comments Daniel, I appreciate the pushback!
I agree with ajskateboarder that the arguments mostly cut against dataset-based methods rather than activation space per se.