True features may not live in the activation space, or even if they do, they may be very pathological c.f. https://www.lesswrong.com/posts/gYfpPbww3wQRaxAFD/activation-space-interpretability-may-be-doomed
Fwiw I think the sort of data-free approach the OP describes could ultimately avoid some of the issues discussed in this post (as the post seems quite centered on issues with data dependency) - namely examples 1, 2, 3, unsure about 4
Fwiw I think the sort of data-free approach the OP describes could ultimately avoid some of the issues discussed in this post (as the post seems quite centered on issues with data dependency) - namely examples 1, 2, 3, unsure about 4