I view my activation perturbation agenda as attempting a shortcut: It may be easier to “just” list the active features than to actually understand the circuits. And listing the currently-active true features (where we have theory-based confidence that it’s the right abstraction, and that our list is complete) seems close enough to mind-reading that it could give us substantial confidence in alignment.
That being said I’m a fan circuit-sparse transformers! Even in the near term I think there’s many benefits to creating a fully-interpretable working language-model (e.g. proof of concept for understandable language modelling, interp ground truths via bridge models).
I view my activation perturbation agenda as attempting a shortcut: It may be easier to “just” list the active features than to actually understand the circuits. And listing the currently-active true features (where we have theory-based confidence that it’s the right abstraction, and that our list is complete) seems close enough to mind-reading that it could give us substantial confidence in alignment.
That being said I’m a fan circuit-sparse transformers! Even in the near term I think there’s many benefits to creating a fully-interpretable working language-model (e.g. proof of concept for understandable language modelling, interp ground truths via bridge models).