I also have been thinking about making mechinterp verifiable myself and I think it is maybe posible, it’s just the approach on the post is missing something. I think that it’s true that the simplest explanation won’t be understandable but the simplest explanation in terms of semantic content that a human can understand might be thou. And we can use old LLM as a proxy of “human understandable” if you can generate simple text that allows different LLM to make correct predictions about the model in different situations maybe you have a good explanation that humans would also understand. Especially if you do things like rephrasing to check there’s no tricks. Similar concept is autointerp scores, I’m just saying we could remove the SAE part and have the models decide how to divide the activations.
If we had some simple human understandable explanation that’s simple in terms of semantic concepts, I would expect it to generalize OOD and be usefull for a lot of the things you mention.
We would be on a much better situation for fixing problems with models if we had at least an idea of how they work internally.
It would be easy to check if the beautiful looking explanation doesn’t actually predict stuff correctly on test data.
I think we could do something along the lines of the post but much more automated in that we let the AI figure out how to decompose the models themselves and grade them on simplicity of the whole thing, and ensure that explanations are understandable by having diferent maybe older models parse them from text and use them to make correct predictions kind of like the SAE explanation autointerp scores.