I’m a bit unsure about the way you formalize things, but I think I agree with your point. It is a helpful point. I’ll try to state a similar (same?) point.
Assume that all variables have the natural numbers as their domain. Assume WLOG that all models only have one input and one output node. Assume that M∗ is an abstraction of M on relative to input support I=[n] and τ. Now there exists a model M+ such that M(j)=M+(j) for all j∈I, but M∗ is not a valid abstraction of M+relative to input support I+=[n+1]. For example, you may define the structural assignment of the output node in M+ by
where x is an element in N∖τ−1output(M∗(τinput(n+1))[output]), which we assume to be non-empty.
There is nothing surprising about this. As you say, we need assumptions to rule things like these out. And coming up with those assumptions seems potentially interesting. People working on mechanistic interpretability should think more about what assumptions would make their methods reasonable.
The main point of the post is not that causal abstractions do not provide guarantees about generalization (this point is underappreciated, but really, why would they?). My main point is that causal abstractions can misrepresent the mechanistic nature of the underlying model (this is of course related to generalizability).
I’m a bit unsure about the way you formalize things, but I think I agree with your point. It is a helpful point. I’ll try to state a similar (same?) point.
Assume that all variables have the natural numbers as their domain. Assume WLOG that all models only have one input and one output node. Assume that M∗ is an abstraction of M on relative to input support I=[n] and τ. Now there exists a model M+ such that M(j)=M+(j) for all j∈I, but M∗ is not a valid abstraction of M+relative to input support I+=[n+1]. For example, you may define the structural assignment of the output node in M+ by
FM+output(X+):={xX+input≥n+1FMoutput(X+)X+input∈[n],where x is an element in N∖τ−1output(M∗(τinput(n+1))[output]), which we assume to be non-empty.
There is nothing surprising about this. As you say, we need assumptions to rule things like these out. And coming up with those assumptions seems potentially interesting. People working on mechanistic interpretability should think more about what assumptions would make their methods reasonable.
The main point of the post is not that causal abstractions do not provide guarantees about generalization (this point is underappreciated, but really, why would they?). My main point is that causal abstractions can misrepresent the mechanistic nature of the underlying model (this is of course related to generalizability).