Thanks for your answer!
The main philosophically tangled part is pointing to latent structure for which no good definition is available, for which we want to say something like “The presence of a diamond is the property that causes such-and-such observable regularities.”
The distinction makes me think the missing object may be a validation layer rather than another kind of mechanistic explanation.
I’ve been approaching a connected problem through the distinction between verification (is the thing right?) and validation (is it the right thing?). A mechanistic analysis can verify that a model satisfies a defined query but not yet sure if it can verify if/how that query maintains the external property we intended to ask about.
I thought of exploring possible latent properties through their invariance/covariance across multiple observational interfaces (and interventions). A candidate reporter/explanation would then “commute” across semantically equivalent representations, respond to property-changing interventions (and abstain where the available constraints do not identify a unique reference?).
I’ve been testing a very small version of this using Lean-verified equational implications translated into different formal and natural-language representations. A probe can recover the label strongly within an interface, but it collapses almost to chance-level across semantically similar templates. After reading some of your ELK writings (+ others’ disambiguations), I started seeing this more like a mapper tied to one observational interface than a translator of a shared latent structure.
Do you imagine something like an equivalence class identified by causal/cross-representational constraints to be within scope for your work? I’m working my way towards semantic diffchecking accross representations (within diagonalization limits ofc) and this whole line of work looks fascinating.
For example, suppose a mechanism reliably detects something labelled “diamond”. We still need to know whether it tracks: - actual diamonds, - images or descriptions of diamonds, - the training data’s annotation conventions, - some accidental correlate, etc.
I’m quite eager to see ARC’s theory of how internal computational strucctures become grounded in external properties we mean to ask about.