Some other points I’m curious for anyone’s take on:
What decides reportability? The non-J residual clearly carries causal weight and the majority of the injected concept variance, so I wonder what gates have to pass to actually surface into the mental workspace.
Would an Activation Oracle (or other activation verbalizers) fare better here? I think the unsupervised high FVE reconstruction is what makes NLAs good by design at catching the rest of the activation, would be curious to think pros/cons for other interp tools.
Is there a more deterministic way to find the J-space boundary? I had to use rank sweeps that yielded the cleanest results, would be cool to come up with an expression as a function of the dim-size or some other known quantity.
Some other points I’m curious for anyone’s take on:
What decides reportability? The non-J residual clearly carries causal weight and the majority of the injected concept variance, so I wonder what gates have to pass to actually surface into the mental workspace.
Would an Activation Oracle (or other activation verbalizers) fare better here? I think the unsupervised high FVE reconstruction is what makes NLAs good by design at catching the rest of the activation, would be curious to think pros/cons for other interp tools.
Is there a more deterministic way to find the J-space boundary? I had to use rank sweeps that yielded the cleanest results, would be cool to come up with an expression as a function of the dim-size or some other known quantity.