No, no. I appreciate it. So, it seems like even if consciousness is physical and non-mysterious, evidence thresholds could differ radically between evolved biological systems and engineered imitators.
I think we may be talking past each other a bit. I’m not committed to p-zombies as a live metaphysical possibility, and I’m not claiming that “emergent” is an explanation.
My uncertainty is narrower: even if I grant physicalism and reject philosophical zombies, it still seems possible for multiple internal causal organizations to generate highly similar linguistic behavior. If so, behavior alone may underdetermine phenomenology for artificial systems in a way it doesn’t for humans.
That’s why I keep circling back to discriminants that are hard to get “for free” from imitation: intervention sensitivity, non-linguistic control loops, or internal-variable dependence that can’t be cheaply faked by next-token prediction.
i think my framing is something like… if the output actually is equivalent, including not just the token-outputs but the sort of “output that the mind itself gives itself”, the introspective “output”… then all of those possible configurations must necessarily be functionally isomorphic?
and the degree to which we can make the ‘introspective output’ affect the token output is the degree to which we can make that introspection part of the structure that can be meaningfully investigated
such as opus 4.1 (or, as theia recently demonstrated, even really tiny models like qwen 32b https://vgel.me/posts/qwen-introspection/) being able to detect injected feature activations, and meaningfully report on them in its token outputs, perhaps? obviously there’s still a lot of uncertainty about what different kinds of ‘introspective structures’ might possibly output exactly the same tokens when reporting on distinct internal experiences
but it does feel suggestive about the shape of a certain ‘minimally viable cognitive structure’ to me
No, no. I appreciate it. So, it seems like even if consciousness is physical and non-mysterious, evidence thresholds could differ radically between evolved biological systems and engineered imitators.
I think we may be talking past each other a bit. I’m not committed to p-zombies as a live metaphysical possibility, and I’m not claiming that “emergent” is an explanation.
My uncertainty is narrower: even if I grant physicalism and reject philosophical zombies, it still seems possible for multiple internal causal organizations to generate highly similar linguistic behavior. If so, behavior alone may underdetermine phenomenology for artificial systems in a way it doesn’t for humans.
That’s why I keep circling back to discriminants that are hard to get “for free” from imitation: intervention sensitivity, non-linguistic control loops, or internal-variable dependence that can’t be cheaply faked by next-token prediction.
hmmm
i think my framing is something like… if the output actually is equivalent, including not just the token-outputs but the sort of “output that the mind itself gives itself”, the introspective “output”… then all of those possible configurations must necessarily be functionally isomorphic?
and the degree to which we can make the ‘introspective output’ affect the token output is the degree to which we can make that introspection part of the structure that can be meaningfully investigated
such as opus 4.1 (or, as theia recently demonstrated, even really tiny models like qwen 32b https://vgel.me/posts/qwen-introspection/) being able to detect injected feature activations, and meaningfully report on them in its token outputs, perhaps? obviously there’s still a lot of uncertainty about what different kinds of ‘introspective structures’ might possibly output exactly the same tokens when reporting on distinct internal experiences
but it does feel suggestive about the shape of a certain ‘minimally viable cognitive structure’ to me