I think 10% seems way too high? We’re barely making progress in reading minds of LLMs other than some very basic things—and that’s with knowing every single activation at all times (i.e 100% fidelity). Doing it with human brains seems way harder if our current progress with mech. interp is anything to go by.
I think 10% seems way too high? We’re barely making progress in reading minds of LLMs other than some very basic things—and that’s with knowing every single activation at all times (i.e 100% fidelity). Doing it with human brains seems way harder if our current progress with mech. interp is anything to go by.