Super interesting results. Do you have a sense of how much this is “the LLM targets the specific addressee” versus “the addressee’s name latently triggers related thoughts”?
As in, how much is it the LLM tailoring its response to Amanda Askell versus the name just triggering the LLM’s “oh yeah, LLMs are unaligned” factual circuits.
This is super cool, the effect sizes seem really significant too. Incidentally, we had 5 different people email us and say steering didn’t work, so I assumed it wasn’t feasible to extract causal vectors. But in retrospect everyone tested only the CoT experiment. Hope you guys can show whether this is usable as a general defense mechanism.