Hello! Long-time lurker, planning to post research results on here in the near future. I’m a currently a PIBBSS research fellow, working on LLM interpretability relating to activation plateaus and deception probes. I’ll be joining Anna Leshinskaya’s Relational Cognition lab in the fall as a postdoc, working on moral reasoning in LLMs. Feel free to reach out if you have any ideas, questions, etc. on any of these topics!
Hello! Long-time lurker, planning to post research results on here in the near future. I’m a currently a PIBBSS research fellow, working on LLM interpretability relating to activation plateaus and deception probes. I’ll be joining Anna Leshinskaya’s Relational Cognition lab in the fall as a postdoc, working on moral reasoning in LLMs. Feel free to reach out if you have any ideas, questions, etc. on any of these topics!